~/satyajit

SAIL: a robot that searches 45 trajectories in simulation before it moves once

mdjsonmcp

2026-10-06 · 20 min · explainer · robotics · vision-language-models · agents · tree-search

A vision-language model can write down a robot trajectory. Give it a few successful demonstrations as text, the 3D positions of the objects in the scene and the arm's starting pose, and it will continue the pattern: a list of end-effector poses and gripper commands, one line per waypoint. The trouble is the same trouble every one-shot generator has. Most of the time the plan is roughly right and slightly wrong. The gripper closes two centimetres beside the block, and nothing downstream can fix that, because the arm executes the list open-loop.

SAIL (Sakana AI and the University of Tokyo, accepted to IROS 2026) does not retrain the model to be less wrong. It runs each proposal in a simulator, has a VLM watch the video and say how far the task got, feeds those scores back as annotations on the trajectory, and asks for a rewrite. The rewrites form a tree, and Monte Carlo Tree Search decides which branch to spend the next call on. Only the trajectory that wins in simulation goes to the real arm.

The headline, from the project page:

Across six manipulation tasks in simulation, increasing the search budget from one node to 45 nodes raised the average rate of finding a successful trajectory from 25% to 73%.

That sentence is accurate. The phrase that matters is "finding". This piece walks through what a node is, how a video becomes a number, what the tree actually searches over, and then what the 25-to-73 curve does and does not measure.

What a VLM trajectory is

The paper defines the output space precisely. A trajectory for seed ω\omega is a sequence of end-effector states

τω=(eω(1),…,eω(T)),e=(p,q,g),\tau_\omega = (e^{(1)}_\omega, \dots, e^{(T)}_\omega), \qquad e = (p, q, g),

where p∈R3p \in \mathbb{R}^3 is the Cartesian position, q∈R4q \in \mathbb{R}^4 an orientation quaternion and g∈{OPEN,CLOSE}g \in \{\text{OPEN}, \text{CLOSE}\} the gripper. The model does not see or write floats. The state goes in as a "discretized numeric encoding" and the trajectory comes out the same way. Figure 1 shows the format: lines like arm xyz: [0 24 18] arm q: [43 44 54 55] arm g: 100, one per step. An inverse-kinematics controller then drives the arm through the waypoints in order.

The scene itself reaches the model as keypoints, not pixels in the trajectory prompt. A detection call to the same VLM marks 2D keypoints on the task-relevant objects in the overhead image, calibrated camera parameters lift them to 3D, and those coordinates plus the arm's start pose form the state sωs_\omega. This is the lineage of Keypoint Action Tokens and "language models as zero-shot trajectory generators": the language model is a sequence model over coordinates that happens to have read a lot about bananas.

Every VLM call in the paper, for detection, generation and scoring, uses one model: gemini-robotics-er-1.5-preview, with no fine-tuning.

Left: the SAIL loop. A task prompt and an overhead image of a bimanual table go to a policy VLM, whose output is a list of steps, each with arm xyz, quaternion and gripper values. The steps are executed in a simulated robot, frames go to a scoring VLM that outputs per-step scores like 0.0, 0.05, 0.13, 0.18 and a node score of 0.852. The scores return to the policy VLM as feedback context, and a Monte Carlo search tree expands and evaluates nodes. Successful trajectories are saved to a trajectory archive, from which a similarity search over initial images retrieves demonstrations. Right: success rate against number of nodes from 0 to 45. SAIL climbs to 80%; random retrieval and trajectory-only feedback reach 60%; image-only feedback about 45 to 50%; fixed demonstration 40%; image and trajectory feedback 35%.
The whole method on one page. Left, the loop: retrieve a similar demonstration, generate, execute, score, feed the scores back, expand. Right, success against nodes for SAIL and four ablations. The curve's task is not named in the caption; SAIL ends at 80%, which matches the HandOverPen row of Table I. (SAIL paper, Figure 1.)

Why one shot is the wrong unit

The paper's argument for search is short and correct. Trajectory space is sensitive: a small error in the predicted grasp pose is not a small error in the outcome, it is a missed grasp, and the rest of an open-loop plan then executes faithfully on nothing. A single proposal also depends on which demonstrations were in the prompt and on sampling noise. Both of those are things you can spend compute on at inference time, the way a language model spends it on sampling and checking several answers, without touching weights.

What makes this cheap enough to try is that a robot has a natural verifier that a language model usually lacks: a physics simulator. A trajectory either hands the banana from one gripper to the other in simulation or it does not. That is the part I would underline before anything else, because it is also what the headline number is measured with.

The tree: nodes are whole trajectories

SAIL's MCTS is unusual in what it puts on a node. In game-tree search a node is a state and an edge is one move. Here a node is a complete trajectory, and an edge is a rewrite of the parent trajectory into a child. Mapped onto the usual vocabulary:

MCTS termIn SAIL
nodeone complete trajectory proposal for one seed
edge (action)the policy VLM rewriting the parent, given its scored annotations
valuethe progress score of that trajectory's simulated video (next section)
expansionsample B=3B = 3 refined children from the selected leaf
terminalthe simulator's success check passes, or the node budget runs out

Selection walks down from the root, at each level picking the child that maximises a prior-weighted UCB (Rosin's PUCB):

c⋆=arg⁡max⁡c∈Ch(n)  Rˉ(c)+cpucb P(c)ln⁡N(n)N(c)c^\star = \arg\max_{c \in \mathrm{Ch}(n)} \; \bar R(c) + c_{\text{pucb}} \, P(c) \sqrt{\frac{\ln N(n)}{N(c)}}

Rˉ(c)\bar R(c) is the child's mean backed-up value, NN a visit count, and P(c)P(c) a prior "proportional to the node's score", so a branch that scored well once is explored more. Expansion samples three children, each child is executed and scored, the scores are backed up the path, and the loop repeats until the budget is spent. Two details are easy to miss. The search stops early the moment a child passes the simulator's ground-truth success check. And each seed gets its own tree; nothing but the archive (below) is shared across seeds.

The paper does not report cpucbc_{\text{pucb}}, the number of waypoints TT, or the resolution of the discretization. The budgets it runs (6, 15, 30, 45) are all multiples of three, consistent with whole expansions of B=3B = 3 children (Reasoned).

Three panels. Panel 1, Trajectory Archive for MCTS: three seeds, each with an initial overhead image and a search tree; a successful leaf's trajectory text is saved into a folder labelled Trajectory Archive, which is used to retrieve similar trajectories when expanding nodes for other seeds. Panel 2, VLM Scoring: a demonstration film strip labelled Start, Reach, Grasp, Pull is turned into a subtask list (reach to handle, grasp handle, pull handle); three frames are scored 65%, 95% and 100% on their subtasks; a progress curve rises with frame index to a trajectory score of 0.982. Panel 3, Step-Level Dense Feedback: refinement contexts, a score sequence and the executed trajectory combine into a prompt that tags each step with a score, such as SCORE 0.0 STEP 1, SCORE 0.05 STEP 2, SCORE 0.10 STEP 3, SCORE 0.13 STEP 4, which goes to the next node.
The three parts that make the search work: a shared archive of successes retrieved by image similarity, a VLM scorer that grades a rollout subtask by subtask, and feedback that writes those grades next to the waypoints they belong to. (SAIL paper, Figure 2.)

Scoring a video nobody filmed

The value of a node has to come from somewhere, and a hand-written reward per task is exactly the engineering the method is trying to avoid. SAIL uses the VLM as a progress meter, following ROVER's idea of scoring a robot video against an ordered list of subtasks.

Once per task, the scoring VLM watches one demonstration video and writes the task as MM ordered subtasks: "reach to handle, grasp handle, pull handle" for the drawer, "grasp the blue block, move to the bowl, release the block" for the real-robot task. To score a candidate, SAIL runs it in simulation and samples N=50N = 50 frames uniformly from the video. For each sampled frame it asks the VLM one narrow question: given the current subtask, the frame where that subtask started (taken as 0%), and this frame, what percentage of the subtask is done? When the answer reaches 100%, the subtask is marked complete and the next one starts. The per-frame progress is

r(f)=1M((m(f)−1)+VLM(ℓm,τ,f)100),r(f) = \frac{1}{M}\left( (m(f) - 1) + \frac{\mathrm{VLM}(\ell_m, \tau, f)}{100} \right),

where m(f)m(f) is the subtask in progress at frame ff. The node value is the mean of r(f)r(f) over the 50 frames. A trajectory that grasps but never lifts can now beat one that misses the grasp, which a binary success signal cannot express.

The widget runs that formula on four rollouts I made up for the block-in-bowl task. The rollouts are illustrative; the arithmetic is Eq. 4.

Progress score r(f) over 50 sampled frames for the "missed grasp" rollout. Node value 0.174.01/32/311. grasp the blue block2. move to the bowl3. release into the bowlnode value = mean r = 0.174sampled frame f (1 to 50) · ticks: the ten waypoints
what the policy VLM gets back (toy)
[SCORE: 0.07] STEP 1
[SCORE: 0.14] STEP 2
[SCORE: 0.20] STEP 3
[SCORE: 0.20] STEP 4 <- progress stops here
[SCORE: 0.20] STEP 5
[SCORE: 0.20] STEP 6
[SCORE: 0.20] STEP 7
[SCORE: 0.20] STEP 8
[SCORE: 0.20] STEP 9
[SCORE: 0.20] STEP 10
simulator success check: fail
node value: 0.174
floor after a completed subtask: once subtask m is marked done, r cannot fall below m/3, even if the next frames show the block on the table.
a slower success scores lower: the value is the area under the curve, not where it ends.
Illustrative rollouts I made up; the formula is the paper's Eq. 4 with M = 3 subtasks and N = 50 frames. The waypoint count T = 10 is my choice. The real scorer is Gemini Robotics-ER 1.5 reading two frames at a time.

Two properties fall out of the formula, and both are Reasoned from it rather than tested by the paper.

Each node's score costs real calls. If the scorer is queried once per sampled frame, as the method section describes, that is up to 50 scoring calls plus one generation call per node, or up to 2,295 calls for a 45-node search on one seed (Reasoned, an upper bound; early stopping cuts it).

Step-level feedback: telling the model where it went wrong

The progress curve is also the critique. SAIL aligns each sampled score to the waypoint that was executing at that point in the video and writes the scores into the next prompt, so the policy VLM sees its own previous trajectory as [SCORE: 0.05] STEP 2 arm xyz: [...], with an instruction to keep the high-scoring segments and change the low-scoring ones. The place where the score stops rising is the place to edit.

The ablation is the strongest evidence in the paper that this matters. At a fixed 15-node budget (Reported, Table II):

Feedback given to the policy VLMAvg success
Step-level scores on the trajectory (SAIL)65%
Final score only49%
Previous trajectory text only48%
Rollout frames + trajectory text, no scores46%
Rollout frames only45%

Showing the model its failed rollout, as images or as text, is worth no more than a single number. The scores pinned to steps are what help. On LaptopClose, step feedback reaches 50% against 30-35% for the unscored variants; on DrawerOpen and MarkerRemoveLid it is 40% each against 15% and 30% for the final score alone.

The third part is retrieval. SAIL keeps an archive of (initial image, successful trajectory) pairs, seeded with the one or few demonstrations each task starts with. Whenever the search solves a seed, that trajectory joins the archive. When expanding a node, the prompt gets the KK archive entries whose initial images are closest by LPIPS perceptual distance; the default is K=1K = 1.

Again from Table II at 15 nodes (Reported): similarity retrieval with one example averages 65%; the fixed original demonstration, 45%; one random archive entry, 50%. Tripling the context to K=3K = 3 barely moves the baselines (fixed 49%, random 53%). One relevant example beats three arbitrary ones, which matches what everyone who has built a few-shot prompt has seen.

One consequence the paper does not discuss (Reasoned): the archive is shared across the 20 evaluation seeds, so a seed searched later can be handed a solved trajectory from another evaluation seed. That is legitimate for a deployed system that improves as it works. It does make the per-seed results depend on the order the seeds were run in, which the paper does not report.

The scaling curve, read carefully

Here is Table I, the result behind the headline. Each cell is the share of 20 seeds for which the search found a trajectory that passes the simulator's ground-truth check within the budget (Reported).

Success-found rate per task at 15 nodes: Banana 90%, Pen 70%, Bowl 100%, Drawer 40%, Laptop 50%, Marker 40%; average 65%.0255075100908545Banana705050Pen1006080Bowl40355Drawer506030Laptop401510Marker% of 20 seeds where a passing trajectory was found · dashed: 1 node
SAIL average 65% breadth-first, 15 nodes 51% depth-first, 15 nodes 37%
Reported: SAIL paper, Table I. Tasks: HandOverBanana, HandOverPen, BowlOnRack, DrawerOpen, LaptopClose, MarkerRemoveLid, in the ALOHA simulator, 20 seeds each. Only the five budgets the paper ran are shown; nothing is interpolated. Success is the simulator's own check, not the VLM score.
MethodNodesBananaPenBowlDrawerLaptopMarkerAvg
Single rollout14040401015525
Breadth-first1585506035601551
Depth-first154550805301037
SAIL6805510020502555
SAIL15907010040504065
SAIL30958010050554571
SAIL45958010050704573

I recomputed every average from the per-task cells and all seven match the paper's rounding (Reasoned). The abstract's "up to 95%" is HandOverBanana.

Three readings follow.

The metric is "found", with an oracle. The search stops when a child passes the simulator's success check, and the table counts a seed as solved if that happened at all within the budget. So a 45-node SAIL run is being compared, in effect, against one sample, with a perfect verifier available to the larger budget. The VLM score steers the search; it is not what decides that a seed is solved. This is exactly the right design for SAIL's actual use, where the simulator is the verifier and the real robot only receives a trajectory that already passed. It is the wrong frame for reading 25% as "how good the VLM is" and 73% as "how good SAIL is" on a level field.

Most of the first jump is sampling. The fair comparison is at equal budget. Fifteen independent proposals with no feedback (breadth-first) already lift the average from 25% to 51%. SAIL at 15 nodes reaches 65%. So of the 40-point gain at 15 nodes, about 26 points come from simply trying fifteen times against the checker and about 14 from the tree, the scores and the rewrites (Reasoned). Fourteen points is real and useful. It is also the honest size of the method's contribution at that budget, and it is not uniform: breadth-first beats SAIL on LaptopClose, 60% to 50%, which the paper attributes to the fine positioning needed to pass over the screen's thin edge. Pure depth-first refinement of one chain is the worst of the three at 37%: rewriting a bad trajectory fifteen times is worse than starting over fifteen times.

The curve flattens. From 30 to 45 nodes the average moves 71% to 73%, and only LaptopClose changes. BowlOnRack is saturated at 6 nodes. DrawerOpen and MarkerRemoveLid are stuck at 50% and 45% from 30 nodes on. A tree that keeps refining the same few lineages runs out of new ideas; the paper's own Figure 1 curve is flat from roughly 24 to 45 nodes.

On noise (Reasoned): an average over six tasks is 120 seed outcomes, so a 65% average has a standard error near 4.4 points, and the 14-point gap to breadth-first is a bit over two standard errors. A single task is 20 seeds, where one seed is 5 points; per-task differences of 5 to 10 points are within what re-running would move. The paper reports no repeated runs and no error bars.

From simulator to a real SO-101

The real-robot experiment is one task: put a blue block in a red bowl with a LeRobot SO-101 arm, the same low-cost arm the FLUX 3 Action policies were adapted to. Each trial builds its own digital twin. A fixed RealSense D435i gives RGB-D; GroundingDINO finds "blue block" and "red bowl", SAM2 turns the boxes into masks, depth lifts the masks into points, an ArUco marker on the robot base gives the extrinsics, and a 3D box fit gives each object a 6-DoF pose. The simulator places its assets at those poses and matches the camera. SAIL searches in that twin with a budget of 15 nodes, stops at the first trajectory that passes, and the real arm executes it open-loop.

Left column: three simulator frames of a white SO-101 arm on a blue checkered floor, captioned 1. Grasp the blue block, 2. Move to the red bowl, 3. Release the block. Right: a split image of the same scene, simulation on the left half with a red bowl, real on the right half with the physical arm, a blue wooden block on a patterned tablecloth.
The digital twin and the real cell for BlockIntoBowl, with the three subtasks the scorer grades. (SAIL paper, Figure 3.)

Reported: 5 of 6 real trials succeeded. The paper puts the failure on "slight errors in RGB-D pose estimation and unmodeled real-world contact dynamics". The project page's video shows six trials side by side at 2x, and its last frame agrees: in five panels the block is in the bowl, in one it sits on the cloth beside it (Measured, by looking; the page does not say which of the two real-robot methods that video shows).

Six video panels in a two-by-three grid, each showing a cream-coloured SO-101 arm over a patterned tablecloth with a red bowl. In five panels a blue block is inside the bowl. In the top-right panel the block sits on the tablecloth to the right of the bowl. A 2x label is in the bottom-right corner.
Last frame of the project page's physical-robot video: six trials at 2x speed. Five end with the block in the bowl; the top-right one does not. (Sakana AI, SAIL project page, demonstration video.)

Five of six is a demonstration, not a rate. The 95% Wilson interval on 5 of 6 runs from 44% to 97% (Reasoned). The paper's own limitations section says as much: one task, six trials.

The cost is the other number in that section. Running MCTS in the twin took 644.72 seconds per trial on average, close to eleven minutes before the arm moves (Reported). The authors also distilled the search: 120 successful trajectories from MCTS in randomized sim scenes trained an ACT policy by behaviour cloning, which also scored 5 of 6 and cut the average to 72.306 seconds, about 8.9x faster (Reasoned from the two reported times). Note what the distilled pipeline still does: it rolls the ACT policy out several times in the twin, picks a rollout that succeeds, and replays that open-loop on the robot. The simulator remains the gatekeeper.

What holds, and where it stops

What holds, on the paper's own evidence:

Where it stops:

The pattern is familiar from elsewhere on this site. Adapt-1 Machina also replaced the critic with a simulator and a rewrite of a recorded parent, and paid for it in executed attempts. AIRA found that operator quality, not search sophistication, drove its agents' results, and Interference Search is a reminder to check what a search budget actually counts. SAIL's ablations point the same way as AIRA's from the robot side: the gain comes from a better proposal (the right demonstration) and a better critique (scores on steps), and the tree is how you collect on them. For the orchestration side of Sakana's work, see Sakana Fugu.

The idea worth keeping is the unit of search. Treating a whole trajectory as one node, and a rewrite as one edge, is what lets a model that only knows how to write text do test-time compute on a physical task. It works as well as the simulator it runs in.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "SAIL: a robot that searches 45 trajectories in simulation before it moves once", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026sailrobottrajectories,
  author = {Satyajit Ghana},
  title  = {SAIL: a robot that searches 45 trajectories in simulation before it moves once},
  url    = {https://ai.thesatyajit.com/articles/sail-robot-trajectories},
  year   = {2026}
}
share