2026-10-06 · 22 min · benchmarks · evaluation · video-generation · world-models · vision-language-models · datasets
Pour milk from a sealed carton into a mug that is sitting upside down. A person does it in about twenty seconds and never thinks about the order: turn the mug over, open the carton, pour. A video model given a photo of that table and the sentence "The cup contains some milk from the sealed milk box" will very often pour the milk onto the bottom of the inverted mug, and then, a few frames later, the mug is upright and full.
Two benchmarks that landed in the same week are about that failure. Ego2Act (arXiv 2610.01092) grades video generators on goal-directed, first-person manipulation. EgoTools (arXiv 2609.39378) grades video understanding models on first-person tool use. Both end up measuring the same thing: whether a model tracks the state of objects through a sequence of hand actions, rather than whether its frames look right.
I read both papers, pulled Ego2Act's released score tables and judge traces from Hugging Face, and recomputed its headline numbers. Every number below is labelled: measured (I computed it from a released file), reported (the authors' figure, not re-run), or reasoned (my arithmetic on the other two).
Ego2Act: the task
Each case is a pair: a real egocentric start frame and a goal that names an end state and nothing else. The model must generate a video that reaches using only objects visible in . Every model gets the same prompt template, which adds three constraints: one hand, only objects in the starting frame, no teleportation or duplication. There is no step list. Planning the steps is the point.
The benchmark has 110 cases (reported; the cases table in ego2act/ego2act-bench has 110 rows, measured). Each case was recorded by people three times correctly, with at least three distinct action sequences, and three times incorrectly as negative controls. Six generators produce three seeds per case (101, 202 and 303). The paper totals 660 human videos plus 1,980 generated ones, 2,640 in all (reported). The release is slightly smaller: human/metadata.csv indexes 616 human recordings (311 correct, 305 wrong) and the generated-video repo holds 1,974 clips (both measured). The dataset card says most tasks have three of each and some fewer, so 660 is the design, not the shipped count.
Tasks average 5.17 observed steps in the successful human recordings, and the median correct recording runs 20.3 s (both measured from the released tables; the paper reports a 20.4 s median). The five domains split 27 kitchen, 26 household, 25 office, 21 personal care and 11 other (measured).

Because any valid step order counts, scoring has to be reference-free, and that is the more interesting half of the paper.
The rubric: gates, not a holistic score
A rater, human or model, first decomposes into subgoals , each with one source object, at most one target and one action. Enabling actions get their own subgoal: "open the carton" is a subgoal even though the goal never mentions it.
Each subgoal then climbs two ladders and stops at the first rung it fails.
- Task gates. T1 intent: did a directed attempt on the right object begin? T2 process: was the action carried out as a recognisable trajectory? T3 end state: does the subgoal end in the required state? The task level is the number of gates passed.
- Physics gates. P1 continuity: do objects keep identity, count and a traceable path? P2 causation: does visible contact precede each response? P3 interaction: is the interaction physically possible while it happens? P4 persistence: does the result stay put? The physics level is .
Then
where is the set of subgoals that were actually attempted () and inspectable. A video that attempts nothing has no and gets .
The gating is the design choice the paper defends hardest. The authors rescored their pilot human annotations with gates removed or asked independently and compared each variant against holistic human ratings. Dropping P1 moves physics MAE from 0.432 to 0.922 score points; asking all physics gates independently gives 0.777 (reported, Figure 5 below). Inside the automated judge, asking every gate independently inflates Final scores by 6.8 points and lowers Final from 0.69 to 0.65 (reported).

The geometric mean is chosen so that a zero on either axis zeroes the video. It has a side effect the paper states in its appendix: because averages only over attempted subgoals, a rollout that skips the hard interactions can keep a high . The released judge traces show this on the milk task. Step through it.
P4 persistence: does the result stay put as forces allow?
S1 not reached
S2 not reached
S3 not reached
Wan 2.7 attempts all three subgoals and scores 0, because its two attempted interactions both break continuity at P1. MiniMax H3 only turns the cup over, never touches the milk, and scores 57.7: Physics is averaged over attempted subgoals only, so the one clean interaction carries a perfect P.
Here are the three rollouts in numbers (all measured from traces/ego2act_judge/traces.jsonl). The judge planned the same three subgoals for each: turn the cup upright, open the carton, pour.
| Rollout | Task levels | Physics levels | |||
|---|---|---|---|---|---|
| Seedance 2.0, seed 101 | 3, 3, 3 | 4, 4, 4 | 100.0 | 100.0 | 100.0 |
| MiniMax H3, seed 101 | 3, 0, 0 | 4, n/a, n/a | 33.3 | 100.0 | 57.7 |
| Wan 2.7, seed 101 | 0, 3, 2 | n/a, 0, 0 | 55.6 | 0.0 | 0.0 |
Wan 2.7 attempts every subgoal, skips the prerequisite (the cup is never turned over), and then breaks continuity twice: the carton grows a yellow spout between frames, and around 8.6 s the inverted cup becomes an upright cup full of milk. It scores 0. MiniMax H3 turns the cup over cleanly, never touches the milk, and scores 57.7. Neither video completes the goal, and the rubric ranks the one that did less above the one that did more. On a leaderboard, a generator that learns to do less, cleanly, is rewarded. (Reasoned, from the traces and the rule in the paper's Appendix B.2.)

Ego2ActJudge
The automated judge runs that same rubric through a VLM, Gemini 3.7 Flash, with no fine-tuning and no reference video. Before it sees the video it writes two plans from and : task subgoals with prerequisites, and physics interaction windows. It then works through the gates. Each axis starts from frames sampled uniformly over the video, 8 for Task and 24 for Physics. At any gate it may call an inspection tool that returns 12 frames from a time window and region it names. Every answer is yes, no or unresolved, with frame IDs as evidence. An unresolved answer withholds the axis score instead of guessing (all reported, Appendix B.5). The paper quotes about US$0.04 per video in practice (reported).
The plan is re-derived for every video, which is a source of noise. 91% of Task plans use the case's most common subgoal count (reported). On 358 videos judged twice with fresh plans, the Task score moves 5.2 points when both plans have the same count and 11.5 to 11.9 points when they don't (reported). The test-retest correlation of the Final score is , with a single-run standard deviation of 10.4 points (reported). The authors' own advice follows: use the judge to compare models over hundreds of videos, not to rule on one clip.
What 0.69 actually is
The thread that announced Ego2Act says the judge "reaches best agreement with human consensus (0.69) with human-human agreement at over 0.84". Those are two different statistics.
- 0.69 is a Pearson correlation, per video, between the judge's Final score and the human consensus Final score on the human-rated panel. The consensus is the mean of the available ratings on each axis, combined as . I recomputed it from the released
human_ratings.parquetandjudge_scores.parquet. After applying the paper's rule that a video with mean Task 0 scores , the 432 videos with both a human and a judge Final give , Kendall , Lin's CCC , MAE and a mean signed bias of (all measured; the paper reports 0.69, 0.46, 0.61, 22.0 and +15.4). - 0.84 and up are intraclass correlations among the human raters: 0.863 on Final, 0.831 on Task and 0.834 on Physics (reported). The paper uses a one-way random-effects ICC of the average rating. That measures how reliable a mean of several raters is, not how well one rater agrees with another, and averaging alone pushes it up.
The like-for-like human number is in the paper's Table 21: each individual rating against the consensus of the other raters on the same video gives , CCC 0.76, MAE 15.5 and bias (reported; my leave-one-rater-out recomputation gives over 486 ratings, measured). So the judge's per-video correlation is 0.69 against a human ceiling of about 0.76, not 0.84. Its error, though, is 22.0 points against the humans' 15.5, and its bias is against roughly zero. The judge is close to people on ordering and loose on level.
Two more checks on the 0.69. Cosmos 3 Nano scores near zero on almost everything, which stretches the range. Drop it and falls to 0.48 (reported; I get 0.483 over 368 videos, measured). Among the six generators the judge's ranking matches the humans' with Spearman , one adjacent swap (reported; measured 0.943 from the per-model means on the panel videos).
right column: judge minus human. Mean over the six generators: +15.4 points (final, different video sets).
The judge sits right of the humans on almost every row, and furthest right on Physics. It gets the order nearly right and the level wrong. MiniMax H3 and Cosmos 3 Nano are the two generators the judge scores below the humans on Task, and on the full benchmark it drops from third to fifth on Final.
The comparison in that chart, in numbers:
| Generator | Human Task / Physics / Final (panel) | Judge Task / Physics / Final (all 110 cases) | Judge Final, panel videos |
|---|---|---|---|
| Seedance 2.0 | 67.8 / 63.0 / 64.0 | 79.9 / 91.1 / 83.3 | 88.0 |
| Kling 3.0 Pro | 59.3 / 64.2 / 59.6 | 64.0 / 88.9 / 72.5 | 76.7 |
| MiniMax H3 | 61.2 / 55.8 / 55.3 | 53.7 / 76.2 / 59.3 | 70.7 |
| Grok Imagine 1.5 | 56.6 / 47.9 / 49.9 | 67.8 / 79.4 / 69.5 | 72.9 |
| Wan 2.7 | 55.5 / 41.4 / 44.3 | 67.8 / 80.5 / 71.4 | 68.6 |
| Cosmos 3 Nano | 16.9 / 5.2 / 3.9 | 15.4 / 38.8 / 13.3 | 14.4 |
| Human, wrong attempts | 59.6 / 94.1 / 73.4 | 57.8 / 91.6 / 70.1 |
All reported (paper Tables 2 and 9). The human Final column reproduces exactly from the released ratings: 64.0, 59.6, 55.3, 49.9, 44.3 and 3.9 (measured). It only does so after zero-filling the Final of videos whose mean Task is 0, as the paper's rule says. Without that, MiniMax H3 reads 56.3 and Cosmos 4.1 (measured). Note that the Final average is not of the averages: Seedance's , not 64.0 (reasoned), because Final is computed per video and then averaged.
Three things in the release are worth knowing before you quote this leaderboard.
- The human panel is thinner than "three annotators score 600 videos" suggests. Of 600 rows, 75 are correct human recordings that were never rated: they receive the rubric maximum by construction. Another 77 generated videos have no rating at all. Only 44 videos carry three ratings (all measured).
- MiniMax H3's human score rests on single ratings. 56 of its 75 panel videos were rated, every one by exactly one rater. Every other generator has 35 to 41 videos with two or more ratings (measured). The ICC and the leave-one-out human agreement are computed on videos with at least two ratings, so they include no MiniMax H3 video. Its 55.3, third place, is the least cross-checked number on the board. The judge, for its part, drops it from third to fifth on the full benchmark.
- Panel size. The paper says the panel is 25 cases. The released table holds 600 rows across 26 case IDs, and the
casestable flags 26 casesin_human_panel(measured). The panel is five cases per domain plus a sixth kitchen case. That changes nothing above, but 25 cases is the paper's sample size for every bootstrap interval.
Where the judge is blind
The paper's best section is the one where it attacks its own judge. The authors took 30 correct human recordings that the judge had scored 100 on Physics and corrupted each one in the most active 3 s. Teleport deletes 1.5 s of frames. Swap reverses two adjacent 1.5 s segments. Ghost blends in the frame from 3 s earlier at 50% opacity, so objects appear doubled. Detection rates were 6.9% for teleport, 10.3% for swap and 55.2% for ghost, with zero false alarms on unedited re-encodes (reported).
The pattern is mechanical. A ghost is visible inside one frame. A teleport or swap only exists between frames, and the Physics pass starts from 24 evenly spaced frames, so a 1.5 s cut looks like an ordinary sampling gap. The judge zoomed into the edited window in only 7% (teleport) and 13% (swap) of videos. When it did inspect the ghost window, it caught the duplication in 10 of 13 (reported). On the human panel, the judge scores Physics 24.1 points above people on average and Task only 8.1 above (reported; measured 24.1 and 8.1).

What the generators get wrong
The failure taxonomy is the most useful output of the benchmark, because it says what a world model is missing.

- Skipped prerequisites. Cut the chocolate with the wrapper on, insert into a closed container, pour before the cup is upright. Across the six generators, 88.7% of failed Task judgments are level 0 (the step never happened) or level 2 (it happened but missed its end state). The range per generator is 85.4% to 91.6% (reported, on 294 case-seed units).
- Contact-rich actions are the bottleneck. At the subgoal level, attachment and connection reach 38.6% Task completion and 66.7% Physics validity, against 51.1% and 75.5% for relocation (reported). Task length correlates only weakly with score, Spearman (reported): short tasks still fail.
- Physics shortcuts follow task shortcuts. If the lid was never opened, the beans go through it. The skipped step and the boundary violation are the same failure seen from two axes.
Read together: today's best generators render a convincing hand and convincing objects without carrying a state variable for "the mug is upside down". When the goal needs that state, the model hallucinates its way to the end frame. Seedance 2.0 at 64.0 is the best anyone does.
A caveat on Cosmos 3 Nano's 3.9. The paper's text calls it "Cosmos 3", while the project page, the thread and the dataset card say the Nano variant. It was run from open weights on a local server, like MiniMax H3, while Seedance, Kling, Wan and Grok were called through OpenRouter (reported, dataset card). The judge could not score Physics on 36.4% of its videos because nothing was attempted (reported). A score of 3.9 says this model mostly does not start the task. It does not say how the larger Cosmos 3 would fare. Note also that MiniMax H3 here is the open-weights release I covered in MiniMax H3: open weights, four excluded countries, zero benchmarks. Ego2Act is, as far as I know, the first third-party number on it.
EgoTools: the understanding side
EgoTools turns the camera around. Instead of asking a model to generate a hand using a tool, it shows a model real first-person footage of tool use and asks questions about it.
The corpus, EgoTools-Data, is 646 videos and 100.37 hours across seven domains: kitchen, classroom, research lab, repair workshop, craft, office and household. It was recorded on a head-mounted rig with four synchronized fisheye cameras plus audio and motion signals, and canonicalized to 1024 × 1024 at 20 FPS (reported). On top sit Gemini-3-Flash hierarchical captions, tool-centric narrations recorded by the people who did the tasks, and a 3D layer from Gaussian-splatting reconstruction plus SAM2 object tracks (reported). The counts are 361,332 captions and 6,519 narrations (reported).

EgoTools-Bench is 1,000 eight-way multiple-choice questions, so chance is 12.5%. 900 are human-written and 100 are generated from the 3D annotations and then human-verified. They are drawn from a 40.34-hour pool held out at the source-video level from all training data, and the average question spans 4.19 minutes of video (reported). The four tracks are:
- Affordance & Causality (363 questions): why this tool, what it affords, what an action caused.
- Perception & Grounding (236): which tool, which object, what state, how many.
- Procedural Dynamics (222): order, transitions, staging.
- Spatial Reasoning (179): hand-tool-object geometry.
Questions were hardened against text-only shortcuts. An independent check of 250 items found 89.8% agreement with the answer key (reported).

The results, all reported (paper Table 2; AC, PG, PD and SR are the four tracks):
| Model | Frames | AC | PG | PD | SR | Overall |
|---|---|---|---|---|---|---|
| Human expert | free viewing | 82.1 | 85.2 | 83.3 | 82.7 | 83.2 |
| Gemini-3.1-Pro | 1 fps | 69.7 | 51.7 | 72.1 | 74.9 | 66.9 |
| Gemini-3-Flash | 1 fps | 64.5 | 50.0 | 73.0 | 72.6 | 64.4 |
| GLM4.1V-Thinking | 64 | 46.1 | 50.0 | 57.1 | 53.3 | 50.7 |
| Qwen3-VL-8B-Instruct | 64 | 46.1 | 45.9 | 57.1 | 55.4 | 50.0 |
| EgoTools-8B (Qwen3-VL-8B + SFT) | 64 | 55.4 | 59.7 | 68.1 | 44.4 | 60.9 |
Perception & Grounding is the weak track for every strong model. Gemini-3.1-Pro is 18.0 to 23.2 points lower there than on its other three tracks (reasoned from the table). The question it fails is which visually similar tool is in the hand and what it touched, not why someone would use it. The same Qwen3-VL-8B scores 69.0% on EgoSchema and 50.0% here (reported). Recognizing the activity does not mean recognizing the mechanism.
Fine-tuning on 184,679 examples from the training pool lifts Qwen3-VL-8B from 50.0% to 60.9% (reported). Spatial Reasoning drops 11.0 points, though, and EgoPlan-Bench drops 7.1. The paper reports both and scopes its claim accordingly.
Two caveats that change how far I would lean on these numbers.
Language priors carry most of the score. In the paper's input ablation, Qwen3-VL-8B gets 51.7% with 64 ordered frames and 42.7% with no video at all (reported). Shuffling the frames costs only 2.4 points, and a single middle frame costs 4.8. Text alone recovers about 83% of the full-video score (reasoned: 42.7 / 51.7). Video adds 9.0 points overall and 13.9 on Perception & Grounding (reported). The benchmark does need vision, but less than the "requires visual evidence" heading suggests, and the order of frames barely matters.
The released numbers do not all agree, and the benchmark is not public yet. The paper's error-overlap figure labels its models with accuracies that differ from Table 2: Gemini 3.1 Pro 67.9% against 66.9%, and Qwen3-VL-8B Instruct 40.7% against 50.0% (measured by reading both). The text says all results were recomputed on the quality-controlled benchmark, so the figure is likely from an earlier version. On the release side, the repo's own docs/DATA.md says the public dataset holds 131 recording episodes and no benchmark question table or SFT data yet. The ropedia-ai/egotools-8b checkpoint is gated behind manual approval on Hugging Face (measured from the Hub API, 2026-10-06). The README says the complete paper experiments "have not yet been reproduced with the released resources". So the EgoTools numbers above are reported and not re-run, and for now nobody outside the team can re-run them.

The same blind spot, from both ends
Put the two papers side by side and a loop closes. Ego2Act's judge is a VLM. EgoTools shows that frontier VLMs are weakest at exactly the perception it needs: which object, in what state, touched by what. Gemini-3.1-Pro manages 51.7% on that track. Ego2Act's negative controls show the same weakness from the other side. The judge sees a duplicated object inside a frame half the time, and a 1.5 s cut between frames about one time in ten. The tool that grades world models inherits the perception ceiling of the models that EgoTools grades. (Reasoned; the two papers use different Gemini versions, 3.7 Flash as the judge and 3.1 Pro and 3 Flash as test subjects, so this is a direction, not a measurement.)
The failure underneath is the same in both: carrying object state through time. A generator that pours onto an inverted mug has no "mug is upside down" variable. A VLM that cannot say which of two pliers is in the hand has not bound the tool to its trajectory. Both get credit from plausible-looking frames and lose it on dependencies. That is why first-person video is the right test bed. Hands occlude everything, the camera never stops moving, and every task is a chain of state changes.
What I would take from each:
- Use Ego2Act's rubric, read its leaderboard with care. The gated T/P decomposition and the released traces are the contribution. They let you see why a rollout scored what it did, which no single-number video metric gives you. Treat the judge as a ranking tool, as its authors do: rank correlation 0.94 across models, per-video MAE 22 points. Watch the attempted-only Physics average. On the milk task it scores 57.7 for a video that did one of three subgoals and 0 for one that tried all three.
- Use EgoTools-Bench's Perception & Grounding track as the diagnostic, once it ships. It is the track where vision moves the score most and where frontier models score lowest. And report a text-only baseline next to any number you publish on it.
For video world models used directly as robot policies, where skipping a prerequisite costs a real grasp, see FLUX 3 Action.