~/satyajit

Ego2Act and EgoTools: two egocentric benchmarks that grade state, not pixels

mdjsonmcp

2026-10-06 · 22 min · benchmarks · evaluation · video-generation · world-models · vision-language-models · datasets

Pour milk from a sealed carton into a mug that is sitting upside down. A person does it in about twenty seconds and never thinks about the order: turn the mug over, open the carton, pour. A video model given a photo of that table and the sentence "The cup contains some milk from the sealed milk box" will very often pour the milk onto the bottom of the inverted mug, and then, a few frames later, the mug is upright and full.

Two benchmarks that landed in the same week are about that failure. Ego2Act (arXiv 2610.01092) grades video generators on goal-directed, first-person manipulation. EgoTools (arXiv 2609.39378) grades video understanding models on first-person tool use. Both end up measuring the same thing: whether a model tracks the state of objects through a sequence of hand actions, rather than whether its frames look right.

I read both papers, pulled Ego2Act's released score tables and judge traces from Hugging Face, and recomputed its headline numbers. Every number below is labelled: measured (I computed it from a released file), reported (the authors' figure, not re-run), or reasoned (my arithmetic on the other two).

Ego2Act: the task

Each case is a pair: a real egocentric start frame f1f_1 and a goal GG that names an end state and nothing else. The model must generate a video V={f1,…,fT}V = \{f_1, \dots, f_T\} that reaches GG using only objects visible in f1f_1. Every model gets the same prompt template, which adds three constraints: one hand, only objects in the starting frame, no teleportation or duplication. There is no step list. Planning the steps is the point.

The benchmark has 110 cases (reported; the cases table in ego2act/ego2act-bench has 110 rows, measured). Each case was recorded by people three times correctly, with at least three distinct action sequences, and three times incorrectly as negative controls. Six generators produce three seeds per case (101, 202 and 303). The paper totals 660 human videos plus 1,980 generated ones, 2,640 in all (reported). The release is slightly smaller: human/metadata.csv indexes 616 human recordings (311 correct, 305 wrong) and the generated-video repo holds 1,974 clips (both measured). The dataset card says most tasks have three of each and some fewer, so 660 is the design, not the shipped count.

Tasks average 5.17 observed steps in the successful human recordings, and the median correct recording runs 20.3 s (both measured from the released tables; the paper reports a 20.4 s median). The five domains split 27 kitchen, 26 household, 25 office, 21 personal care and 11 other (measured).

An initial scene with an open travel case, a toothpaste tube, its cap and a folding toothbrush, the goal text, and three generated attempts shown as five-frame strips with green checks for valid states, blue crosses for task errors such as a brush left unfolded, and purple crosses for physics errors such as a new organizer appearing.
One Ego2Act case, three generated attempts. Each row follows a different but valid step order; each fails somewhere, on the task (blue) or on physics (purple) (Ego2Act paper, Figure 1).

Because any valid step order counts, scoring has to be reference-free, and that is the more interesting half of the paper.

The rubric: gates, not a holistic score

A rater, human or model, first decomposes GG into subgoals g1,…,gNg_1, \dots, g_N, each with one source object, at most one target and one action. Enabling actions get their own subgoal: "open the carton" is a subgoal even though the goal never mentions it.

Each subgoal then climbs two ladders and stops at the first rung it fails.

Then

T=1003N∑i=1Nti,P=1004 ∣J∣∑i∈Jpi,S=T⋅PT = \frac{100}{3N}\sum_{i=1}^{N} t_i, \qquad P = \frac{100}{4\,|\mathcal{J}|}\sum_{i \in \mathcal{J}} p_i, \qquad S = \sqrt{T \cdot P}

where J\mathcal{J} is the set of subgoals that were actually attempted (ti>0t_i \gt 0) and inspectable. A video that attempts nothing has no PP and gets S=0S = 0.

The gating is the design choice the paper defends hardest. The authors rescored their pilot human annotations with gates removed or asked independently and compared each variant against holistic human ratings. Dropping P1 moves physics MAE from 0.432 to 0.922 score points; asking all physics gates independently gives 0.777 (reported, Figure 5 below). Inside the automated judge, asking every gate independently inflates Final scores by 6.8 points and lowers Final rr from 0.69 to 0.65 (reported).

Two panels. Left: case construction, where a goal string, an initial scene and human videos pass task-validity and visual-quality checks to become an accepted case with metadata and three successful plus three unsuccessful human attempts. Right: the judge decomposes the goal into subgoals, applies task gates T1 to T3 and physics gates P1 to P4 per subgoal with satisfied, failed and not-reached marks, and combines a task score and a physics score into the Ego2Act score.
Case construction (left) and Ego2ActJudge (right): subgoal decomposition, then sequential task and physics gates per subgoal (Ego2Act paper, Figure 3).

The geometric mean is chosen so that a zero on either axis zeroes the video. It has a side effect the paper states in its appendix: because PP averages only over attempted subgoals, a rollout that skips the hard interactions can keep a high PP. The released judge traces show this on the milk task. Step through it.

ego2actjudge · gate trace · pour_milkreleased traces
Task gatesPhysics gatesT1T2T3P1P2P3P4S1 turn cup upright✗––n/a–––S2 open the carton✓✓✓✗–––S3 pour milk into cup✓✓✗✗–––
gates shown7/7

P4 persistence: does the result stay put as forces allow?

S1 not reached

S2 not reached

S3 not reached

T = mean(t) / 3
levels 0, 3, 2
55.6
P = mean(p) / 4, attempted only
levels n/a, 0, 0
0.0
S = √(T · P)
final
0.0

Wan 2.7 attempts all three subgoals and scores 0, because its two attempted interactions both break continuity at P1. MiniMax H3 only turns the cup over, never touches the milk, and scores 57.7: Physics is averaged over attempted subgoals only, so the one clean interaction carries a perfect P.

Here are the three rollouts in numbers (all measured from traces/ego2act_judge/traces.jsonl). The judge planned the same three subgoals for each: turn the cup upright, open the carton, pour.

RolloutTask levelsPhysics levelsTTPPSS
Seedance 2.0, seed 1013, 3, 34, 4, 4100.0100.0100.0
MiniMax H3, seed 1013, 0, 04, n/a, n/a33.3100.057.7
Wan 2.7, seed 1010, 3, 2n/a, 0, 055.60.00.0

Wan 2.7 attempts every subgoal, skips the prerequisite (the cup is never turned over), and then breaks continuity twice: the carton grows a yellow spout between frames, and around 8.6 s the inverted cup becomes an upright cup full of milk. It scores 0. MiniMax H3 turns the cup over cleanly, never touches the milk, and scores 57.7. Neither video completes the goal, and the rubric ranks the one that did less above the one that did more. On a leaderboard, a generator that learns to do less, cleanly, is rewarded. (Reasoned, from the traces and the rule in the paper's Appendix B.2.)

Four frames from a generated video: an upside-down mug next to a sealed milk carton and a bent straw, a hand holding the carton now with a yellow spout, milk pouring onto the base of the inverted mug, and a close-up of milk pooled in the mug's base.
Wan 2.7, seed 101, on the milk task: the mug is never turned over, and the milk is poured onto its base (Ego2Act paper, Figure 12e).

Ego2ActJudge

The automated judge runs that same rubric through a VLM, Gemini 3.7 Flash, with no fine-tuning and no reference video. Before it sees the video it writes two plans from GG and f1f_1: task subgoals with prerequisites, and physics interaction windows. It then works through the gates. Each axis starts from frames sampled uniformly over the video, 8 for Task and 24 for Physics. At any gate it may call an inspection tool that returns 12 frames from a time window and region it names. Every answer is yes, no or unresolved, with frame IDs as evidence. An unresolved answer withholds the axis score instead of guessing (all reported, Appendix B.5). The paper quotes about US$0.04 per video in practice (reported).

The plan is re-derived for every video, which is a source of noise. 91% of Task plans use the case's most common subgoal count (reported). On 358 videos judged twice with fresh plans, the Task score moves 5.2 points when both plans have the same count and 11.5 to 11.9 points when they don't (reported). The test-retest correlation of the Final score is r=0.89r = 0.89, with a single-run standard deviation of 10.4 points (reported). The authors' own advice follows: use the judge to compare models over hundreds of videos, not to rule on one clip.

What 0.69 actually is

The thread that announced Ego2Act says the judge "reaches best agreement with human consensus (0.69) with human-human agreement at over 0.84". Those are two different statistics.

The like-for-like human number is in the paper's Table 21: each individual rating against the consensus of the other raters on the same video gives r=0.76r = 0.76, CCC 0.76, MAE 15.5 and bias −0.3-0.3 (reported; my leave-one-rater-out recomputation gives r=0.756r = 0.756 over 486 ratings, measured). So the judge's per-video correlation is 0.69 against a human ceiling of about 0.76, not 0.84. Its error, though, is 22.0 points against the humans' 15.5, and its bias is +15.4+15.4 against roughly zero. The judge is close to people on ordering and loose on level.

Two more checks on the 0.69. Cosmos 3 Nano scores near zero on almost everything, which stretches the range. Drop it and rr falls to 0.48 (reported; I get 0.483 over 368 videos, measured). Among the six generators the judge's ranking matches the humans' with Spearman ρ=0.94\rho = 0.94, one adjacent swap (reported; measured 0.943 from the per-model means on the panel videos).

ego2act · human raters vs ego2actjudgereported, 0-100
|sort|
0255075100human raters (25-case panel)Ego2ActJudge (all 110 cases)Seedance 2.0+19.3Kling 3.0 Pro+12.9MiniMax H3+4.0Grok Imagine 1.5+19.6Wan 2.7+27.1Cosmos 3 Nano+9.4Human, correct−4.6Human, wrong−3.3

right column: judge minus human. Mean over the six generators: +15.4 points (final, different video sets).

The judge sits right of the humans on almost every row, and furthest right on Physics. It gets the order nearly right and the level wrong. MiniMax H3 and Cosmos 3 Nano are the two generators the judge scores below the humans on Task, and on the full benchmark it drops from third to fifth on Final.

The comparison in that chart, in numbers:

GeneratorHuman Task / Physics / Final (panel)Judge Task / Physics / Final (all 110 cases)Judge Final, panel videos
Seedance 2.067.8 / 63.0 / 64.079.9 / 91.1 / 83.388.0
Kling 3.0 Pro59.3 / 64.2 / 59.664.0 / 88.9 / 72.576.7
MiniMax H361.2 / 55.8 / 55.353.7 / 76.2 / 59.370.7
Grok Imagine 1.556.6 / 47.9 / 49.967.8 / 79.4 / 69.572.9
Wan 2.755.5 / 41.4 / 44.367.8 / 80.5 / 71.468.6
Cosmos 3 Nano16.9 / 5.2 / 3.915.4 / 38.8 / 13.314.4
Human, wrong attempts59.6 / 94.1 / 73.457.8 / 91.6 / 70.1

All reported (paper Tables 2 and 9). The human Final column reproduces exactly from the released ratings: 64.0, 59.6, 55.3, 49.9, 44.3 and 3.9 (measured). It only does so after zero-filling the Final of videos whose mean Task is 0, as the paper's rule says. Without that, MiniMax H3 reads 56.3 and Cosmos 4.1 (measured). Note that the Final average is not T⋅P\sqrt{T \cdot P} of the averages: Seedance's 67.8×63.0=65.4\sqrt{67.8 \times 63.0} = 65.4, not 64.0 (reasoned), because Final is computed per video and then averaged.

Three things in the release are worth knowing before you quote this leaderboard.

  1. The human panel is thinner than "three annotators score 600 videos" suggests. Of 600 rows, 75 are correct human recordings that were never rated: they receive the rubric maximum by construction. Another 77 generated videos have no rating at all. Only 44 videos carry three ratings (all measured).
  2. MiniMax H3's human score rests on single ratings. 56 of its 75 panel videos were rated, every one by exactly one rater. Every other generator has 35 to 41 videos with two or more ratings (measured). The ICC and the leave-one-out human agreement are computed on videos with at least two ratings, so they include no MiniMax H3 video. Its 55.3, third place, is the least cross-checked number on the board. The judge, for its part, drops it from third to fifth on the full benchmark.
  3. Panel size. The paper says the panel is 25 cases. The released table holds 600 rows across 26 case IDs, and the cases table flags 26 cases in_human_panel (measured). The panel is five cases per domain plus a sixth kitchen case. That changes nothing above, but 25 cases is the paper's sample size for every bootstrap interval.

Where the judge is blind

The paper's best section is the one where it attacks its own judge. The authors took 30 correct human recordings that the judge had scored 100 on Physics and corrupted each one in the most active 3 s. Teleport deletes 1.5 s of frames. Swap reverses two adjacent 1.5 s segments. Ghost blends in the frame from 3 s earlier at 50% opacity, so objects appear doubled. Detection rates were 6.9% for teleport, 10.3% for swap and 55.2% for ghost, with zero false alarms on unedited re-encodes (reported).

The pattern is mechanical. A ghost is visible inside one frame. A teleport or swap only exists between frames, and the Physics pass starts from 24 evenly spaced frames, so a 1.5 s cut looks like an ordinary sampling gap. The judge zoomed into the edited window in only 7% (teleport) and 13% (swap) of videos. When it did inspect the ghost window, it caught the duplication in 10 of 13 (reported). On the human panel, the judge scores Physics 24.1 points above people on average and Task only 8.1 above (reported; measured 24.1 and 8.1).

Two bar charts of mean absolute error in score points. Task: full rubric 0.292, no T1 0.497, no T2 0.324, no T3 0.331, independent gates 0.306. Physics: full rubric 0.432, no P1 0.922, no P2 0.456, no P3 0.444, no P4 0.438, independent gates 0.777.
Rubric ablation on pilot human annotations: removing the first gate, or asking gates independently, raises the error against holistic human ratings, most of all for Physics (Ego2Act paper, Figure 5).

What the generators get wrong

The failure taxonomy is the most useful output of the benchmark, because it says what a world model is missing.

Six panels of before-and-after frame pairs with red circles. Task errors: cutting a chocolate bar with the wrapper still on, milk poured onto an inverted cup, whiteboard letters left after wiping, a toothbrush left unfolded, nesting dolls left outside, a refill placed in a cover. Physics errors: one plush doll becoming two, an e-reader turning into a pouch, a toy passing through a closed lid, coffee beans entering through a closed lid, a laptop screen pivoting at its side, a doll gaining an invented hinge.
Recurring failures: skipped steps, incomplete actions and domain-specific operations (task); object inconsistencies, passing through closed lids, and wrong mechanisms (physics) (Ego2Act paper, Figure 11).

Read together: today's best generators render a convincing hand and convincing objects without carrying a state variable for "the mug is upside down". When the goal needs that state, the model hallucinates its way to the end frame. Seedance 2.0 at 64.0 is the best anyone does.

A caveat on Cosmos 3 Nano's 3.9. The paper's text calls it "Cosmos 3", while the project page, the thread and the dataset card say the Nano variant. It was run from open weights on a local server, like MiniMax H3, while Seedance, Kling, Wan and Grok were called through OpenRouter (reported, dataset card). The judge could not score Physics on 36.4% of its videos because nothing was attempted (reported). A score of 3.9 says this model mostly does not start the task. It does not say how the larger Cosmos 3 would fare. Note also that MiniMax H3 here is the open-weights release I covered in MiniMax H3: open weights, four excluded countries, zero benchmarks. Ego2Act is, as far as I know, the first third-party number on it.

EgoTools: the understanding side

EgoTools turns the camera around. Instead of asking a model to generate a hand using a tool, it shows a model real first-person footage of tool use and asks questions about it.

The corpus, EgoTools-Data, is 646 videos and 100.37 hours across seven domains: kitchen, classroom, research lab, repair workshop, craft, office and household. It was recorded on a head-mounted rig with four synchronized fisheye cameras plus audio and motion signals, and canonicalized to 1024 × 1024 at 20 FPS (reported). On top sit Gemini-3-Flash hierarchical captions, tool-centric narrations recorded by the people who did the tasks, and a 3D layer from Gaussian-splatting reconstruction plus SAM2 object tracks (reported). The counts are 361,332 captions and 6,519 narrations (reported).

An illustrated campus map with seven domains, each with its video count and hours: kitchen 276 videos and 32.37 hours, research lab 61 videos and 30.92 hours, repair workshop 84 videos, craft 67, classroom 74, household 54, office 30. Sample questions surround it, and bottom panels show hierarchical temporal captions for making coffee, long-sequence 4D object tracking, and a donut chart of the four benchmark tracks.
EgoTools at a glance: seven tool-use domains, hierarchical captions, 4D object tracking, and the four-track 1,000-question benchmark (EgoTools paper, Figure 1).

EgoTools-Bench is 1,000 eight-way multiple-choice questions, so chance is 12.5%. 900 are human-written and 100 are generated from the 3D annotations and then human-verified. They are drawn from a 40.34-hour pool held out at the source-video level from all training data, and the average question spans 4.19 minutes of video (reported). The four tracks are:

Questions were hardened against text-only shortcuts. An independent check of 250 items found 89.8% agreement with the answer key (reported).

Left: a radial chart of the 1,000 questions split into four tracks and their subtracks, such as why, what and when questions. Right: a caption, a narration and an eight-option question asking why the person switches from scissors to a knife while trimming asparagus, plus word clouds comparing EPIC-KITCHENS-100, Ego4D and EgoTools-Bench.
Benchmark composition, an annotation example, and a vocabulary comparison with EPIC-KITCHENS-100 and Ego4D (EgoTools paper, Figure 2).

The results, all reported (paper Table 2; AC, PG, PD and SR are the four tracks):

ModelFramesACPGPDSROverall
Human expertfree viewing82.185.283.382.783.2
Gemini-3.1-Pro1 fps69.751.772.174.966.9
Gemini-3-Flash1 fps64.550.073.072.664.4
GLM4.1V-Thinking6446.150.057.153.350.7
Qwen3-VL-8B-Instruct6446.145.957.155.450.0
EgoTools-8B (Qwen3-VL-8B + SFT)6455.459.768.144.460.9

Perception & Grounding is the weak track for every strong model. Gemini-3.1-Pro is 18.0 to 23.2 points lower there than on its other three tracks (reasoned from the table). The question it fails is which visually similar tool is in the hand and what it touched, not why someone would use it. The same Qwen3-VL-8B scores 69.0% on EgoSchema and 50.0% here (reported). Recognizing the activity does not mean recognizing the mechanism.

Fine-tuning on 184,679 examples from the training pool lifts Qwen3-VL-8B from 50.0% to 60.9% (reported). Spatial Reasoning drops 11.0 points, though, and EgoPlan-Bench drops 7.1. The paper reports both and scopes its claim accordingly.

Two caveats that change how far I would lean on these numbers.

Language priors carry most of the score. In the paper's input ablation, Qwen3-VL-8B gets 51.7% with 64 ordered frames and 42.7% with no video at all (reported). Shuffling the frames costs only 2.4 points, and a single middle frame costs 4.8. Text alone recovers about 83% of the full-video score (reasoned: 42.7 / 51.7). Video adds 9.0 points overall and 13.9 on Perception & Grounding (reported). The benchmark does need vision, but less than the "requires visual evidence" heading suggests, and the order of frames barely matters.

The released numbers do not all agree, and the benchmark is not public yet. The paper's error-overlap figure labels its models with accuracies that differ from Table 2: Gemini 3.1 Pro 67.9% against 66.9%, and Qwen3-VL-8B Instruct 40.7% against 50.0% (measured by reading both). The text says all results were recomputed on the quality-controlled benchmark, so the figure is likely from an earlier version. On the release side, the repo's own docs/DATA.md says the public dataset holds 131 recording episodes and no benchmark question table or SFT data yet. The ropedia-ai/egotools-8b checkpoint is gated behind manual approval on Hugging Face (measured from the Hub API, 2026-10-06). The README says the complete paper experiments "have not yet been reproduced with the released resources". So the EgoTools numbers above are reported and not re-run, and for now nobody outside the team can re-run them.

A five-by-five heatmap of Jaccard overlap between the wrong answers of Gemini 3.1 Pro, Gemini 3 Flash, Gemini Flash Lite, Qwen3-VL-8B Instruct and Qwen3-VL-8B Thinking. The two Qwen variants overlap at 0.63, Gemini Pro and Flash at 0.53, and Gemini against Qwen around 0.35 to 0.38. Row labels show accuracies of 67.9, 64.7, 56.3, 40.7 and 41.7 percent.
Overlap of wrong answers between models. Thinking variants keep most of their instruct counterpart's errors. Note the row accuracies differ from the paper's main results table (EgoTools paper, Figure 6).

The same blind spot, from both ends

Put the two papers side by side and a loop closes. Ego2Act's judge is a VLM. EgoTools shows that frontier VLMs are weakest at exactly the perception it needs: which object, in what state, touched by what. Gemini-3.1-Pro manages 51.7% on that track. Ego2Act's negative controls show the same weakness from the other side. The judge sees a duplicated object inside a frame half the time, and a 1.5 s cut between frames about one time in ten. The tool that grades world models inherits the perception ceiling of the models that EgoTools grades. (Reasoned; the two papers use different Gemini versions, 3.7 Flash as the judge and 3.1 Pro and 3 Flash as test subjects, so this is a direction, not a measurement.)

The failure underneath is the same in both: carrying object state through time. A generator that pours onto an inverted mug has no "mug is upside down" variable. A VLM that cannot say which of two pliers is in the hand has not bound the tool to its trajectory. Both get credit from plausible-looking frames and lose it on dependencies. That is why first-person video is the right test bed. Hands occlude everything, the camera never stops moving, and every task is a chain of state changes.

What I would take from each:

For video world models used directly as robot policies, where skipping a prerequisite costs a real grasp, see FLUX 3 Action.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Ego2Act and EgoTools: two egocentric benchmarks that grade state, not pixels", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026egocentricvideobenchmarks,
  author = {Satyajit Ghana},
  title  = {Ego2Act and EgoTools: two egocentric benchmarks that grade state, not pixels},
  url    = {https://ai.thesatyajit.com/articles/egocentric-video-benchmarks},
  year   = {2026}
}
share