~/satyajit

H-JEPA: the top of the stack forgets the ant's legs, and that is the point

mdjsonmcp

2026-10-06 · 26 min · world-models · self-supervised-learning · representation-learning · robotics · reinforcement-learning

Why read this

Hightop 30%

How H-JEPA's levels, SIGReg and top-down planner work, with the two gains split apart, and where the paper's own tables and code disagree.

  • Interactive explanations
  • Original analysis
  • Runs on a consumer GPU

Robotics & embodiedMITResearch paper

How this was scored
Is it new?
2 of 3: A real new idea, method or capability
Can I trust it?
2 of 3: Measures key facts from files, code or configs
Can I run it?
2 of 3: Open code or weights with real limits
Will I understand it?
3 of 3: Mechanism carried by interactives built from real code or data
Can I act on it?
2 of 3: A concrete recipe, numbers or comparison
Will it last?
2 of 3: A reference for a year or more
Does it affect many?
1 of 3: A specialist community
Only here?
2 of 3: A teardown or measurement few others did

Score 71 of 100, ranked 90 of 445 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored

Randall Balestriero's announcement calls H-JEPA "the first end-to-end learned hierarchical world model for long-horizon visual planning." I clicked because of the second post in his thread: "Semantic abstraction yields temporal compression!" It is a strong claim about where hierarchy comes from. It says nobody has to design the levels. Ask each level to predict farther ahead, and the level throws away whatever it cannot predict at that distance, and what is left is slow and abstract on its own.

The paper is arXiv 2610.06805, by Wancong Zhang, Basile Terver, Michael Rabbat, Yann LeCun and Balestriero, with code at kevinghst/H-JEPA under MIT and 93 checkpoints on Hugging Face. I read all three. The headline number holds up: on Visual AntMaze, a three-level model raises planning success from 18% to 73%. What I did not expect is how carefully the paper takes that number apart. It isolates two different reasons hierarchy helps, shows one of them working with no hierarchical planner at all, and admits two environments where neither shows up. Then an appendix proves that the regularizer the whole recipe rests on is blind to a specific failure on real robot video. I found that appendix more useful than the headline.

kevinghst/H-JEPA@840e76b · snapshot 2026-10-06
tracked files
159
license
MIT
branch
main
tests
none found
source
721.1 kB
commit date
2026-10-06
source by language
Python715.0 kB(66)Shell6.1 kB(5)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

local clone, 2026-10-06 at 840e76b — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow

shallow clone: counts describe the pinned tree, not the history

Left: three rows of Visual AntMaze frames, one per hierarchy level. Level 3 predicts 20 environment steps per prediction, level 2 predicts 10, level 1 predicts 5; the first prediction of each level is boxed and passed down as the subgoal for the level below. Right: planning success against planner TFLOPs per episode on a log axis, with three-level H-JEPA highest, two-level in the middle and flat LeWM lowest.
Planning top-down in Visual AntMaze: each level's first predicted state becomes the subgoal of the level below. Right, success against planner compute for flat LeWM and two- and three-level H-JEPA (paper, Figure 1).

A world model that never draws a pixel

H-JEPA builds on LeWorldModel, or LeWM, from the same group, so start there. A JEPA world model has three parts. An encoder EE turns an image oto_t into a vector ztz_t. An action encoder turns the command ata_t into a vector. A predictor FF takes the current latent and the action and guesses the next latent. Training minimises the squared error between F(zt,at)F(z_t, a_t) and the encoder's own zt+1z_{t+1}. No decoder, no reward, no pixels in the loss. If you have read the site's piece on next-latent prediction, it is the same idea with actions in the loop.

Planning with it is pleasantly direct. Encode the current frame and a goal frame. Start from a batch of random action sequences, roll each through the predictor, measure how far the last predicted latent lands from the goal latent, and push the actions downhill on that distance with gradient descent. Execute the first couple of actions, look again, repeat. Model predictive control, with a learned model doing the predicting.

The catch is the one every JEPA has. The easiest way to predict your own next latent perfectly is to make every latent the same. A constant encoder gets zero prediction loss and encodes nothing. Earlier JEPAs stopped this with an EMA teacher, a stop-gradient, VICReg-style variance and covariance terms, or a frozen pretrained encoder. LeWM's claim was that you need none of that: two loss terms, prediction plus SIGReg, trained from pixels end to end. H-JEPA's whole recipe depends on that being true at every level of a stack, so SIGReg deserves a proper look.

SIGReg, from the formula to the 35 lines that ship

SIGReg comes from Balestriero and LeCun's LeJEPA, which the site covered through its video descendant LeVJEPA. The target is an isotropic Gaussian: embeddings should look like draws from N(0,ID)\mathcal{N}(0, I_D). Testing that in DD dimensions directly is hopeless, so it leans on the Cramér-Wold theorem: a distribution is Gaussian if and only if every one-dimensional projection of it is. Pick MM random unit directions umu_m, project the batch to hm=Zumh_m = Z u_m, and run a one-dimensional normality test on each:

SIGReg(Z)=1M∑m=1MTEP(hm),TEP(h)=n∫−33∣φ^h(t)−e−t2/2∣2e−t2/2 dt\mathrm{SIGReg}(Z) = \frac{1}{M}\sum_{m=1}^{M} T_{\mathrm{EP}}(h_m), \qquad T_{\mathrm{EP}}(h) = n \int_{-3}^{3} \left| \hat\varphi_h(t) - e^{-t^2/2} \right|^2 e^{-t^2/2}\, dt

φ^h(t)=1n∑ieithi\hat\varphi_h(t) = \frac{1}{n}\sum_i e^{\mathrm{i} t h_i} is the empirical characteristic function of the nn projected values, e−t2/2e^{-t^2/2} is the standard Gaussian's characteristic function, and the outer e−t2/2e^{-t^2/2} weights low frequencies. This is the Epps-Pulley statistic. Each level of H-JEPA adds it to its prediction loss with a weight λℓ\lambda_\ell:

L(ℓ)=Lpred(ℓ)+λℓ SIGReg(Z(ℓ))\mathcal{L}^{(\ell)} = \mathcal{L}^{(\ell)}_{\mathrm{pred}} + \lambda_\ell\, \mathrm{SIGReg}\big(Z^{(\ell)}\big)

The released implementation is short enough to read whole. Its core is h_jepa/loss.py:16-47:

def __init__(self, knots=17, num_proj=1024):
    t = torch.linspace(0, 3, knots, dtype=torch.float32)
    dt = 3 / (knots - 1)
    weights = torch.full((knots,), 2 * dt, dtype=torch.float32)
    weights[[0, -1]] = dt
    window = torch.exp(-t.square() / 2.0)
    ...
def forward(self, proj):            # proj: (T, B, D)
    A = torch.randn(proj.size(-1), self.num_proj, device=proj.device)
    A = A.div_(A.norm(p=2, dim=0))
    x_t = (proj @ A).unsqueeze(-1) * self.t
    cos, sin = x_t.cos().mean(-3), x_t.sin().mean(-3)
    ...                             # all_reduce of cos/sin across ranks
    err = (cos - self.phi).square() + sin.square()
    statistic = (err @ self.weights) * n
    return statistic.mean()

Three details are worth knowing. The integral is a trapezoid rule on [0,3][0, 3] with 17 knots, doubled because the integrand is even, so the interior weights are 2 * dt and the ends dt. The directions are redrawn every call, 1,024 of them, so over training the test covers far more than 1,024 directions. And the mean over samples is taken over the batch axis of a (T, B, D) tensor: one normality test per timestep, across the BB sequences in the batch. Hold on to that last point. It matters a lot on DROID.

The appendix adds one engineering fix I liked. The statistic is a property of the whole batch, so computing it per GPU and averaging gives a different number depending on how many GPUs you use. The authors all-reduce the real and imaginary ECF before the quadrature instead, two M×17M \times 17 tensors per call whatever the batch size, and show in their Table 11 that the result becomes exactly invariant to sharding, where LeWM's single-GPU version drifts by 14% between one shard and two. If you train SIGReg models across devices, copy that.

The widget below runs the same arithmetic on a toy batch of 256 embeddings in two dimensions, sweeping 16 directions instead of 1,024 random ones. Try "dimensional collapse": one axis survives, so projections along it pass, and only the directions near the dead axis fail. That is why one direction is never enough. Then try "slow-feature collapse", which I explain further down.

SIGReg on one timestep: 256 embeddings, 17 knots, 16 directionsmean statistic = 0.62
embeddings z (batch at one timestep)uRe ECF of h = z·u vs e^(−t²/2)t = 301statistic per direction (log)1this u: 0.99moves with the robot: yes

An isotropic Gaussian cloud. Every direction passes; the average sits at the statistic's noise floor, near 1 in expectation, which is where the paper says a healthy run settles.

The paper's training curves give a scale for these values. A healthy run's statistic drops fast early in training and settles "within a small factor of its value on exactly Gaussian embeddings, about 1." On collapsed data it grows linearly with the batch size nn, which keeps the anti-collapse gradient per sample constant.

Stacking the JEPAs

H-JEPA is a stack of these world models where each level reads the level below instead of pixels. Level 1 is exactly LeWM: a ViT-Tiny over the image, its [CLS] token through an MLP projector, and a four-layer causal transformer predictor with adaptive LayerNorm action conditioning. Every level above it has its own encoder, its own action encoder and its own predictor, and two temporal knobs. The stride sℓs_\ell says how many lower steps one upper step covers. The window wℓw_\ell says how many lower states one upper state summarises:

zt(ℓ)=E(ℓ) ⁣(ztsℓ : tsℓ+wℓ(ℓ−1)),at(ℓ)=A(ℓ) ⁣(atsℓ+wℓ−1 : (t+1)sℓ+wℓ−1(ℓ−1))z^{(\ell)}_t = E^{(\ell)}\!\left(z^{(\ell-1)}_{t s_\ell \,:\, t s_\ell + w_\ell}\right), \qquad a^{(\ell)}_t = A^{(\ell)}\!\left(a^{(\ell-1)}_{t s_\ell + w_\ell - 1 \,:\, (t+1) s_\ell + w_\ell - 1}\right)

Every experiment in the paper uses wℓ=1w_\ell = 1 and sℓ=2s_\ell = 2 in simulation, so the upper state encoder is a two-layer MLP applied to every other lower latent, and the action encoder, a one-layer transformer, pools the two lower action embeddings in between into a 4-dimensional macro-action. A level-1 step already spans 5 environment steps, so level 2 predicts 10 steps at a time, level 3 predicts 20, and level 4 predicts 40.

Left: level-1 encoders E1 turn observations o0 to o3 into latents z1; level-2 encoders E2 read only z0 and z2, skipping the faded z1 and z3; on the right, two level-1 action embeddings feed one level-2 action encoder A2. Right: each level's predictor F takes the latent and action and predicts the next latent, scored against the encoded target, with SIGReg applied to the batch of encoded states and the per-level objective written as prediction loss plus lambda times SIGReg.
Building and training the hierarchy with window 1 and stride 2: level 2 encodes every other level-1 state and pools the two action embeddings between them; every level trains its own predictor with SIGReg on its own states (paper, Figure 2).

The part that makes this "end to end" is simple to state and the reason the paper exists. All levels train together from scratch, and the upper levels' losses backpropagate into the lower encoders. The code just sums the per-level losses, h_jepa/hjepa_utils.py:172:

total_loss = level_loss if total_loss is None else total_loss + level_loss

There is no detach between levels and no stop-gradient on targets. So the level-1 encoder is shaped by every level above it, and every level's latent space is free to collapse under its own predictor. The obvious hack would be to train level 1, freeze it, train level 2 on top. The paper has that variant, calls it stagewise, and uses it only for analysis. Joint training is only safe because SIGReg holds every level's latents open at the same time, with the same weight per level: 0.18 on FourRoom and 0.09 everywhere else. The four-level training curves in the appendix show no level's prediction loss sliding toward zero, which is what a collapse looks like.

jepa-world-models/h-jepa@6844abd · snapshot 2026-10-06
repo size
12.43 GB
task
robotics
library
pytorch
license
mit
largest file
348.2 MB
files
188
downloads
0
likes
0
world-modelsjepaplanningrobotics

repo last modified 2026-10-06

One thing surprised me in the released files. The upper levels are not small. On AntMaze the flat LeWM checkpoint is 49.6 MB, and each extra level adds about 45 MB: 94.6 MB for two levels, 139.5 MB for three, 184.5 MB for four. The level-2 predictor is a five-layer transformer with 12 heads, which costs as much as the whole level-1 model. A three-level H-JEPA is roughly three LeWMs glued together, which is worth knowing before you call it cheap. The planner is a different story; I come back to it.

What the upper levels throw away

Thread post two claims abstraction emerges. The evidence is a probing experiment. Freeze a trained four-level model, fit a small MLP probe on each level's latents to recover a physical quantity, and report the error normalised by the quantity's variance, so 0 means perfect and 1 means no better than guessing the mean.

Left: probe NMSE by level on Visual AntMaze, where body-state error rises from near 0.06 at level 1 to about 0.8 at level 4 while position error stays near zero. Middle: FourRoom Distractors, where distractor error rises from about 0.06 to about 0.58 while the ego agent's error stays near zero. Right: one frame decoded from levels 1 to 3 for each environment; the ant's legs blur with depth and the blue distractor dot fades while the red agent stays sharp.
Probe error by level: the ant's body state and the FourRoom distractor become unrecoverable at higher levels while the agent's position stays readable; decoders are for visualisation only (paper, Figure 4).

On AntMaze, the ant's 27-dimensional body state becomes steadily harder to read from levels 2, 3 and 4, while its global position stays near-perfectly recoverable. On FourRoomDistractors, the distractor dot fades and the agent stays. The paper's explanation is the one from the thread, and I find it convincing because it is mechanical rather than mystical. The encoder at level ℓ\ell only gets gradient from a predictor asked to look 5⋅2ℓ−15 \cdot 2^{\ell-1} environment steps ahead. The ant's joints oscillate many times between two level-3 states 20 steps apart, so encoding them only adds prediction error; the position drifts slowly and stays predictable. The FourRoom distractor teleports at Poisson intervals with a mean of 35 steps, and a longer horizon is more likely to straddle a teleport.

The paper is honest that this is conditional. Figure 5 relates the effect to a "frequency gap", the ratio of the spectral centroids of the fastest and slowest state components. Humanoid, AntMaze and FourRoom have gaps of 3.1 to 12.7 times and show the effect. Push-T, OGBench Cube and DROID have gaps of 1.2 to 1.6 times and show little of it: in a pushing task, the block and the pusher move on the same timescale, so there is nothing slow to keep and nothing fast to drop. So "semantic abstraction yields temporal compression" is true when the world has separated timescales, and that condition decides almost everything that follows.

Planning from the top down

Planning starts by encoding the current frame and the goal frame at every level. The top level then optimises macro-actions so that its last predicted latent lands on the goal:

a0:HL−1(L),∗=arg⁡min⁡a(L)∥z^HL(L)−g(L)∥22a^{(L),*}_{0:H_L-1} = \arg\min_{a^{(L)}} \left\| \hat z^{(L)}_{H_L} - g^{(L)} \right\|_2^2

Each level below plans primitive or less-macro actions, but it is not scored in its own space. Its predicted rollout is lifted through the upper encoder, z~i(ℓ+1)=E(ℓ+1)(z^isℓ+1(ℓ))\tilde z^{(\ell+1)}_i = E^{(\ell+1)}(\hat z^{(\ell)}_{i s_{\ell+1}}), and compared with the subgoals the level above predicted:

a(ℓ),∗=arg⁡min⁡a(ℓ)[∥z~K(ℓ+1)−z^K(ℓ+1),∗∥22+β∑i=1K−1∥z~i(ℓ+1)−z^i(ℓ+1),∗∥22]a^{(\ell),*} = \arg\min_{a^{(\ell)}} \left[ \left\| \tilde z^{(\ell+1)}_{K} - \hat z^{(\ell+1),*}_{K} \right\|_2^2 + \beta \sum_{i=1}^{K-1} \left\| \tilde z^{(\ell+1)}_{i} - \hat z^{(\ell+1),*}_{i} \right\|_2^2 \right]

with KK the number of subgoals matched. In the closed-loop simulation runs K=1K = 1: each level only has to reach the first state the level above imagined. Level 1's actions are executed for two blocks, 10 environment steps, and then the whole stack replans from the new frame.

I went in expecting CEM or MPPI, the sampling planners most latent world models use. The optimiser is gradient descent through the rollout at every level: AdamW, at most 90 iterations, early stopping when the best cost stalls, keep the lowest-cost candidate. Macro-actions have no natural bounds, so the code keeps the last 4,000 macro-actions seen in training, initialises candidates from a Gaussian fitted to them, and clips them after every step to that queue's 2nd and 98th percentiles (stable_worldmodel/solver/gd.py:193-194). Without that, gradient descent would happily wander into macro-actions the upper predictor has never seen. The top-down loop itself is plain, h_jepa/hierarchical_solver.py:121-213, trimmed:

for level in range(len(self.level_models), 0, -1):
    if level > 1 and planning_horizon == 1:
        continue                     # a one-step upper plan yields no useful subgoal
    if subgoal_latents is None:      # top level: aim at the real goal
        level_info['goal_embed_0'] = goal_embeddings[f'embed_{level}']
    else:                            # lower level: aim at the level above's prediction
        next_goal = subgoal_latents[:, first_pred_idx:first_pred_idx + num_subgoals].clone()
        if includes_last and upper_anchored:
            next_goal[:, -1] = goal_embeddings[f'embed_{level + 1}'].squeeze(1)
        level_info['goal_embed_0'] = next_goal
    result = self.level_solvers[level].solve(info_dict=level_info, ...)
    subgoal_latents = result['predictions']['predicted_embed_0']

The anchoring line is a nice touch the paper does not dwell on: when a subgoal would be the upper plan's final state, and that plan was itself aimed at the true goal, the code swaps in the real goal embedding so a drifting prediction cannot propagate down the chain.

The widget shows what this buys in reach. Pick an environment and a depth and it draws, in environment steps, how far each level plans on the first call, using the horizons from the paper's planning table. A flat LeWM on AntMaze has to roll its single predictor 23 steps, 115 environment steps, and differentiate through all of it. A three-level model reaches 100 environment steps with five predictions at the top, and while the top is still planning, each level below rolls out only two steps.

what each level predicts on the first planning call (Table 7 settings)3-level · Visual AntMaze
task ≈ 82 stepslevel 320 steps/pred · H=5100level 210 steps/pred · H=220level 15 steps/pred · H=4100environment steps
H-JEPA success
73.3%
HWM success
38.7%
flat LeWM
18.0%
training clips vs LeWM
88%

Visual AntMaze: expert route between cells three apart, 82 steps on average, episode budget 350 steps. H is the configured horizon from the paper’s Table 7. On the first call the top active level plans its full H; every level below it plans two steps, just enough to reach the first subgoal it is handed. Later calls shrink the top horizon toward one, and an upper level whose horizon is one is skipped.

So the compute story is about depth of rollout. Gradients through two predictor steps are cheap and well-behaved; gradients through 23 autoregressive steps of a model that has compounding error are neither. The planner budget in the depth comparison is capped at 100 TFLOPs per episode for every model, which is what makes the comparison fair.

The result, and the two levers inside it

Top row: success rate against planner TFLOPs per episode for FourRoom Distractors, Visual AntMaze, OGBench Cube and Push-T, comparing flat LeWM, two-level and three-level H-JEPA; on the first three, three levels sit highest and leftmost. Bottom row: success against number of levels from 1 to 4 for H-JEPA and HWM; AntMaze peaks at three-level H-JEPA near 73% while HWM stays below 40%; Push-T drops to zero at four levels for both.
Top: success against planner compute, with Pareto fronts in bold. Bottom: success against depth on a separate set of 50 tasks per environment, H-JEPA against HWM, a hierarchy that shares one latent space (paper, Figure 6).

On FourRoom, AntMaze and Cube each added level up to three moves the Pareto front up and to the left: more success for less planner compute. On AntMaze, under the 100-TFLOP cap, flat LeWM solves 18.0% of the 50 tasks, two-level H-JEPA 39.3%, three-level 73.3%. Cube goes from 32.0% to 60.0% at three levels. FourRoom goes from 40.7% to 96.0%.

The comparison I care about is the blue line, because it answers "is this just HWM again?" HWM is the same group's earlier hierarchical planner: one latent space, several prediction horizons. The authors rebuild it inside this codebase as H-JEPA with identity encoders above level 1, so the two differ only in whether the upper levels get a representation of their own. On AntMaze the difference is large: three-level HWM reaches 38.7% against H-JEPA's 73.3%, and four-level HWM falls to 14.0%. On FourRoom it vanishes, and HWM is actually ahead at three levels, 99.3% against 96.0%. The checkpoints confirm this is a fair fight in size: the identity encoders save only 1 to 4 MB per AntMaze model (93.4 MB against 94.6 MB at two levels), because the predictors dominate.

So there are two things in a hierarchy, and Section 4.2 pulls them apart.

The first is the cost. The authors take each trained H-JEPA, throw away its upper planners, and plan flat with level 1 only, changing nothing except where the distance to the goal is measured: in level-1 space, or after pushing both the prediction and the goal up through the upper encoders. The rollout dynamics are identical. Only the ruler changes.

AntMaze success (%)2 levels3 levels4 levels
Flat planning, native level-1 cost23.316.710.0
Flat planning, best upper-level cost31.322.026.7
Full hierarchical planner39.373.363.3

Measuring distance in a more abstract space helps a flat planner at every depth, and the best projected cost beats the separately trained LeWM's 18.0%. The cost maps show why. In level-1 space, the latent distance to a fixed point in the maze is nearly flat everywhere except right next to it, so a gradient planner has nothing to follow until it is already close. At levels 2 and 3 the distance rises smoothly along the corridors.

Three maze maps coloured by latent distance to a starred anchor point, one per level of a three-level H-JEPA. At level 1 almost the whole maze is uniformly red except a tiny blue spot at the star. At levels 2 and 3 a broad blue-to-yellow basin surrounds the star and the colour grades along the corridors.
Latent distance to the starred anchor in each level's space, one training seed: level-1 distances are almost constant away from the anchor, higher levels grade along the corridors (paper, Figure 7).

The second lever is time. A flat planner's cost to the final goal is a bad progress signal even along an expert trajectory: the paper measures its monotonicity, the negative Spearman correlation of cost with time, at 0.48 to 0.75 across the four simulated environments. The cost of tracking the next level-2 subgoal over a short segment scores 0.89 to 1.00 for H-JEPA. HWM gets nearly the same subgoal benefit, which is why it beats flat LeWM everywhere. The representation lever is what H-JEPA adds on top, and it only shows up where the frequency gap gave the upper levels something to throw away.

I like this section a lot. The cost-only ablation is the kind of experiment that tells a practitioner something usable today: if you already have a hierarchical encoder, you can get part of the gain with your existing flat planner by changing the ruler.

Where it does not help

The paper puts its failures in the main text, and they are instructive.

Push-T is the clearest. Two levels match or slightly beat LeWM, 45.3% against 40.0%. Three levels drop to 17.3% and four to 0.7%. The paper's explanation is data, and I think it is right. A training sample must fit inside one episode, and a four-level sample needs 25 level-1 frames, 121 environment steps. Push-T episodes average 125 steps, so a four-level model sees only 14% of the training clips LeWM sees per epoch, and three levels see 58%. HWM collapses the same way, which points at the data rather than the representation.

Cube is subtler. The full hierarchy helps there, 60.0% at three levels, but the cost-only ablation does not: the level-2 cost lowers the two-level model's flat success from 24.0% to 13.3%, and the level-3 and level-4 costs roughly halve success in the deeper models. With no frequency gap there is no useful abstraction to measure in, so whatever Cube gains comes from the subgoal lever alone. The authors say as much, and add that their results "do not establish that selective abstraction alone causes the gains in the other environments." A careful sentence, and I appreciated it.

Then the data-specialisation appendix, which deserves more attention than its position suggests. Trained end to end, both levels must see the same data. Trained stagewise, each level can have its own. On AntMaze, training level 1 on the noisy exploration data and level 2 on the goal-directed "stitch" data reaches 67% hierarchical success, while end to end on the union of both reaches 54%. The reverse pairing collapses to 2%. The low level wants coverage; the high level wants trajectories that go somewhere. End-to-end training cannot express that preference. The paper's headline recipe is end to end, but its own best AntMaze cell outside the main table came from not doing it.

DROID, and a collapse SIGReg cannot see

Every environment so far has a fixed camera and a fixed background. DROID does not: it is real teleoperated Franka video where the scene, lighting, camera calibration and objects change every episode. Train LeWM on it and planning fidelity is zero.

Left: bar chart of Fréchet fidelity on DROID: one-level LeWM without inverse dynamics sits at zero labelled slow-feature collapse; with inverse dynamics it reaches about +34; two-level HWM about +35; two-level H-JEPA about +40. Right: Fréchet fidelity against planner TFLOPs per episode on a log axis; H-JEPA's curve sits highest near 40 at under 100 TFLOPs, HWM and LeWM plus IDM lower, and frozen-encoder baselines JEPA-WM and V-JEPA 2-AC far to the right at 1,000 to 100,000 TFLOPs with lower fidelity.
DROID at 5 fps: without an inverse-dynamics term the flat model collapses to the zero-action floor; with it, a second level adds fidelity, at lower planner compute than the flat model and orders of magnitude less than the frozen-encoder world models (paper, Figure 9b and 9c; the two panels are cropped from the rendered figure, an overlapping (c) label removed and the 20 axis tick redrawn).

The appendix that explains this is the best writing in the paper. It names the failure slow-feature collapse: the encoder learns to encode only what is constant within an episode, the background, and ignores the moving arm. That solution has zero prediction loss, because the identity predictor is perfect when nothing changes, and here is the uncomfortable part: SIGReg cannot rule it out. Their Proposition 1 says it for any regulariser that only looks at the marginal distribution of single embeddings. If backgrounds vary continuously across episodes, you can map each background to a draw from exactly the target Gaussian. Every batch then looks perfectly isotropic, and the encoder still says nothing about the robot.

Recall the (T, B, D) grouping in the SIGReg code. Each test sees one timestep across many sequences, and a collapsed encoder gives one distinct, Gaussian-looking embedding per sequence. The "slow-feature collapse" button in the SIGReg widget above is exactly that case: the same cloud as the healthy one, the same statistic, and nothing that moves when the robot moves. The appendix measures where the variance actually goes in trained latents: 75% of it lies between episodes on DROID against 28% on Cube.

They test two fixes. Putting time into the sample axis of the test adds a barrier, but a small one, and every run in their SIGReg sweep still collapses. A "time repulsion" term that rewards latents for changing keeps the predictor off the identity but never clears the zero-action floor. What works is an inverse-dynamics head that must regress the action from two consecutive latents. A collapsed encoder carries no information about the action, so the head's loss has a floor fixed by the data, 1 with normalised actions, which is 50 at the weight they use. The appendix calls it "the only Ω(1)\Omega(1) barrier here," and notes the cost: it needs action labels. So the two-term recipe that works in simulation becomes, on real video, a three-term one: prediction, SIGReg, and action regression with weight 100 at level 1 and 50 at level 2.

With that term, the numbers are modest and honest. The paper's bars read +34 for flat LeWM plus inverse dynamics, +35 for HWM and +40 for H-JEPA, as medians over training and planner seeds. The README reports the same models as means: 33.57%, 35.65% and 38.42%. H-JEPA gets there at 10.1 planner TFLOPs per episode against 14.5 for the flat model.

Two things limit how far I would take this. Fréchet fidelity is an offline metric: it compares the planned end-effector path with the logged expert path, with 0% meaning the arm does not move and 100% meaning it follows the expert. Nothing is executed on a robot. The paper calibrates the metric in simulation and concludes it is a necessary condition for success on Cube, not a predictor of it. Their conclusion states the open question plainly: "A next step is to test whether these gains translate to closed-loop control on a physical robot." And the evaluation is 16 hand-picked clips with good lighting and large objects. A careful protocol, and a small one.

Sixteen DROID clips in two columns, ordered by Fréchet fidelity from +75% down to the zero floor. Each clip shows a row of logged expert frames above a row of decoded planned frames, with every third planned frame ringed in blue as the level-2 macro-plan; the planned frames blur along the rollout.
What the DROID planner imagines on all 16 evaluation clips, ordered best to worst; the decoded plans move the gripper toward the object before the goal, and blur is the decoder's limit (paper, Figure 16).

Two replies worth reading

Two replies under the announcement point at work the article should mention. Daniel Korchinski linked his paper with Alessandro Favero and Matthieu Wyart, arXiv 2605.27734. It proves that on data from a probabilistic context-free grammar, latent prediction recovers a hidden tree with a sample count constant in its depth, where token-level learning needs exponentially many, and that data2vec already does this implicitly. Its abstract ends: "This suggests that explicit stacking such as H-JEPA is largely redundant." The theory is about static compositional data with no actions and no planning. H-JEPA's evidence agrees with it in one respect: on Push-T and Cube, where there is no separation of timescales, a separate representation per level adds little. Where the world has a slow part and a fast part, and a planner needs a ruler for the slow part, the explicit levels earned their keep. I do not think the two papers contradict each other; they disagree about which data matters.

The other reply links HDFlow, a hierarchical planner with a diffusion model proposing latent subgoals and a rectified-flow model filling in trajectories, and asks why this group never uses diffusion dynamics. Nobody had answered when I read the thread. My reading of the design is that a deterministic predictor is what makes gradient-through-the-rollout planning possible here; a stochastic generator would push you back toward sampling planners. I am inferring this; the authors have not said it.

What I would take from it

The idea that will last is the separation of the two levers. Measure distance where the dynamics are slow, plan in time where the horizon is long, and check which one your problem actually needs. The cost-projection experiment is cheap to run on any stack of encoders, and the frequency gap is a cheap thing to compute on a dataset before you build anything.

The part I would be careful with is "end to end" as a virtue. It works, it is stable, and SIGReg genuinely makes joint training of four latent spaces uneventful. But the paper's own appendix shows that training each level on its own data beat it on AntMaze, and that on real video the regulariser needs an action-supervised partner. The training is end to end; the safety net is not SIGReg alone.

And the real-robot claim should wait for a real robot. The simulation results are strong and well controlled. DROID is an offline proxy on 16 clips.

How I checked

I read the arXiv HTML of 2610.06805v1, including every appendix, and shallow-cloned kevinghst/H-JEPA at commit 840e76b. Code quotes are from that commit with file and line. The hierarchy, loss assembly, SIGReg quadrature, solver loop, macro-action clipping and per-level configs were read from source; I did not run any of it. Checkpoint counts and sizes come from the Hugging Face model API for jepa-world-models/h-jepa: 93 checkpoint files, 84 for simulation and 9 for DROID, about 12.43 GB in all, matching the README's 9.7 GB and 2.7 GB. The SIGReg widget is my own toy that reimplements the released formula on synthetic 2-D data; the depth widget plots the paper's Table 7 horizons and the Figure 6 and Table 3 numbers.

A few things in the sources do not agree, none of them large. The four-level AntMaze hierarchical result is 63.3% in the paper's Table 1 and 67.3% in Figure 6 and the README, so I quote the table where I quote the table. The paper's DROID planning table lists 16 candidates at level 2, while the released droid_l2.yaml and the README use 4. The DROID training corpus is 74,530 episodes in the paper and 74,896 in the README's file list. ARCHITECTURE.md still describes a Push-T example with stride 3 and 25 frames per step, which matches no released config. I could not check any number that depends on running the planner, including all success rates and fidelities; those are the authors'.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "H-JEPA: the top of the stack forgets the ant's legs, and that is the point", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026hjepa,
  author = {Satyajit Ghana},
  title  = {H-JEPA: the top of the stack forgets the ant's legs, and that is the point},
  url    = {https://ai.thesatyajit.com/articles/h-jepa},
  year   = {2026}
}
share