# H-JEPA: the top of the stack forgets the ant's legs, and that is the point

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/h-jepa
> date: 2026-10-06
> tags: world-models, self-supervised-learning, representation-learning, robotics, reinforcement-learning

Randall Balestriero's announcement calls H-JEPA "the first end-to-end learned hierarchical world
model for long-horizon visual planning." I clicked because of the second post in his thread:
"Semantic abstraction yields temporal compression!" It is a strong claim about where hierarchy
comes from. It says nobody has to design the levels. Ask each level to predict farther ahead,
and the level throws away whatever it cannot predict at that distance, and what is left is slow
and abstract on its own.

The paper is [arXiv 2610.06805](https://arxiv.org/abs/2610.06805), by Wancong Zhang, Basile
Terver, Michael Rabbat, Yann LeCun and Balestriero, with code at
[kevinghst/H-JEPA](https://github.com/kevinghst/H-JEPA) under MIT and 93 checkpoints on
[Hugging Face](https://huggingface.co/jepa-world-models/h-jepa). I read all three. The headline
number holds up: on Visual AntMaze, a three-level model raises planning success from 18% to
73%. What I did not expect is how carefully the paper takes that number apart. It isolates two
different reasons hierarchy helps, shows one of them working with no hierarchical planner at
all, and admits two environments where neither shows up. Then an appendix proves that the
regularizer the whole recipe rests on is blind to a specific failure on real robot video. I
found that appendix more useful than the headline.

<RepoCard repo="kevinghst/H-JEPA" />

<Figure
  src="https://ai.thesatyajit.com/articles/h-jepa/fig1.png"
  alt="Left: three rows of Visual AntMaze frames, one per hierarchy level. Level 3 predicts 20 environment steps per prediction, level 2 predicts 10, level 1 predicts 5; the first prediction of each level is boxed and passed down as the subgoal for the level below. Right: planning success against planner TFLOPs per episode on a log axis, with three-level H-JEPA highest, two-level in the middle and flat LeWM lowest."
  caption="Planning top-down in Visual AntMaze: each level's first predicted state becomes the subgoal of the level below. Right, success against planner compute for flat LeWM and two- and three-level H-JEPA (paper, Figure 1)."
/>

## A world model that never draws a pixel

H-JEPA builds on LeWorldModel, or [LeWM](https://arxiv.org/abs/2603.19312), from the same
group, so start there. A JEPA world model has three parts. An encoder $E$ turns an image $o_t$
into a vector $z_t$. An action encoder turns the command $a_t$ into a vector. A predictor $F$
takes the current latent and the action and guesses the next latent. Training minimises the
squared error between $F(z_t, a_t)$ and the encoder's own $z_{t+1}$. No decoder, no reward, no
pixels in the loss. If you have read the site's piece on
[next-latent prediction](/articles/next-latent-world-models), it is the same idea with
actions in the loop.

Planning with it is pleasantly direct. Encode the current frame and a goal frame. Start from a
batch of random action sequences, roll each through the predictor, measure how far the last
predicted latent lands from the goal latent, and push the actions downhill on that distance
with gradient descent. Execute the first couple of actions, look again, repeat. Model
predictive control, with a learned model doing the predicting.

The catch is the one every JEPA has. The easiest way to predict your own next latent perfectly
is to make every latent the same. A constant encoder gets zero prediction loss and encodes
nothing. Earlier JEPAs stopped this with an EMA teacher, a stop-gradient, VICReg-style variance
and covariance terms, or a frozen pretrained encoder. LeWM's claim was that you need none of
that: two loss terms, prediction plus SIGReg, trained from pixels end to end. H-JEPA's whole
recipe depends on that being true at every level of a stack, so SIGReg deserves a proper look.

## SIGReg, from the formula to the 35 lines that ship

SIGReg comes from Balestriero and LeCun's [LeJEPA](https://arxiv.org/abs/2511.08544), which the
site covered through its video descendant [LeVJEPA](/articles/levjepa). The target is an
isotropic Gaussian: embeddings should look like draws from $\mathcal{N}(0, I_D)$. Testing that in
$D$ dimensions directly is hopeless, so it leans on the Cramér-Wold theorem: a distribution is
Gaussian if and only if every one-dimensional projection of it is. Pick $M$ random unit
directions $u_m$, project the batch to $h_m = Z u_m$, and run a one-dimensional normality test on
each:

$$
\mathrm{SIGReg}(Z) = \frac{1}{M}\sum_{m=1}^{M} T_{\mathrm{EP}}(h_m),
\qquad
T_{\mathrm{EP}}(h) = n \int_{-3}^{3} \left| \hat\varphi_h(t) - e^{-t^2/2} \right|^2 e^{-t^2/2}\, dt
$$

$\hat\varphi_h(t) = \frac{1}{n}\sum_i e^{\mathrm{i} t h_i}$ is the empirical characteristic
function of the $n$ projected values, $e^{-t^2/2}$ is the standard Gaussian's characteristic
function, and the outer $e^{-t^2/2}$ weights low frequencies. This is the Epps-Pulley statistic.
Each level of H-JEPA adds it to its prediction loss with a weight $\lambda_\ell$:

$$
\mathcal{L}^{(\ell)} = \mathcal{L}^{(\ell)}_{\mathrm{pred}} + \lambda_\ell\, \mathrm{SIGReg}\big(Z^{(\ell)}\big)
$$

The released implementation is short enough to read whole. Its core is
`h_jepa/loss.py:16-47`:

```python
def __init__(self, knots=17, num_proj=1024):
    t = torch.linspace(0, 3, knots, dtype=torch.float32)
    dt = 3 / (knots - 1)
    weights = torch.full((knots,), 2 * dt, dtype=torch.float32)
    weights[[0, -1]] = dt
    window = torch.exp(-t.square() / 2.0)
    ...
def forward(self, proj):            # proj: (T, B, D)
    A = torch.randn(proj.size(-1), self.num_proj, device=proj.device)
    A = A.div_(A.norm(p=2, dim=0))
    x_t = (proj @ A).unsqueeze(-1) * self.t
    cos, sin = x_t.cos().mean(-3), x_t.sin().mean(-3)
    ...                             # all_reduce of cos/sin across ranks
    err = (cos - self.phi).square() + sin.square()
    statistic = (err @ self.weights) * n
    return statistic.mean()
```

Three details are worth knowing. The integral is a trapezoid rule on $[0, 3]$ with 17 knots,
doubled because the integrand is even, so the interior weights are `2 * dt` and the ends `dt`.
The directions are redrawn every call, 1,024 of them, so over training the test covers far more
than 1,024 directions. And the mean over samples is taken over the batch axis of a
`(T, B, D)` tensor: one normality test per timestep, across the $B$ sequences in the batch.
Hold on to that last point. It matters a lot on DROID.

The appendix adds one engineering fix I liked. The statistic is a property of the whole batch,
so computing it per GPU and averaging gives a different number depending on how many GPUs you
use. The authors all-reduce the real and imaginary ECF before the quadrature instead, two
$M \times 17$ tensors per call whatever the batch size, and show in their Table 11 that the
result becomes exactly invariant to sharding, where LeWM's single-GPU version drifts by 14%
between one shard and two. If you train SIGReg models across devices, copy that.

The widget below runs the same arithmetic on a toy batch of 256 embeddings in two dimensions,
sweeping 16 directions instead of 1,024 random ones. Try "dimensional collapse": one axis
survives, so projections along it pass, and only the directions near the dead axis fail. That
is why one direction is never enough. Then try "slow-feature collapse", which I explain further
down.

<SigregProbe />

The paper's training curves give a scale for these values. A healthy run's statistic drops
fast early in training and settles "within a small factor of its value on exactly Gaussian
embeddings, about 1." On collapsed data it grows linearly with the batch size $n$, which keeps
the anti-collapse gradient per sample constant.

## Stacking the JEPAs

H-JEPA is a stack of these world models where each level reads the level below instead of
pixels. Level 1 is exactly LeWM: a ViT-Tiny over the image, its `[CLS]` token through an MLP
projector, and a four-layer causal transformer predictor with adaptive LayerNorm action
conditioning. Every level above it has its own encoder, its own action encoder and its own
predictor, and two temporal knobs. The stride $s_\ell$ says how many lower steps one upper step
covers. The window $w_\ell$ says how many lower states one upper state summarises:

$$
z^{(\ell)}_t = E^{(\ell)}\!\left(z^{(\ell-1)}_{t s_\ell \,:\, t s_\ell + w_\ell}\right),
\qquad
a^{(\ell)}_t = A^{(\ell)}\!\left(a^{(\ell-1)}_{t s_\ell + w_\ell - 1 \,:\, (t+1) s_\ell + w_\ell - 1}\right)
$$

Every experiment in the paper uses $w_\ell = 1$ and $s_\ell = 2$ in simulation, so the upper
state encoder is a two-layer MLP applied to every other lower latent, and the action encoder, a
one-layer transformer, pools the two lower action embeddings in between into a 4-dimensional
macro-action. A level-1 step already spans 5 environment steps, so level 2 predicts 10 steps
at a time, level 3 predicts 20, and level 4 predicts 40.

<Figure
  src="https://ai.thesatyajit.com/articles/h-jepa/fig2.png"
  alt="Left: level-1 encoders E1 turn observations o0 to o3 into latents z1; level-2 encoders E2 read only z0 and z2, skipping the faded z1 and z3; on the right, two level-1 action embeddings feed one level-2 action encoder A2. Right: each level's predictor F takes the latent and action and predicts the next latent, scored against the encoded target, with SIGReg applied to the batch of encoded states and the per-level objective written as prediction loss plus lambda times SIGReg."
  caption="Building and training the hierarchy with window 1 and stride 2: level 2 encodes every other level-1 state and pools the two action embeddings between them; every level trains its own predictor with SIGReg on its own states (paper, Figure 2)."
/>

The part that makes this "end to end" is simple to state and the reason the paper exists. All
levels train together from scratch, and the upper levels' losses backpropagate into the lower
encoders. The code just sums the per-level losses, `h_jepa/hjepa_utils.py:172`:

```python
total_loss = level_loss if total_loss is None else total_loss + level_loss
```

There is no detach between levels and no stop-gradient on targets. So the level-1 encoder is
shaped by every level above it, and every level's latent space is free to collapse under its
own predictor. The obvious hack would be to train level 1, freeze it, train level 2 on top. The
paper has that variant, calls it stagewise, and uses it only for analysis. Joint training is
only safe because SIGReg holds every level's latents open at the same time, with the same
weight per level: 0.18 on FourRoom and 0.09 everywhere else. The four-level training curves in
the appendix show no level's prediction loss sliding toward zero, which is what a collapse
looks like.

<ModelCard repo="jepa-world-models/h-jepa" />

One thing surprised me in the released files. The upper levels are not small. On AntMaze the
flat LeWM checkpoint is 49.6 MB, and each extra level adds about 45 MB: 94.6 MB for two levels,
139.5 MB for three, 184.5 MB for four. The level-2 predictor is a five-layer transformer with
12 heads, which costs as much as the whole level-1 model. A three-level H-JEPA is roughly three
LeWMs glued together, which is worth knowing before you call it cheap. The planner is a
different story; I come back to it.

## What the upper levels throw away

Thread post two claims abstraction emerges. The evidence is a probing experiment. Freeze a
trained four-level model, fit a small MLP probe on each level's latents to recover a physical
quantity, and report the error normalised by the quantity's variance, so 0 means perfect and 1
means no better than guessing the mean.

<Figure
  src="https://ai.thesatyajit.com/articles/h-jepa/fig3.png"
  alt="Left: probe NMSE by level on Visual AntMaze, where body-state error rises from near 0.06 at level 1 to about 0.8 at level 4 while position error stays near zero. Middle: FourRoom Distractors, where distractor error rises from about 0.06 to about 0.58 while the ego agent's error stays near zero. Right: one frame decoded from levels 1 to 3 for each environment; the ant's legs blur with depth and the blue distractor dot fades while the red agent stays sharp."
  caption="Probe error by level: the ant's body state and the FourRoom distractor become unrecoverable at higher levels while the agent's position stays readable; decoders are for visualisation only (paper, Figure 4)."
/>

On AntMaze, the ant's 27-dimensional body state becomes steadily harder to read from levels 2,
3 and 4, while its global position stays near-perfectly recoverable. On FourRoomDistractors,
the distractor dot fades and the agent stays. The paper's explanation is the one from the thread,
and I find it convincing because it is mechanical rather than mystical. The encoder at level
$\ell$ only gets gradient from a predictor asked to look $5 \cdot 2^{\ell-1}$ environment steps
ahead. The ant's joints oscillate many times between two level-3 states 20 steps apart, so
encoding them only adds prediction error; the position drifts slowly and stays predictable. The
FourRoom distractor teleports at Poisson intervals with a mean of 35 steps, and a longer horizon
is more likely to straddle a teleport.

The paper is honest that this is conditional. Figure 5 relates the effect to a "frequency gap",
the ratio of the spectral centroids of the fastest and slowest state components. Humanoid,
AntMaze and FourRoom have gaps of 3.1 to 12.7 times and show the effect. Push-T, OGBench Cube
and DROID have gaps of 1.2 to 1.6 times and show little of it: in a pushing task, the block and
the pusher move on the same timescale, so there is nothing slow to keep and nothing fast to
drop. So "semantic abstraction yields temporal compression" is true when the world has
separated timescales, and that condition decides almost everything that follows.

## Planning from the top down

Planning starts by encoding the current frame and the goal frame at every level. The top level
then optimises macro-actions so that its last predicted latent lands on the goal:

$$
a^{(L),*}_{0:H_L-1} = \arg\min_{a^{(L)}} \left\| \hat z^{(L)}_{H_L} - g^{(L)} \right\|_2^2
$$

Each level below plans primitive or less-macro actions, but it is not scored in its own space.
Its predicted rollout is lifted through the upper encoder, $\tilde z^{(\ell+1)}_i =
E^{(\ell+1)}(\hat z^{(\ell)}_{i s_{\ell+1}})$, and compared with the subgoals the level above
predicted:

$$
a^{(\ell),*} = \arg\min_{a^{(\ell)}} \left[ \left\| \tilde z^{(\ell+1)}_{K} - \hat z^{(\ell+1),*}_{K} \right\|_2^2 + \beta \sum_{i=1}^{K-1} \left\| \tilde z^{(\ell+1)}_{i} - \hat z^{(\ell+1),*}_{i} \right\|_2^2 \right]
$$

with $K$ the number of subgoals matched. In the closed-loop simulation runs $K = 1$: each level
only has to reach the first state the level above imagined. Level 1's actions are executed for
two blocks, 10 environment steps, and then the whole stack replans from the new frame.

I went in expecting CEM or MPPI, the sampling planners most latent world models use. The
optimiser is gradient descent through the rollout at every level: AdamW, at most 90 iterations,
early stopping when the best cost stalls, keep the lowest-cost candidate. Macro-actions have no
natural bounds, so the code keeps the last 4,000 macro-actions seen in training, initialises
candidates from a Gaussian fitted to them, and clips them after every step to that queue's 2nd
and 98th percentiles (`stable_worldmodel/solver/gd.py:193-194`). Without that, gradient descent
would happily wander into macro-actions the upper predictor has never seen. The top-down loop
itself is plain, `h_jepa/hierarchical_solver.py:121-213`, trimmed:

```python
for level in range(len(self.level_models), 0, -1):
    if level > 1 and planning_horizon == 1:
        continue                     # a one-step upper plan yields no useful subgoal
    if subgoal_latents is None:      # top level: aim at the real goal
        level_info['goal_embed_0'] = goal_embeddings[f'embed_{level}']
    else:                            # lower level: aim at the level above's prediction
        next_goal = subgoal_latents[:, first_pred_idx:first_pred_idx + num_subgoals].clone()
        if includes_last and upper_anchored:
            next_goal[:, -1] = goal_embeddings[f'embed_{level + 1}'].squeeze(1)
        level_info['goal_embed_0'] = next_goal
    result = self.level_solvers[level].solve(info_dict=level_info, ...)
    subgoal_latents = result['predictions']['predicted_embed_0']
```

The anchoring line is a nice touch the paper does not dwell on: when a subgoal would be the
upper plan's final state, and that plan was itself aimed at the true goal, the code swaps in the
real goal embedding so a drifting prediction cannot propagate down the chain.

The widget shows what this buys in reach. Pick an environment and a depth and it draws, in
environment steps, how far each level plans on the first call, using the horizons from the
paper's planning table. A flat LeWM on AntMaze has to roll its single predictor 23 steps, 115
environment steps, and differentiate through all of it. A three-level model reaches 100
environment steps with five predictions at the top, and while the top is still planning, each
level below rolls out only two steps.

<DepthLadder />

So the compute story is about depth of rollout. Gradients through two predictor steps are cheap and well-behaved;
gradients through 23 autoregressive steps of a model that has compounding error are neither.
The planner budget in the depth comparison is capped at 100 TFLOPs per episode for every model,
which is what makes the comparison fair.

## The result, and the two levers inside it

<Figure
  src="https://ai.thesatyajit.com/articles/h-jepa/fig5.png"
  alt="Top row: success rate against planner TFLOPs per episode for FourRoom Distractors, Visual AntMaze, OGBench Cube and Push-T, comparing flat LeWM, two-level and three-level H-JEPA; on the first three, three levels sit highest and leftmost. Bottom row: success against number of levels from 1 to 4 for H-JEPA and HWM; AntMaze peaks at three-level H-JEPA near 73% while HWM stays below 40%; Push-T drops to zero at four levels for both."
  caption="Top: success against planner compute, with Pareto fronts in bold. Bottom: success against depth on a separate set of 50 tasks per environment, H-JEPA against HWM, a hierarchy that shares one latent space (paper, Figure 6)."
/>

On FourRoom, AntMaze and Cube each added level up to three moves the Pareto front up and to the
left: more success for less planner compute. On AntMaze, under the 100-TFLOP cap, flat LeWM
solves 18.0% of the 50 tasks, two-level H-JEPA 39.3%, three-level 73.3%. Cube goes from 32.0% to
60.0% at three levels. FourRoom goes from 40.7% to 96.0%.

The comparison I care about is the blue line, because it answers "is this just HWM again?"
[HWM](https://arxiv.org/abs/2604.03208) is the same group's earlier hierarchical planner: one
latent space, several prediction horizons. The authors rebuild it inside this codebase as
H-JEPA with identity encoders above level 1, so the two differ only in whether the upper levels
get a representation of their own. On AntMaze the difference is large: three-level HWM reaches
38.7% against H-JEPA's 73.3%, and four-level HWM falls to 14.0%. On FourRoom it vanishes, and
HWM is actually ahead at three levels, 99.3% against 96.0%. The checkpoints confirm this is a
fair fight in size: the identity encoders save only 1 to 4 MB per AntMaze model (93.4 MB against 94.6 MB at two
levels), because the predictors dominate.

So there are two things in a hierarchy, and Section 4.2 pulls them apart.

The first is the cost. The authors take each trained H-JEPA, throw away its upper planners, and
plan flat with level 1 only, changing nothing except where the distance to the goal is measured:
in level-1 space, or after pushing both the prediction and the goal up through the upper
encoders. The rollout dynamics are identical. Only the ruler changes.

| AntMaze success (%) | 2 levels | 3 levels | 4 levels |
|---|---|---|---|
| Flat planning, native level-1 cost | 23.3 | 16.7 | 10.0 |
| Flat planning, best upper-level cost | 31.3 | 22.0 | 26.7 |
| Full hierarchical planner | 39.3 | 73.3 | 63.3 |

Measuring distance in a more abstract space helps a flat planner at every depth, and the best
projected cost beats the separately trained LeWM's 18.0%. The cost maps show why. In level-1
space, the latent distance to a fixed point in the maze is nearly flat everywhere except right
next to it, so a gradient planner has nothing to follow until it is already close. At levels 2
and 3 the distance rises smoothly along the corridors.

<Figure
  src="https://ai.thesatyajit.com/articles/h-jepa/fig4.png"
  alt="Three maze maps coloured by latent distance to a starred anchor point, one per level of a three-level H-JEPA. At level 1 almost the whole maze is uniformly red except a tiny blue spot at the star. At levels 2 and 3 a broad blue-to-yellow basin surrounds the star and the colour grades along the corridors."
  caption="Latent distance to the starred anchor in each level's space, one training seed: level-1 distances are almost constant away from the anchor, higher levels grade along the corridors (paper, Figure 7)."
/>

The second lever is time. A flat planner's cost to the final goal is a bad progress signal even
along an expert trajectory: the paper measures its monotonicity, the negative Spearman
correlation of cost with time, at 0.48 to 0.75 across the four simulated environments. The cost
of tracking the next level-2 subgoal over a short segment scores 0.89 to 1.00 for H-JEPA. HWM
gets nearly the same subgoal benefit, which is why it beats flat LeWM everywhere. The
representation lever is what H-JEPA adds on top, and it only shows up where the frequency gap
gave the upper levels something to throw away.

I like this section a lot. The cost-only ablation is the kind of experiment that tells a
practitioner something usable today: if you already have a hierarchical encoder, you can get
part of the gain with your existing flat planner by changing the ruler.

## Where it does not help

The paper puts its failures in the main text, and they are instructive.

Push-T is the clearest. Two levels match or slightly beat LeWM, 45.3% against 40.0%. Three
levels drop to 17.3% and four to 0.7%. The paper's explanation is data, and I think it is right.
A training sample must fit inside one episode, and a four-level sample needs 25 level-1 frames,
121 environment steps. Push-T episodes average 125 steps, so a four-level model sees only 14% of
the training clips LeWM sees per epoch, and three levels see 58%. HWM collapses the same way,
which points at the data rather than the representation.

Cube is subtler. The full hierarchy helps there, 60.0% at three levels, but the cost-only
ablation does not: the level-2 cost lowers the two-level model's flat success from 24.0% to
13.3%, and the level-3 and level-4 costs roughly halve success in the deeper models. With no
frequency gap there is no useful abstraction to measure in, so whatever Cube gains comes from
the subgoal lever alone. The authors say as much, and add that their results "do not establish
that selective abstraction alone causes the gains in the other environments." A careful
sentence, and I appreciated it.

Then the data-specialisation appendix, which deserves more attention than its position suggests.
Trained end to end, both levels must see the same data. Trained stagewise, each level can have
its own. On AntMaze, training level 1 on the noisy exploration data and level 2 on the
goal-directed "stitch" data reaches 67% hierarchical success, while end to end on the union of
both reaches 54%. The reverse pairing collapses to 2%. The low level wants coverage; the high
level wants trajectories that go somewhere. End-to-end training cannot express that preference.
The paper's headline recipe is end to end, but its own best AntMaze cell outside the main table
came from not doing it.

## DROID, and a collapse SIGReg cannot see

Every environment so far has a fixed camera and a fixed background. DROID does not: it is real
teleoperated Franka video where the scene, lighting, camera calibration and objects change every
episode. Train LeWM on it and planning fidelity is zero.

<Figure
  src="https://ai.thesatyajit.com/articles/h-jepa/fig6.png"
  alt="Left: bar chart of Fréchet fidelity on DROID: one-level LeWM without inverse dynamics sits at zero labelled slow-feature collapse; with inverse dynamics it reaches about +34; two-level HWM about +35; two-level H-JEPA about +40. Right: Fréchet fidelity against planner TFLOPs per episode on a log axis; H-JEPA's curve sits highest near 40 at under 100 TFLOPs, HWM and LeWM plus IDM lower, and frozen-encoder baselines JEPA-WM and V-JEPA 2-AC far to the right at 1,000 to 100,000 TFLOPs with lower fidelity."
  caption="DROID at 5 fps: without an inverse-dynamics term the flat model collapses to the zero-action floor; with it, a second level adds fidelity, at lower planner compute than the flat model and orders of magnitude less than the frozen-encoder world models (paper, Figure 9b and 9c; the two panels are cropped from the rendered figure, an overlapping (c) label removed and the 20 axis tick redrawn)."
/>

The appendix that explains this is the best writing in the paper. It names the failure
slow-feature collapse: the encoder learns to encode only what is constant within an episode,
the background, and ignores the moving arm. That solution has zero prediction loss, because
the identity predictor is perfect when nothing changes, and here is the uncomfortable part:
SIGReg cannot rule it out. Their Proposition 1 says it for any regulariser that only looks at
the marginal distribution of single embeddings. If backgrounds vary continuously across
episodes, you can map each background to a draw from exactly the target Gaussian. Every batch
then looks perfectly isotropic, and the encoder still says nothing about the robot.

Recall the `(T, B, D)` grouping in the SIGReg code. Each test sees one timestep across many
sequences, and a collapsed encoder gives one distinct, Gaussian-looking embedding per sequence.
The "slow-feature collapse" button in the SIGReg widget above is exactly that case: the same
cloud as the healthy one, the same statistic, and nothing that moves when the robot moves. The
appendix measures where the variance actually goes in trained latents: 75% of it lies between
episodes on DROID against 28% on Cube.

They test two fixes. Putting time into the sample axis of the test adds a barrier, but a small
one, and every run in their SIGReg sweep still collapses. A "time repulsion" term that rewards
latents for changing keeps the predictor off the identity but never clears the zero-action
floor. What works is an inverse-dynamics head that must regress the action from two consecutive
latents. A collapsed encoder carries no information about the action, so the head's loss has a
floor fixed by the data, 1 with normalised actions, which is 50 at the weight they use. The
appendix calls it "the only $\Omega(1)$ barrier here," and notes the cost: it needs action
labels. So the two-term recipe that works in simulation becomes, on real video, a three-term
one: prediction, SIGReg, and action regression with weight 100 at level 1 and 50 at level 2.

With that term, the numbers are modest and honest. The paper's bars read +34 for flat LeWM plus
inverse dynamics, +35 for HWM and +40 for H-JEPA, as medians over training and planner seeds.
The README reports the same models as means: 33.57%, 35.65% and 38.42%. H-JEPA gets there at
10.1 planner TFLOPs per episode against 14.5 for the flat model.

Two things limit how far I would take this. Fréchet fidelity is an offline metric: it compares
the planned end-effector path with the logged expert path, with 0% meaning the arm does not move
and 100% meaning it follows the expert. Nothing is executed on a robot. The paper calibrates the
metric in simulation and concludes it is a necessary condition for success on Cube, not a
predictor of it. Their conclusion states the open question plainly: "A next step is to test
whether these gains translate to closed-loop control on a physical robot." And the evaluation is
16 hand-picked clips with good lighting and large objects. A careful protocol, and a
small one.

<Figure
  src="https://ai.thesatyajit.com/articles/h-jepa/fig7.png"
  alt="Sixteen DROID clips in two columns, ordered by Fréchet fidelity from +75% down to the zero floor. Each clip shows a row of logged expert frames above a row of decoded planned frames, with every third planned frame ringed in blue as the level-2 macro-plan; the planned frames blur along the rollout."
  caption="What the DROID planner imagines on all 16 evaluation clips, ordered best to worst; the decoded plans move the gripper toward the object before the goal, and blur is the decoder's limit (paper, Figure 16)."
/>

## Two replies worth reading

Two replies under the announcement point at work the article should mention. Daniel Korchinski
linked his paper with Alessandro Favero and Matthieu Wyart,
[arXiv 2605.27734](https://arxiv.org/abs/2605.27734). It proves that on data from a
probabilistic context-free grammar, latent prediction recovers a hidden tree with a sample count
constant in its depth, where token-level learning needs exponentially many, and that data2vec
already does this implicitly. Its abstract ends: "This suggests that explicit stacking such as
H-JEPA is largely redundant." The theory is about static compositional data with no actions
and no planning. H-JEPA's evidence agrees with it in one respect: on Push-T and Cube, where there
is no separation of timescales, a separate representation per level adds little. Where the world
has a slow part and a fast part, and a planner needs a ruler for the slow part, the explicit
levels earned their keep. I do not think the two papers contradict each other; they disagree
about which data matters.

The other reply links [HDFlow](https://arxiv.org/abs/2605.04525), a hierarchical planner with a
diffusion model proposing latent subgoals and a rectified-flow model filling in trajectories,
and asks why this group never uses diffusion dynamics. Nobody had answered when I read the
thread. My reading of the design is that a deterministic predictor is what makes
gradient-through-the-rollout planning possible here; a stochastic generator would push you back
toward sampling planners. I am inferring this; the authors have not said it.

## What I would take from it

The idea that will last is the separation of the two levers. Measure distance where the
dynamics are slow, plan in time where the horizon is long, and check which one your problem
actually needs. The cost-projection experiment is cheap to run on any stack of encoders, and
the frequency gap is a cheap thing to compute on a dataset before you build anything.

The part I would be careful with is "end to end" as a virtue. It works, it is stable, and SIGReg
genuinely makes joint training of four latent spaces uneventful. But the paper's own appendix
shows that training each level on its own data beat it on AntMaze, and that on real video the
regulariser needs an action-supervised partner. The training is end to end; the safety net is
not SIGReg alone.

And the real-robot claim should wait for a real robot. The simulation results are strong and
well controlled. DROID is an offline proxy on 16 clips.

## How I checked

I read the arXiv HTML of 2610.06805v1, including every appendix, and shallow-cloned
`kevinghst/H-JEPA` at commit `840e76b`. Code quotes are from that commit with file and line. The
hierarchy, loss assembly, SIGReg quadrature, solver loop, macro-action clipping and per-level
configs were read from source; I did not run any of it. Checkpoint counts and sizes come from the
Hugging Face model API for `jepa-world-models/h-jepa`: 93 checkpoint files, 84 for simulation
and 9 for DROID, about 12.43 GB in all, matching the README's 9.7 GB and 2.7 GB. The SIGReg
widget is my own toy that reimplements the released formula on synthetic 2-D data; the depth
widget plots the paper's Table 7 horizons and the Figure 6 and Table 3 numbers.

A few things in the sources do not agree, none of them large. The four-level AntMaze hierarchical
result is 63.3% in the paper's Table 1 and 67.3% in Figure 6 and the README, so I quote the table
where I quote the table. The paper's DROID planning table lists 16 candidates at level 2, while
the released `droid_l2.yaml` and the README use 4. The DROID training corpus is 74,530 episodes
in the paper and 74,896 in the README's file list. `ARCHITECTURE.md` still describes a Push-T
example with stride 3 and 25 frames per step, which matches no released config. I could not check
any number that depends on running the planner, including all success rates and fidelities;
those are the authors'.
