# JEPA-Anything: one recipe for seven worlds, not one model

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/jepa-anything
> date: 2026-10-06
> tags: world-models, self-supervised-learning, representation-learning, pde, robotics, explainer

The launch post for JEPA-Anything says "one world model that works across molecules, cells, fluids,
patients, robots, or even weather." The paper ([arXiv 2609.20800](https://arxiv.org/abs/2609.20800),
from PhAI Labs, CUHK and others, code at [Gen-Verse/JEPA-Anything](https://github.com/Gen-Verse/JEPA-Anything))
is more careful than the post. Its own README says it in one line: "domains need not share a single
encoder or one set of model weights."

So the first thing to get right is what is shared. It is a training recipe and a latent interface,
not a network. I counted the released weights on the
[Hugging Face dataset](https://huggingface.co/datasets/Gen-Verse/jepa-anything): 36 checkpoint files,
one or more per task, from 1.9 MB to 419.2 MB, about 0.91 GB in all (measured, from the Hub API). A
weather checkpoint cannot read a molecule.

That is not a criticism of the idea, which is a small change to *how a JEPA predicts* that bolts onto
any existing JEPA. This piece explains it from the ground up, then checks the numbers.

<RepoCard repo="Gen-Verse/JEPA-Anything" />

## A JEPA in four parts

A joint-embedding predictive architecture learns by predicting one part of the world from another,
in representation space rather than in raw observation space. It has four moving parts:

1. A **context encoder** $f_\theta$ reads what is visible (unmasked patches, the past, the current
   state) and produces $z_c$.
2. A **target encoder** $f_{\bar\theta}$ reads the thing to be predicted and produces $z_t$.
3. A **predictor** maps $z_c$, plus a descriptor of *which* target is wanted (a patch position, a
   time step, an action), to a guess $\hat z_t$.
4. A **loss** pulls $\hat z_t$ toward $z_t$.

Why predict in latent space? Because a pixel loss pays for every pixel. Predict the next frame of a
pond and most of the loss is in the ripples, which are unpredictable and useless. A latent target lets
the target encoder decide what level of detail is worth predicting.

The catch is collapse. If the target encoder and the predictor are both learned, the loss has a
trivial minimum: map everything to the same vector. The common defences:

- **EMA target and stop-gradient** (I-JEPA, V-JEPA): the target encoder is an exponential moving
  average of the context encoder and gets no gradient, so it cannot chase the predictor.
- **Variance and covariance regularisers** (VICReg): penalise any embedding coordinate whose standard
  deviation across the batch falls below a floor, and penalise correlated coordinates.
- **Distributional regularisers** (SIGReg in [LeVJEPA](/articles/levjepa)): push the embedding
  distribution toward an isotropic Gaussian and drop the EMA branch entirely.

JEPA-Anything uses the first and borrows the spirit of the second. The target encoder is an EMA,
$\bar\theta \leftarrow m\bar\theta + (1-m)\theta$, with a stop-gradient on its output. On top of that
it adds two hinge-style variance floors, described below. [SiamJEPA](/articles/siamjepa) and
[NextLat](/articles/next-latent-world-models) are two other recent takes on the same family, if you
want a comparison of collapse machinery.

## The problem OPF is aimed at

A standard JEPA predicts its target with one predictor and one target vector. The paper's argument is
that a real target mixes several predictive modes: a future patient state carries the risks of
hundreds of events, a physical field has slow and fast modes. Squeezed into one regression target,
the easy or high-variance part dominates the gradient, several latent directions do the same job, and
the weaker structure gets conflicting updates.

Their fix is to make the predictor's capacity allocation explicit. They call it **orthogonal
predictive factorization** (OPF).

<Figure
  src="https://ai.thesatyajit.com/articles/jepa-anything/fig1.jpg"
  alt="Overview figure. Top: seven domain icons (vision, cells, clinical, robotics, molecules, physical fields, weather). Middle left: a standard JEPA with context and target encoders feeding one predictor and one monolithic target embedding. Middle right: JEPA-Anything, where the target representation is split into four factors F1 to F4, each with its own predictor, recombined into a complete latent world state. Bottom: three evaluation groups, terminal readout, latent world dynamics and scientific analysis, with example panels."
  caption="Standard JEPA predicts one monolithic target; JEPA-Anything projects the target into K orthogonal factors, predicts each with its own head, and synthesises a complete latent state. The lower panels are the paper's three evaluation groups (paper, Figure 1)."
/>

## How OPF works

Take a latent target $z_t \in \mathbb{R}^d$ from the EMA encoder and stop its gradient. Learn $K$
projector matrices $P_k \in \mathbb{R}^{d \times r}$ with $Kr = d$, so together they form a square
$d \times d$ matrix $P = [P_1, \ldots, P_K]$. Each projector reads off one block of coordinates:

$$
z_t^{(k)} = P_k^\top \operatorname{sg}(z_t), \qquad k = 1, \ldots, K.
$$

Each block gets its own predictor $q_k$, which sees the same context representation and the same
target descriptor: $\hat z_t^{(k)} = q_k(z_c, s_t)$. The prediction loss is a plain mean squared error
per block. It regresses magnitude as well as direction, because the blocks have to be put back
together later:

$$
\mathcal{L}_{\text{pred}} = \frac{1}{K|T|r} \sum_{t \in T} \sum_{k=1}^{K} \left\| \hat z_t^{(k)} - z_t^{(k)} \right\|_2^2 .
$$

Nothing so far stops all $K$ projectors from learning the same directions. That is the job of the
orthogonality loss, which asks each projector's columns to be orthonormal and different projectors to
be orthogonal to each other:

$$
\mathcal{L}_{\text{orth}} = \sum_{k} \left\| P_k^\top P_k - I_r \right\|_F^2 + \sum_{i < j} \left\| P_i^\top P_j \right\|_F^2 .
$$

The two activity terms are hinge penalties on standard deviations, measured across the mini-batch.
$\mathcal{L}_{\text{fac}}$ penalises any projected target coordinate whose standard deviation drops
below $\gamma_{\text{fac}}$; it shapes the projectors, since the target itself is stopped.
$\mathcal{L}_{\text{enc}}$ does the same for every coordinate of the online context representation,
which sends a direct anti-collapse gradient into the encoder. Each domain keeps whatever loss its
original implementation already had and adds the OPF terms:

$$
\mathcal{L}_{\text{train}} = \mathcal{L}_{\text{base}} + \mathcal{L}_{\text{pred}} + \lambda_{\text{orth}} \mathcal{L}_{\text{orth}} + \lambda_{\text{fac}} \mathcal{L}_{\text{fac}} + \lambda_{\text{enc}} \mathcal{L}_{\text{enc}} .
$$

The paper's appendix gives $\lambda_{\text{orth}}, \lambda_{\text{fac}}, \lambda_{\text{enc}} = 0.10, 0.05, 0.02$
for every dynamics task and 0.02, 0.10, 0.02 for the three locomotion tasks (reported, Table 10). The
released library's `jepa_anything_objective()` defaults all three weights to 1.0 and both floors to
0.1 (measured, `losses.py`), so if you use it, set them yourself.

### Putting the state back together

The predicted blocks are concatenated into $\hat u_t \in \mathbb{R}^d$ and turned back into one state
with the Moore-Penrose pseudoinverse of the analysis map:

$$
\hat z_t = \left(P^\top\right)^{\dagger} \hat u_t .
$$

When $P$ is exactly orthogonal, $(P^\top)^\dagger = P$ and this is just $\sum_k P_k \hat z_t^{(k)}$:
add up each block's contribution. The paper's Corollary 1 is the reason orthogonality matters here.
If the factor predictions are off by an error $e$, the state is off by $(P^\top)^{-1} e$, whose size is
at most $\|e\| / \sigma_{\min}(P)$. An orthogonal $P$ has every singular value equal to 1, so the
state error equals the factor error. A $P$ with two nearly parallel columns has a tiny
$\sigma_{\min}$, and a small factor error becomes a large state error.

The widget below is that argument in two dimensions: two projector directions, a fixed true state
$z$, and a factor error that pushes one coordinate up and the other down. Close the angle and watch
the pseudoinverse estimate fly off while the condition number climbs.

<SynthesisGeometry />

In words: at 90 degrees $\kappa_2(P)$ is 1 and a factor error of 0.10 is a state error of 0.10.
Shrink the angle and the worst-case amplification $1/\sigma_{\min}$ grows without bound. Transpose
synthesis, just summing the blocks, is wrong even with perfect factors once the projectors stop being
orthogonal, which is why the paper uses the pseudoinverse.

### Where actions and interventions go

For a world model, the next state depends on something the agent or experimenter does: an action, an
intervention label, a known forcing. The paper calls it $\xi_t$ and lets the domain adapter put it
either into the context tokens or into the target descriptor. A transition is then

$$
\hat z_{t+1} = \left(P^\top\right)^{\dagger} \left[ q_1(z_t, \xi_t, s_{t+1}); \ldots; q_K(z_t, \xi_t, s_{t+1}) \right] ,
$$

and applying it repeatedly gives a latent rollout. A planner scores rollouts with the domain's own
reward. In the locomotion experiments that planner is CEM: sample many action sequences, roll each
forward through the model, keep the first action of the best one.

After training, **representation mode** discards the EMA encoder, projectors and heads and puts a
readout on the online encoder; **operational mode** keeps them and feeds the synthesised state to a
decoder, a planner or the next rollout step.

## One interface, many adapters

Here is where "Anything" comes from. Each domain $\delta$ writes an adapter $\mathcal{A}_\delta$ that
turns a raw observation into content tokens $H$ and structural descriptors $S$ (a patch coordinate, a
time stamp, a graph position, or nothing), and a view sampler $\mathcal{V}_\delta$ that chooses which
tokens are context and which are target. Everything after that follows the one algorithm. The paper's
Table 1 is explicit about the boundary: raw representation, tokenisation, context-target semantics,
encoder architecture (ViT, Transformer, GNN, MLP) and the readout are domain-specific. The EMA scheme,
the factorised predictors, the regularisers and the additive loss are shared.

<DomainAtlas />

The table that matters, condensed from the paper's Table 2:

| Domain | Context → target | Backbone | $d;\ K \times r$ |
|---|---|---|---|
| Vision (MuJoCo scenes) | visible patches → masked patch states | DINOv3 ViT-S, SigLIP2 Base | 384; 4×96 and 768; 4×192 |
| Single cell (kidney, PBMC-10K, Adamson, Norman) | masked expression → complete cell state | scGPT | 512; 4×128 (Norman) |
| Clinical (UK Biobank cohort) | patient history → future patient state | GPT-2 small | 768; 4×192 |
| Interventions (CITRIS Pong) | observation + intervention → next state | matched encoder and predictor | 160; 5×32 |
| Dynamics (CausalWorld, DMC, PDEBench, WeatherBench2) | state, frames, action or forcing → next state or field | matched task-specific | 128; 4×32 |
| Locomotion (Hopper, Walker2d, HalfCheetah) | state + action → future state | latent dynamics + CEM | 32; 4×8 |
| Molecules (water, quartz, paracetamol, benzene) | atomic state → future state | TrajCast-style $O(3)$-equivariant | 64; 4×16 channels |

The molecular case is the neatest adaptation. An equivariant network carries vector features as
$l = 1$ irreps with three components each, and a projector that mixed those components would break
rotation equivariance. So OPF acts only on the 64-channel multiplicity axis, as $P_k \otimes I_3$:
the same channel projection applied identically to x, y and z.

## What is in the released weights

I did not run the authors' code. I read the checkpoint pickles with a restricted unpickler that never
imports a class, and the projector tensors as raw float32 arrays.

The dynamics checkpoints are small multilayer perceptrons. The CausalWorld model is an encoder from a
56-number state through two 256-wide layers to $d = 128$, an identical EMA teacher, a shared trunk
whose first layer reads 137 numbers (the 128-wide latent plus a 9-number action), four 128 → 32 heads,
a decoder, and a projector tensor of shape (4, 128, 32). That is 393,400 online parameters
(measured). The WeatherBench2 model has the same layout with a 6,144-number flattened field as input:
3,514,496 online parameters, a 20.8 MB file (measured). That is the paper's "shared trunk followed by
branch-specific heads" option, and it is a long way from a weather model in the GraphCast sense. The
paper compares it only with a matched monolithic JEPA, never with a weather-specific forecaster.

The projectors are where it gets interesting. The paper's mechanism audit (Table 5) reports, on
CITRIS Pong, a cross-factor overlap of $5.18 \times 10^{-16}$, $\sigma_{\min}(P) = 0.999989$ and
$\kappa_2(P) = 1.00005$ for orthogonal factorization, against an overlap of 0.4550 and $\kappa_2 = 438.52$
for unconstrained multi-head prediction (reported). An overlap at the level of float rounding error
is what an exactly orthonormalised basis gives, and the paper does say that "the controlled geometry
analysis uses this strict decomposition", a QR factorisation. The predictive models instead use
soft-regularised projectors and pseudoinverse synthesis. So I measured those:

| Released checkpoint | $\sigma_{\min}(P)$ | $\kappa_2(P)$ | worst-case error amplification $1/\sigma_{\min}$ |
|---|---|---|---|
| HalfCheetah, 20-step | 1.0000 | 1.000 | 1.0x |
| CausalWorld pushing | 0.9312 | 1.140 | 1.07x |
| CITRIS six-step, five seeds | 0.6742 – 0.7280 | 1.683 – 1.819 | 1.37 – 1.48x |
| PDEBench shallow water | 0.4788 | 3.988 | 2.1x |
| PDEBench Burgers | 0.2596 | 6.679 | 3.9x |
| APEBench Burgers | 0.2764 | 7.969 | 3.6x |
| WeatherBench2 | 0.1784 | 9.175 | 5.6x |
| Synthetic PDE | 0.0218 | 63.287 | 46x |
| CITRIS legacy (three-channel protocol) | 0.0082 | 227.491 | 122x |

Singular values and condition numbers are measured; the amplification column is reasoned, from the
paper's own bound. The locomotion projector is orthogonal to float precision, which suggests it was
retracted with QR (the library has a `qr_retraction` mode). CausalWorld is close behind; the rest drift further. Five of the nine have a
condition number above 6, and the Synthetic PDE projector is closer to the paper's "unconstrained
heads" than to its "orthogonal factorization" column. The CITRIS legacy checkpoint is labelled on its
card as an earlier protocol, so it is not evidence against the paper, but it is in the release.

Two honest readings. One: the soft penalty with $\lambda_{\text{orth}} = 0.10$ is too weak to hold
these projectors near orthogonal, and Table 5 describes the idealised geometry rather than the trained
models. Two: the method still beats the matched JEPA on Burgers and WeatherBench2 with $\kappa$ near
7 to 9, so whatever is doing the work there is not exact orthogonality. More heads, the variance
floors and the per-block loss are all candidates. The paper does not separate them, apart from the
CITRIS audit against an unconstrained multi-head baseline, and that audit reports geometry, not
prediction error.

## Results, against the matched JEPA

Every quantitative comparison holds the adapter, encoder, sampler, budget, split and readout fixed
and swaps the monolithic JEPA term for the OPF terms. That is a clean ablation, and it is the right
experiment for the paper's claim. It is not a comparison against the best model in each field, and
the post's "works across" should be read with that in mind.

### Terminal readout: vision, cells, patients

- **Vision binding** (Table 3, reported). On DINOv3, held-out-cell accuracy goes .572 → .581 and
  collapse rate .426 → .417 against standard JEPA; the frozen checkpoint with the same learned-grid
  readout already scores .569. On SigLIP2 it is .483 → .490. About one point, over ten seeds, no spread reported.
- **Single cell** (Table 4, reported). PBMC zero-shot AvgBIO is 0.5288 for scGPT, 0.7194 for
  Cell-JEPA and 0.7752 for JEPA-Anything; Norman perturbation Pearson is 0.631, 0.787 and 0.814. Most
  of the jump comes from going latent at all (scGPT to Cell-JEPA); OPF adds a smaller but consistent
  step on top.
- **Clinical** (Figure 3, reported). Mean PRAUC over more than 1,000 future events is 0.718 against
  0.711 for the matched JEPA, with a cross-attention fusion model at 0.702, Delphi at 0.689 and
  XGBoost at 0.667. The 0.007 gap (reasoned) is real money in a clinical ranking but small next to the
  spread between methods; no seed variance is shown.

<Figure
  src="https://ai.thesatyajit.com/articles/jepa-anything/fig3.png"
  alt="Horizontal lollipop chart of mean PRAUC across more than 1,000 clinical events: JEPA-Anything 0.718, Standard JEPA 0.711, Cross-attention fusion 0.702, Delphi 0.689, Prophet 0.680, XGBoost 0.667, Qwen3-0.8B 0.651, Qwen2.5-0.5B 0.648, Random Forest 0.602."
  caption="Mean PRAUC over the full clinical event vocabulary on one fixed patient-level split; JEPA-Anything leads the matched JEPA by 0.007 (paper, Figure 3)."
/>

### Interventions and dynamics

On CITRIS Interventional Pong the model sees a frame plus a label of which game variables were
externally changed, and predicts the next frame. Single-intervention one-step MSE falls from 0.009541
to 0.006218, combined interventions (a combination withheld in training) from 0.009441 to 0.008223,
and six-step free rollout from 0.009478 to 0.008665: reductions of 34.83%, 12.90% and 8.58% (reported).

<Figure
  src="https://ai.thesatyajit.com/articles/jepa-anything/fig4.png"
  alt="Dumbbell chart of five-seed mean four-channel MSE on CITRIS Interventional Pong. Single intervention one-step: standard JEPA 0.009541, OPF 0.006218, minus 34.83 percent. Combined intervention one-step: 0.009441 versus 0.008223, minus 12.90 percent. Six-step free rollout: 0.009478 versus 0.008665, minus 8.58 percent."
  caption="The headline 34.8% is the single-intervention case; on withheld combinations and on six-step rollout the gain is 12.90% and 8.58% (paper, Figure 4)."
/>

The released CITRIS weights are the five-seed models trained with six-step feedback. Their card
reports one-step MSE of 0.00952107 against 0.00682598 for single interventions, which is a 28.3%
reduction (reasoned), and 0.00942106 against 0.00822350 for combinations. The paper's 34.8% comes from
a training condition whose weights are not released.

The ten-task dynamics benchmark is the broadest evidence. Relative MSE reduction against the matched
JEPA ranges from 39.7% on PDEBench Burgers and 39.3% on shallow water down to 3.6% on a synthetic PDE
and 0.4% on DMC pixel dynamics (reported, Figure 5). The tenth task is CausalWorld closed-loop control,
where the mean return is −0.055 against −0.087: both negative, with error bars that overlap almost
entirely. The abstract's "improves reported metrics on all 10 dynamics tasks" is true of the means.

<Figure
  src="https://ai.thesatyajit.com/articles/jepa-anything/fig5.png"
  alt="Left: bar chart of relative MSE reduction versus standard JEPA for nine prediction tasks: Burgers 39.7 percent, shallow water 39.3, CausalWorld state 20.6, reaction-diffusion 15.2, WeatherBench2 10.5, synthetic pixel dynamics 8.1, DMC-VB held-out episodes 7.5, synthetic PDE 3.6, DMC pixel dynamics 0.4. Right: CausalWorld closed-loop return, standard JEPA minus 0.087 and JEPA-Anything minus 0.055, with wide overlapping error bars."
  caption="Nine prediction tasks improve by 0.4% to 39.7%; the control task's two returns overlap within their error bars (paper, Figure 5)."
/>

The rollout table shows the errors do not compound faster: on PDEBench Burgers, MSE goes 0.001830 →
0.006369 from step 1 to step 6 for the JEPA and 0.001101 → 0.004014 for JEPA-Anything (reported, Table
6). On the APEBench protocol, Burgers held-out late-state error drops by about 49.5% and six-step
rollout by 44.7%, and Kuramoto–Sivashinsky by 13.2%, in every seed (reported). The capacity-matched
50-step Burgers test is the sobering one: the gain is −5.84% at 20 steps and −3.15% at 50, narrowing as
the horizon grows.

The Hugging Face folder for this benchmark holds 15 task checkpoints, including CLEVRER, Navier-Stokes
and a 2-D reaction-diffusion run that the paper does not report (measured). The paper does not say
whether those were evaluated.

### Planning and locomotion

This is the one place the paper reports a loss. Paired CEM return favours standard JEPA on Hopper
(−1.24), is +0.55 on Walker2d with a confidence interval that crosses zero, and +10.26 on HalfCheetah
(reported, Figure 6). The authors say it plainly: planning performance is environment-dependent.

<Figure
  src="https://ai.thesatyajit.com/articles/jepa-anything/fig6.png"
  alt="Forest plot of paired CEM-return difference, JEPA-Anything minus standard JEPA, with 95 percent confidence intervals over five seeds: Hopper minus 1.24, Walker2d plus 0.55 with interval crossing zero, HalfCheetah plus 10.26."
  caption="Capacity-matched planning: one loss, one tie and one clear win (paper, Figure 6)."
/>

The factor ablation is better evidence that the four blocks are each doing something. Replacing any
one factor by its training-set mean raises 20-step rollout error in all three environments, and on
HalfCheetah masking F1 or F3 lowers return in all five seeds (reported). The single-factor probes
find each block linearly predicts a different body variable best on Hopper, with $R^2$ between 0.294
and 0.424: modest, and the paper is careful to call them "motion preferences" rather than labels.

### Molecules

With a TrajCast-style equivariant forecaster, JEPA-Anything has the lowest one-step displacement MAE
and 100-step RMSD on all four systems (reported, Table 9). Paracetamol 100-step RMSD is 3.155 Å from
scratch, 1.868 Å with JEPA pretraining and 1.776 Å with OPF. Water goes 2.536 → 2.459 Å. As elsewhere,
the big step is from scratch to any JEPA; OPF is a smaller step on top. The released molecular
checkpoint is a 3.1 MB TiTo/PaiNN model with a latent flow-matching module (measured, from its file
list and card), not the TrajCast-style setup in the table.

## The science claims

Group III is where the paper reaches furthest, and where I can check least.

The cancer case says factor analysis "nominated IL-18 combined with NT5E/CD73 blockade", and wet-lab
studies in co-culture, patient-derived organoids, tumour fragments and mice supported it. The paper
gives one figure of selected organoid and fragment results and no protocol, no list of other
candidates the factors ranked, and no description of the "domain-specific intervention analysis"
that turned factor coordinates into a drug pair. I cannot evaluate it from what is published.

The orbital case is checkable in principle. From simulated position-velocity trajectories, spectral
analysis of the latent factor coordinates finds frequencies that scale with semimajor axis as
$f \propto a^{-3/2}$, with a fitted slope of −1.4991 and $R^2 = 0.9999999$ (reported).

<Figure
  src="https://ai.thesatyajit.com/articles/jepa-anything/fig9.png"
  alt="Three panels. Left: frequency-domain ridge of learned spectral modes lying on the Kepler line f equals one over two pi times a to the minus three halves. Top right: recovered power law with fit slope minus 1.4991 and R squared 0.9999999. Bottom right: scale-compensated invariant staying within 0.169 percent of 1. Header: seed 0 selected by the lowest run-level median frequency error."
  caption="Learned latent modes on simulated orbits follow Kepler's third law; the header notes the run shown was the best of its seeds (paper, Figure 9)."
/>

Two things temper it. The figure's own header says the run shown was "selected by the lowest run-level
median frequency error", so it is the best seed, not a typical one. And the simulated orbits obey
Kepler's law by construction, so a model that predicts them well must carry each orbit's period
somewhere in its state (reasoned). What the result shows is that OPF's factor coordinates expose that
period cleanly enough for a spectral fit. That is a useful property of the coordinates. It is not the
model discovering a law it was not shown.

## What the repository is

The GitHub repo is a library, not a training codebase. `jepa-anything-core` has 2,069 lines of Python
across five modules: the projection and pseudoinverse synthesis, the four loss terms, matched
baselines with parameter and FLOP accounting, and geometry audits (measured). A
`jepa-anything-skill` folder holds a design validator and a scaffold generator meant to be driven by
a coding agent, and one synthetic recipe that checks interfaces. Its README is direct about it: the
example "does not train a model or produce benchmark performance", and the checkpoint manifest
records an untrained example with no performance claims. The domain training code is not there. The
per-task model code and checkpoints are on the Hugging Face dataset, Apache-2.0, alongside inference
scripts.

Two more mismatches between paper and release. The vision
checkpoints are a 192-wide ViT on 64×64 Shapes3D images, 13.4 MB, not the DINOv3 ViT-S and SigLIP2
backbones in Table 3 (measured). And the clinical model has no checkpoint at all; its row on the
dataset card reads "progressive checkpoint release".

## What to take from it

Strip away the "world model everything" framing and JEPA-Anything is a clean, portable change to the
predictor of any JEPA: split the target into $K$ blocks with learned, softly orthogonal projectors;
give each block its own small head; keep every coordinate alive with variance floors; stitch the
blocks back together with a pseudoinverse so the full state survives for rollouts and planning. It
costs almost nothing. The locomotion models differ from their dense baselines by under 0.3% in
parameters and 2% in prediction-head FLOPs (reported).

The evidence that it helps is broad and mostly shallow. It beats a matched monolithic JEPA in nearly
every comparison, sometimes by 40%, often by a point or less, and once it loses. It is never compared
with the strongest specialist model in a field. The geometry that justifies the design is exact in
the paper's audit and approximate, sometimes loose, in the shipped weights. If you already run a JEPA
for dynamics or rollouts, OPF is a cheap ablation to try. If you read the launch post as one network
that forecasts weather and molecules, that is not what the paper built, and the paper says so.

For a different direction on the same family, [LeVJEPA](/articles/levjepa) removes the EMA target
altogether, and [Cosmos 3](/articles/cosmos-world-model) is what a world model looks like when it is
built as one large generative network rather than as a recipe.
