~/satyajit

JEPA-Anything: one recipe for seven worlds, not one model

mdjsonmcp

2026-10-06 · 20 min · world-models · self-supervised-learning · representation-learning · pde · robotics · explainer

The launch post for JEPA-Anything says "one world model that works across molecules, cells, fluids, patients, robots, or even weather." The paper (arXiv 2609.20800, from PhAI Labs, CUHK and others, code at Gen-Verse/JEPA-Anything) is more careful than the post. Its own README says it in one line: "domains need not share a single encoder or one set of model weights."

So the first thing to get right is what is shared. It is a training recipe and a latent interface, not a network. I counted the released weights on the Hugging Face dataset: 36 checkpoint files, one or more per task, from 1.9 MB to 419.2 MB, about 0.91 GB in all (measured, from the Hub API). A weather checkpoint cannot read a molecule.

That is not a criticism of the idea, which is a small change to how a JEPA predicts that bolts onto any existing JEPA. This piece explains it from the ground up, then checks the numbers.

Gen-Verse/JEPA-Anything@c6e6c88 · snapshot 2026-10-06
tracked files
56
license
Apache-2.0
branch
main
tests
8 files
source
264.3 kB
commit date
2026-09-18
source by language
Python262.7 kB(15)Makefile1.7 kB(1)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

local clone, 2026-10-06 at c6e6c88 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount

shallow clone: counts describe the pinned tree, not the history

A JEPA in four parts

A joint-embedding predictive architecture learns by predicting one part of the world from another, in representation space rather than in raw observation space. It has four moving parts:

  1. A context encoder fθf_\theta reads what is visible (unmasked patches, the past, the current state) and produces zcz_c.
  2. A target encoder fθˉf_{\bar\theta} reads the thing to be predicted and produces ztz_t.
  3. A predictor maps zcz_c, plus a descriptor of which target is wanted (a patch position, a time step, an action), to a guess z^t\hat z_t.
  4. A loss pulls z^t\hat z_t toward ztz_t.

Why predict in latent space? Because a pixel loss pays for every pixel. Predict the next frame of a pond and most of the loss is in the ripples, which are unpredictable and useless. A latent target lets the target encoder decide what level of detail is worth predicting.

The catch is collapse. If the target encoder and the predictor are both learned, the loss has a trivial minimum: map everything to the same vector. The common defences:

JEPA-Anything uses the first and borrows the spirit of the second. The target encoder is an EMA, θˉ←mθˉ+(1−m)θ\bar\theta \leftarrow m\bar\theta + (1-m)\theta, with a stop-gradient on its output. On top of that it adds two hinge-style variance floors, described below. SiamJEPA and NextLat are two other recent takes on the same family, if you want a comparison of collapse machinery.

The problem OPF is aimed at

A standard JEPA predicts its target with one predictor and one target vector. The paper's argument is that a real target mixes several predictive modes: a future patient state carries the risks of hundreds of events, a physical field has slow and fast modes. Squeezed into one regression target, the easy or high-variance part dominates the gradient, several latent directions do the same job, and the weaker structure gets conflicting updates.

Their fix is to make the predictor's capacity allocation explicit. They call it orthogonal predictive factorization (OPF).

Overview figure. Top: seven domain icons (vision, cells, clinical, robotics, molecules, physical fields, weather). Middle left: a standard JEPA with context and target encoders feeding one predictor and one monolithic target embedding. Middle right: JEPA-Anything, where the target representation is split into four factors F1 to F4, each with its own predictor, recombined into a complete latent world state. Bottom: three evaluation groups, terminal readout, latent world dynamics and scientific analysis, with example panels.
Standard JEPA predicts one monolithic target; JEPA-Anything projects the target into K orthogonal factors, predicts each with its own head, and synthesises a complete latent state. The lower panels are the paper's three evaluation groups (paper, Figure 1).

How OPF works

Take a latent target zt∈Rdz_t \in \mathbb{R}^d from the EMA encoder and stop its gradient. Learn KK projector matrices Pk∈Rd×rP_k \in \mathbb{R}^{d \times r} with Kr=dKr = d, so together they form a square d×dd \times d matrix P=[P1,…,PK]P = [P_1, \ldots, P_K]. Each projector reads off one block of coordinates:

zt(k)=Pk⊤sg⁡(zt),k=1,…,K.z_t^{(k)} = P_k^\top \operatorname{sg}(z_t), \qquad k = 1, \ldots, K.

Each block gets its own predictor qkq_k, which sees the same context representation and the same target descriptor: z^t(k)=qk(zc,st)\hat z_t^{(k)} = q_k(z_c, s_t). The prediction loss is a plain mean squared error per block. It regresses magnitude as well as direction, because the blocks have to be put back together later:

Lpred=1K∣T∣r∑t∈T∑k=1K∥z^t(k)−zt(k)∥22.\mathcal{L}_{\text{pred}} = \frac{1}{K|T|r} \sum_{t \in T} \sum_{k=1}^{K} \left\| \hat z_t^{(k)} - z_t^{(k)} \right\|_2^2 .

Nothing so far stops all KK projectors from learning the same directions. That is the job of the orthogonality loss, which asks each projector's columns to be orthonormal and different projectors to be orthogonal to each other:

Lorth=∑k∥Pk⊤Pk−Ir∥F2+∑i<j∥Pi⊤Pj∥F2.\mathcal{L}_{\text{orth}} = \sum_{k} \left\| P_k^\top P_k - I_r \right\|_F^2 + \sum_{i < j} \left\| P_i^\top P_j \right\|_F^2 .

The two activity terms are hinge penalties on standard deviations, measured across the mini-batch. Lfac\mathcal{L}_{\text{fac}} penalises any projected target coordinate whose standard deviation drops below γfac\gamma_{\text{fac}}; it shapes the projectors, since the target itself is stopped. Lenc\mathcal{L}_{\text{enc}} does the same for every coordinate of the online context representation, which sends a direct anti-collapse gradient into the encoder. Each domain keeps whatever loss its original implementation already had and adds the OPF terms:

Ltrain=Lbase+Lpred+λorthLorth+λfacLfac+λencLenc.\mathcal{L}_{\text{train}} = \mathcal{L}_{\text{base}} + \mathcal{L}_{\text{pred}} + \lambda_{\text{orth}} \mathcal{L}_{\text{orth}} + \lambda_{\text{fac}} \mathcal{L}_{\text{fac}} + \lambda_{\text{enc}} \mathcal{L}_{\text{enc}} .

The paper's appendix gives λorth,λfac,λenc=0.10,0.05,0.02\lambda_{\text{orth}}, \lambda_{\text{fac}}, \lambda_{\text{enc}} = 0.10, 0.05, 0.02 for every dynamics task and 0.02, 0.10, 0.02 for the three locomotion tasks (reported, Table 10). The released library's jepa_anything_objective() defaults all three weights to 1.0 and both floors to 0.1 (measured, losses.py), so if you use it, set them yourself.

Putting the state back together

The predicted blocks are concatenated into u^t∈Rd\hat u_t \in \mathbb{R}^d and turned back into one state with the Moore-Penrose pseudoinverse of the analysis map:

z^t=(P⊤)†u^t.\hat z_t = \left(P^\top\right)^{\dagger} \hat u_t .

When PP is exactly orthogonal, (P⊤)†=P(P^\top)^\dagger = P and this is just ∑kPkz^t(k)\sum_k P_k \hat z_t^{(k)}: add up each block's contribution. The paper's Corollary 1 is the reason orthogonality matters here. If the factor predictions are off by an error ee, the state is off by (P⊤)−1e(P^\top)^{-1} e, whose size is at most ∥e∥/σmin⁡(P)\|e\| / \sigma_{\min}(P). An orthogonal PP has every singular value equal to 1, so the state error equals the factor error. A PP with two nearly parallel columns has a tiny σmin⁡\sigma_{\min}, and a small factor error becomes a large state error.

The widget below is that argument in two dimensions: two projector directions, a fixed true state zz, and a factor error that pushes one coordinate up and the other down. Close the angle and watch the pseudoinverse estimate fly off while the condition number climbs.

state synthesis in 2-D: d = 2, K = 2, r = 1κ₂(P) = 1.000
p₁p₂zP(u+e)(Pᵀ)⁺(u+e)angle between p₁ and p₂90.00°cross-factor overlap ‖p₁ᵀp₂‖²0.0000σ_min(P)1.0000condition number κ₂(P)1.000factor error ‖e‖0.10pinv state error ‖ẑ − z‖0.100 amplification ‖ẑ − z‖ / ‖e‖1.00× worst case 1 / σ_min1.00×transpose state error0.100

The error vector pushes factor 1 up and factor 2 down by the same amount. At 90° both syntheses land on the same point and the state error equals ‖e‖. Close the angle and the pseudoinverse still recovers z exactly when e = 0, but every unit of factor error becomes up to 1/σ_min units of state error; transpose synthesis is wrong even with perfect factors. Preset: the zero-penalty limit of Proposition 1; the toy matches its condition number only. In two dimensions a κ of 438.52 is an angle of about a quarter of a degree.

In words: at 90 degrees κ2(P)\kappa_2(P) is 1 and a factor error of 0.10 is a state error of 0.10. Shrink the angle and the worst-case amplification 1/σmin⁡1/\sigma_{\min} grows without bound. Transpose synthesis, just summing the blocks, is wrong even with perfect factors once the projectors stop being orthogonal, which is why the paper uses the pseudoinverse.

Where actions and interventions go

For a world model, the next state depends on something the agent or experimenter does: an action, an intervention label, a known forcing. The paper calls it ξt\xi_t and lets the domain adapter put it either into the context tokens or into the target descriptor. A transition is then

z^t+1=(P⊤)†[q1(zt,ξt,st+1);…;qK(zt,ξt,st+1)],\hat z_{t+1} = \left(P^\top\right)^{\dagger} \left[ q_1(z_t, \xi_t, s_{t+1}); \ldots; q_K(z_t, \xi_t, s_{t+1}) \right] ,

and applying it repeatedly gives a latent rollout. A planner scores rollouts with the domain's own reward. In the locomotion experiments that planner is CEM: sample many action sequences, roll each forward through the model, keep the first action of the best one.

After training, representation mode discards the EMA encoder, projectors and heads and puts a readout on the online encoder; operational mode keeps them and feeds the synthesised state to a decoder, a planner or the next rollout step.

One interface, many adapters

Here is where "Anything" comes from. Each domain δ\delta writes an adapter Aδ\mathcal{A}_\delta that turns a raw observation into content tokens HH and structural descriptors SS (a patch coordinate, a time stamp, a graph position, or nothing), and a view sampler Vδ\mathcal{V}_\delta that chooses which tokens are context and which are target. Everything after that follows the one algorithm. The paper's Table 1 is explicit about the boundary: raw representation, tokenisation, context-target semantics, encoder architecture (ViT, Transformer, GNN, MLP) and the readout are domain-specific. The EMA scheme, the factorised predictors, the regularisers and the additive loss are shared.

one recipe, seven pipelinesdomain-specific·shared OPF core
raw inputMuJoCo scenesadapter A_δpatches + 2-D positionencoderDINOv3 ViT-SOPF cored; K×r = 384; 4×96q1q2q3q4ẑEMA target · pinv synthesisreadoutbinding probecontext → target: masked patch states
INJ held-out-cell accuracy (Table 3)(higher is better, reported)
matched JEPA .572
JEPA-Anything .581

released weights, measured: Shapes3D checkpoints, a 192-wide ViT on 64×64 images, 13.4 MB: not the Table 3 backbones

Pick a domain. Everything drawn in purple changes: the input, the adapter that turns it into tokens, the encoder family, the readout. The green block is drawn the same every time because it is the paper’s actual claim: an EMA target, K projectors, K small predictors and a pseudoinverse that stitches their outputs back into one state. Only its widths change. Every domain trains its own weights.

The table that matters, condensed from the paper's Table 2:

DomainContext → targetBackboned; K×rd;\ K \times r
Vision (MuJoCo scenes)visible patches → masked patch statesDINOv3 ViT-S, SigLIP2 Base384; 4×96 and 768; 4×192
Single cell (kidney, PBMC-10K, Adamson, Norman)masked expression → complete cell statescGPT512; 4×128 (Norman)
Clinical (UK Biobank cohort)patient history → future patient stateGPT-2 small768; 4×192
Interventions (CITRIS Pong)observation + intervention → next statematched encoder and predictor160; 5×32
Dynamics (CausalWorld, DMC, PDEBench, WeatherBench2)state, frames, action or forcing → next state or fieldmatched task-specific128; 4×32
Locomotion (Hopper, Walker2d, HalfCheetah)state + action → future statelatent dynamics + CEM32; 4×8
Molecules (water, quartz, paracetamol, benzene)atomic state → future stateTrajCast-style O(3)O(3)-equivariant64; 4×16 channels

The molecular case is the neatest adaptation. An equivariant network carries vector features as l=1l = 1 irreps with three components each, and a projector that mixed those components would break rotation equivariance. So OPF acts only on the 64-channel multiplicity axis, as Pk⊗I3P_k \otimes I_3: the same channel projection applied identically to x, y and z.

What is in the released weights

I did not run the authors' code. I read the checkpoint pickles with a restricted unpickler that never imports a class, and the projector tensors as raw float32 arrays.

The dynamics checkpoints are small multilayer perceptrons. The CausalWorld model is an encoder from a 56-number state through two 256-wide layers to d=128d = 128, an identical EMA teacher, a shared trunk whose first layer reads 137 numbers (the 128-wide latent plus a 9-number action), four 128 → 32 heads, a decoder, and a projector tensor of shape (4, 128, 32). That is 393,400 online parameters (measured). The WeatherBench2 model has the same layout with a 6,144-number flattened field as input: 3,514,496 online parameters, a 20.8 MB file (measured). That is the paper's "shared trunk followed by branch-specific heads" option, and it is a long way from a weather model in the GraphCast sense. The paper compares it only with a matched monolithic JEPA, never with a weather-specific forecaster.

The projectors are where it gets interesting. The paper's mechanism audit (Table 5) reports, on CITRIS Pong, a cross-factor overlap of 5.18×10−165.18 \times 10^{-16}, σmin⁡(P)=0.999989\sigma_{\min}(P) = 0.999989 and κ2(P)=1.00005\kappa_2(P) = 1.00005 for orthogonal factorization, against an overlap of 0.4550 and κ2=438.52\kappa_2 = 438.52 for unconstrained multi-head prediction (reported). An overlap at the level of float rounding error is what an exactly orthonormalised basis gives, and the paper does say that "the controlled geometry analysis uses this strict decomposition", a QR factorisation. The predictive models instead use soft-regularised projectors and pseudoinverse synthesis. So I measured those:

Released checkpointσmin⁡(P)\sigma_{\min}(P)κ2(P)\kappa_2(P)worst-case error amplification 1/σmin⁡1/\sigma_{\min}
HalfCheetah, 20-step1.00001.0001.0x
CausalWorld pushing0.93121.1401.07x
CITRIS six-step, five seeds0.6742 – 0.72801.683 – 1.8191.37 – 1.48x
PDEBench shallow water0.47883.9882.1x
PDEBench Burgers0.25966.6793.9x
APEBench Burgers0.27647.9693.6x
WeatherBench20.17849.1755.6x
Synthetic PDE0.021863.28746x
CITRIS legacy (three-channel protocol)0.0082227.491122x

Singular values and condition numbers are measured; the amplification column is reasoned, from the paper's own bound. The locomotion projector is orthogonal to float precision, which suggests it was retracted with QR (the library has a qr_retraction mode). CausalWorld is close behind; the rest drift further. Five of the nine have a condition number above 6, and the Synthetic PDE projector is closer to the paper's "unconstrained heads" than to its "orthogonal factorization" column. The CITRIS legacy checkpoint is labelled on its card as an earlier protocol, so it is not evidence against the paper, but it is in the release.

Two honest readings. One: the soft penalty with λorth=0.10\lambda_{\text{orth}} = 0.10 is too weak to hold these projectors near orthogonal, and Table 5 describes the idealised geometry rather than the trained models. Two: the method still beats the matched JEPA on Burgers and WeatherBench2 with κ\kappa near 7 to 9, so whatever is doing the work there is not exact orthogonality. More heads, the variance floors and the per-block loss are all candidates. The paper does not separate them, apart from the CITRIS audit against an unconstrained multi-head baseline, and that audit reports geometry, not prediction error.

Results, against the matched JEPA

Every quantitative comparison holds the adapter, encoder, sampler, budget, split and readout fixed and swaps the monolithic JEPA term for the OPF terms. That is a clean ablation, and it is the right experiment for the paper's claim. It is not a comparison against the best model in each field, and the post's "works across" should be read with that in mind.

Terminal readout: vision, cells, patients

Horizontal lollipop chart of mean PRAUC across more than 1,000 clinical events: JEPA-Anything 0.718, Standard JEPA 0.711, Cross-attention fusion 0.702, Delphi 0.689, Prophet 0.680, XGBoost 0.667, Qwen3-0.8B 0.651, Qwen2.5-0.5B 0.648, Random Forest 0.602.
Mean PRAUC over the full clinical event vocabulary on one fixed patient-level split; JEPA-Anything leads the matched JEPA by 0.007 (paper, Figure 3).

Interventions and dynamics

On CITRIS Interventional Pong the model sees a frame plus a label of which game variables were externally changed, and predicts the next frame. Single-intervention one-step MSE falls from 0.009541 to 0.006218, combined interventions (a combination withheld in training) from 0.009441 to 0.008223, and six-step free rollout from 0.009478 to 0.008665: reductions of 34.83%, 12.90% and 8.58% (reported).

Dumbbell chart of five-seed mean four-channel MSE on CITRIS Interventional Pong. Single intervention one-step: standard JEPA 0.009541, OPF 0.006218, minus 34.83 percent. Combined intervention one-step: 0.009441 versus 0.008223, minus 12.90 percent. Six-step free rollout: 0.009478 versus 0.008665, minus 8.58 percent.
The headline 34.8% is the single-intervention case; on withheld combinations and on six-step rollout the gain is 12.90% and 8.58% (paper, Figure 4).

The released CITRIS weights are the five-seed models trained with six-step feedback. Their card reports one-step MSE of 0.00952107 against 0.00682598 for single interventions, which is a 28.3% reduction (reasoned), and 0.00942106 against 0.00822350 for combinations. The paper's 34.8% comes from a training condition whose weights are not released.

The ten-task dynamics benchmark is the broadest evidence. Relative MSE reduction against the matched JEPA ranges from 39.7% on PDEBench Burgers and 39.3% on shallow water down to 3.6% on a synthetic PDE and 0.4% on DMC pixel dynamics (reported, Figure 5). The tenth task is CausalWorld closed-loop control, where the mean return is −0.055 against −0.087: both negative, with error bars that overlap almost entirely. The abstract's "improves reported metrics on all 10 dynamics tasks" is true of the means.

Left: bar chart of relative MSE reduction versus standard JEPA for nine prediction tasks: Burgers 39.7 percent, shallow water 39.3, CausalWorld state 20.6, reaction-diffusion 15.2, WeatherBench2 10.5, synthetic pixel dynamics 8.1, DMC-VB held-out episodes 7.5, synthetic PDE 3.6, DMC pixel dynamics 0.4. Right: CausalWorld closed-loop return, standard JEPA minus 0.087 and JEPA-Anything minus 0.055, with wide overlapping error bars.
Nine prediction tasks improve by 0.4% to 39.7%; the control task's two returns overlap within their error bars (paper, Figure 5).

The rollout table shows the errors do not compound faster: on PDEBench Burgers, MSE goes 0.001830 → 0.006369 from step 1 to step 6 for the JEPA and 0.001101 → 0.004014 for JEPA-Anything (reported, Table 6). On the APEBench protocol, Burgers held-out late-state error drops by about 49.5% and six-step rollout by 44.7%, and Kuramoto–Sivashinsky by 13.2%, in every seed (reported). The capacity-matched 50-step Burgers test is the sobering one: the gain is −5.84% at 20 steps and −3.15% at 50, narrowing as the horizon grows.

The Hugging Face folder for this benchmark holds 15 task checkpoints, including CLEVRER, Navier-Stokes and a 2-D reaction-diffusion run that the paper does not report (measured). The paper does not say whether those were evaluated.

Planning and locomotion

This is the one place the paper reports a loss. Paired CEM return favours standard JEPA on Hopper (−1.24), is +0.55 on Walker2d with a confidence interval that crosses zero, and +10.26 on HalfCheetah (reported, Figure 6). The authors say it plainly: planning performance is environment-dependent.

Forest plot of paired CEM-return difference, JEPA-Anything minus standard JEPA, with 95 percent confidence intervals over five seeds: Hopper minus 1.24, Walker2d plus 0.55 with interval crossing zero, HalfCheetah plus 10.26.
Capacity-matched planning: one loss, one tie and one clear win (paper, Figure 6).

The factor ablation is better evidence that the four blocks are each doing something. Replacing any one factor by its training-set mean raises 20-step rollout error in all three environments, and on HalfCheetah masking F1 or F3 lowers return in all five seeds (reported). The single-factor probes find each block linearly predicts a different body variable best on Hopper, with R2R^2 between 0.294 and 0.424: modest, and the paper is careful to call them "motion preferences" rather than labels.

Molecules

With a TrajCast-style equivariant forecaster, JEPA-Anything has the lowest one-step displacement MAE and 100-step RMSD on all four systems (reported, Table 9). Paracetamol 100-step RMSD is 3.155 Å from scratch, 1.868 Å with JEPA pretraining and 1.776 Å with OPF. Water goes 2.536 → 2.459 Å. As elsewhere, the big step is from scratch to any JEPA; OPF is a smaller step on top. The released molecular checkpoint is a 3.1 MB TiTo/PaiNN model with a latent flow-matching module (measured, from its file list and card), not the TrajCast-style setup in the table.

The science claims

Group III is where the paper reaches furthest, and where I can check least.

The cancer case says factor analysis "nominated IL-18 combined with NT5E/CD73 blockade", and wet-lab studies in co-culture, patient-derived organoids, tumour fragments and mice supported it. The paper gives one figure of selected organoid and fragment results and no protocol, no list of other candidates the factors ranked, and no description of the "domain-specific intervention analysis" that turned factor coordinates into a drug pair. I cannot evaluate it from what is published.

The orbital case is checkable in principle. From simulated position-velocity trajectories, spectral analysis of the latent factor coordinates finds frequencies that scale with semimajor axis as f∝a−3/2f \propto a^{-3/2}, with a fitted slope of −1.4991 and R2=0.9999999R^2 = 0.9999999 (reported).

Three panels. Left: frequency-domain ridge of learned spectral modes lying on the Kepler line f equals one over two pi times a to the minus three halves. Top right: recovered power law with fit slope minus 1.4991 and R squared 0.9999999. Bottom right: scale-compensated invariant staying within 0.169 percent of 1. Header: seed 0 selected by the lowest run-level median frequency error.
Learned latent modes on simulated orbits follow Kepler's third law; the header notes the run shown was the best of its seeds (paper, Figure 9).

Two things temper it. The figure's own header says the run shown was "selected by the lowest run-level median frequency error", so it is the best seed, not a typical one. And the simulated orbits obey Kepler's law by construction, so a model that predicts them well must carry each orbit's period somewhere in its state (reasoned). What the result shows is that OPF's factor coordinates expose that period cleanly enough for a spectral fit. That is a useful property of the coordinates. It is not the model discovering a law it was not shown.

What the repository is

The GitHub repo is a library, not a training codebase. jepa-anything-core has 2,069 lines of Python across five modules: the projection and pseudoinverse synthesis, the four loss terms, matched baselines with parameter and FLOP accounting, and geometry audits (measured). A jepa-anything-skill folder holds a design validator and a scaffold generator meant to be driven by a coding agent, and one synthetic recipe that checks interfaces. Its README is direct about it: the example "does not train a model or produce benchmark performance", and the checkpoint manifest records an untrained example with no performance claims. The domain training code is not there. The per-task model code and checkpoints are on the Hugging Face dataset, Apache-2.0, alongside inference scripts.

Two more mismatches between paper and release. The vision checkpoints are a 192-wide ViT on 64×64 Shapes3D images, 13.4 MB, not the DINOv3 ViT-S and SigLIP2 backbones in Table 3 (measured). And the clinical model has no checkpoint at all; its row on the dataset card reads "progressive checkpoint release".

What to take from it

Strip away the "world model everything" framing and JEPA-Anything is a clean, portable change to the predictor of any JEPA: split the target into KK blocks with learned, softly orthogonal projectors; give each block its own small head; keep every coordinate alive with variance floors; stitch the blocks back together with a pseudoinverse so the full state survives for rollouts and planning. It costs almost nothing. The locomotion models differ from their dense baselines by under 0.3% in parameters and 2% in prediction-head FLOPs (reported).

The evidence that it helps is broad and mostly shallow. It beats a matched monolithic JEPA in nearly every comparison, sometimes by 40%, often by a point or less, and once it loses. It is never compared with the strongest specialist model in a field. The geometry that justifies the design is exact in the paper's audit and approximate, sometimes loose, in the shipped weights. If you already run a JEPA for dynamics or rollouts, OPF is a cheap ablation to try. If you read the launch post as one network that forecasts weather and molecules, that is not what the paper built, and the paper says so.

For a different direction on the same family, LeVJEPA removes the EMA target altogether, and Cosmos 3 is what a world model looks like when it is built as one large generative network rather than as a recipe.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "JEPA-Anything: one recipe for seven worlds, not one model", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026jepaanything,
  author = {Satyajit Ghana},
  title  = {JEPA-Anything: one recipe for seven worlds, not one model},
  url    = {https://ai.thesatyajit.com/articles/jepa-anything},
  year   = {2026}
}
share