Does each hierarchy level need its own latent space?
Explores whether separate learned representations at different hierarchy levels improve planning performance compared to sharing a single latent space across all levels.
H-JEPA trains a hierarchy of action-conditioned JEPA world models in which "each level predicts in its own learned latent space," at "a coarser temporal stride and a higher degree of abstraction than the level below." Planning runs top-down: the top level sets a coarse plan toward the goal, and each level's predicted trajectory supplies subgoals for the level below, down to level 1, which "returns primitive actions for execution." Across four simulated navigation and manipulation environments, "hierarchical planning improves over a flat JEPA; on Visual AntMaze, a three-level hierarchy raises success from 18% to 73% using less planner compute." With an added inverse-dynamics term the method extends to DROID, a real-robot manipulation video corpus, improving offline planning fidelity at lower planner compute.
The paper attributes the gain to two mechanisms it isolates by ablation. First, a distinct, more abstract representation at each level lets a higher-level planner "score candidate futures in a more abstract latent space" — useful because a goal is often "more abstract than the state (a location to reach rather than a pose to match)," so a flat latent carrying every fast-varying detail gives a poor goal-matching cost. Second, independent of representation, temporal decomposition breaks a long-horizon task into shorter subgoal problems: tracking a nearby subgoal yields a more consistently improving cost than tracking the full-horizon goal (Spearman monotonicity 0.89–1.00 for subgoal tracking versus 0.48–0.75 for full-horizon cost). The paper credits HWM's prior planning gains over LeWM to this second mechanism alone, since HWM shares one latent space across its prediction horizons; H-JEPA's further gain on Visual AntMaze is attributed to the first mechanism — distinct per-level abstraction — which HWM lacks.
The architecture directly inherits its anti-collapse method from Can a single regularizer prevent JEPA representation collapse?: H-JEPA trains every level "using LeJEPA's SIGReg regularizer," the same Gaussian-distribution constraint LeWM used to stabilize a single-level JEPA, now applied per level across a hierarchy. It also sits in contrast with Can looped computation replace parameter count in world models?: LoopWM scales world-model computation by looping one shared block to adaptive depth within a single latent space, while H-JEPA scales by stacking distinct latent spaces, each at its own timescale and abstraction level — two different answers to where a world model should spend extra computation, temporal depth within one representation versus representational depth across several.
The excerpt reports gains only on simulated navigation/manipulation benchmarks and on offline, open-loop planning fidelity on real-robot video; the paper itself flags closed-loop control on a physical robot as untested ("a next step is to test whether these gains translate to closed-loop control on a physical robot"). It also does not establish that abstraction helps uniformly: on Push-T and Cube "we observe no selective abstraction" and "projected costs provide no clear benefit," so the abstraction mechanism's contribution looks environment-dependent — strongest where state and goal differ in abstraction level (AntMaze, FourRoom), negligible otherwise.
Inquiring lines that read this note 6
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How does decomposing tasks into separate stages affect reasoning quality and safety? How do AI systems determine and balance multiple competing objectives? Can recurrent computation unlock reasoning capabilities that fixed-depth models cannot? Can AI systems achieve real improvement without external human feedback? How do thinking tokens exhibit diminishing returns in reasoning?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can a single regularizer prevent JEPA representation collapse?
JEPAs traditionally need complex loss stacks and auxiliary tricks to avoid collapse. Can a single Gaussian-distribution constraint on latent embeddings do the same stabilization work, and would that simplify training?
H-JEPA reuses LeWM's SIGReg Gaussian-latent regularizer at every level of its hierarchy
-
Can looped computation replace parameter count in world models?
Does iteratively refining latent states through a shared transformer block achieve comparable performance to larger models while adapting computation depth per prediction step? This matters because world models struggle with long-horizon rollout error and computational cost.
contrasting scaling axis: looped depth within one latent space versus stacked distinct latent spaces
-
Can trajectory structure replace hand-annotated process rewards?
Recent methods extract step-level supervision directly from how agent trajectories are structured—trees, expert alignments, tool calls—rather than training separate reward models. Can this structural approach consistently avoid annotation costs?
parallel move of deriving dense, decomposed signal from trajectory structure rather than a single outcome measure
-
Why does latent chain-of-thought fail so easily in training?
Explores why latent reasoning is fragile compared to textual chain-of-thought, focusing on how outcome-only supervision creates gradient starvation and representational drift in learned reasoning trajectories.
evidence for — B's account of latent-space drift under shared representations explains why H-JEPA's separate per-level latents avoid collapse
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning
- Agent S: An Open Agentic Framework that Uses Computers Like a Human
- Emergent Hierarchical Reasoning In LLMs Through Reinforcement Learning
- Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
- Latent Collaboration in Multi-Agent Systems
- LLM+P: Empowering Large Language Models with Optimal Planning Proficiency
- NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction
- Base Models Know How to Reason, Thinking Models Learn When
Original note title
giving each hierarchy level its own latent space, not a shared one, drives H-JEPAs planning gains over flat and shared-latent world models