H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning
Long-horizon planning with latent world models requires reasoning across timescales and levels of abstraction. Existing task-agnostic JEPA world models predict and plan in a single latent space, often at a single timescale. We introduce H-JEPA, an end-to-end recipe for training a hierarchy of action-conditioned JEPAs in which each level predicts farther ahead in its own learned latent space. Planning proceeds top-down: the top level optimizes progress toward the goal, and each level’s predictions become subgoals for the planner below it. When factors in the data evolve at separated timescales, higher levels discard fast, unpredictable detail and retain slower task-relevant state. Across four simulated navigation and manipulation environments, hierarchical planning improves over a flat JEPA; on Visual AntMaze, a three-level hierarchy raises success from 18% to 73% using less planner compute. Ablations attribute these gains to both temporal decomposition and higher-level goal representations. With inverse-dynamics supervision, the approach extends to diverse real-robot videos from DROID, where hierarchy improves offline planning fidelity at lower planner compute.
Introduction. World models learn the dynamics of an environment from experience so that an agent can predict, understand, and plan [29, 30, 31]. Joint-Embedding Predictive Architectures (JEPAs) learn such models by predicting future latent states rather than pixels [49, 1, 9], avoiding both reward and pixel reconstruction [5, 59], and recent action-conditioned, task-agnostic JEPA world models plan zero-shot in the learned latent space [93, 2, 73, 53, 78]. These models, however, predict and plan within a single latent space, often at a single timescale. This has two limitations. First, at a single timescale, long-horizon prediction rolls out many fine-grained steps: prediction error compounds and the action search space grows. Second, a single shared latent space must support both low-level dynamics and goal matching. A latent that carries every fast-varying detail may model low-level dynamics well, but gives a poor goal-matching cost when the goal is more abstract than the state (a location to reach rather than a pose to match; §4.2).
A hierarchical world model divides this work. Each higher level predicts over a longer horizon and keeps in its latent space what remains predictable. Higher levels operate on slow features over long strides and serve long-range prediction and high-level planning; lower levels operate on fast features over short strides and serve short-range prediction and low-level control. The brain has such hierarchical structure: cortical areas form a hierarchy of intrinsic timescales, with slower-varying representations higher up [37, 43, 56, 3]. Hierarchical world models have been studied in video prediction and reward-driven control (§M, Table 13). Hierarchical latent video models reconstruct pixels and are not used for control [68, 44, 55], while hierarchical model-based reinforcement learning learns task-specific policies from reward [32, 36, 27]. Closest to our setting, HWM [92] is a taskagnostic JEPA world model that plans hierarchically, but its predictors over different horizons share one latent space. It decomposes the horizon but cannot match objectives in more abstract spaces.
We introduce H-JEPA, a hierarchical JEPA architecture in which each level predicts the future within its own latent space, at a coarser temporal stride and a higher degree of abstraction than the level below, without reward or reconstruction objectives. Fig. 1 shows the learned hierarchy and its tradeoff between planner compute and success. We make four contributions:
We introduce an end-to-end training method for hierarchical JEPA world models (§2).
We show that when factors in the data evolve at separated timescales, higher levels discard fast detail they cannot predict over their horizons and retain slower, predictable state (§3).
We show that H-JEPA’s hierarchical planning outperforms single-level planning at a fraction of the test-time compute (§4.1). We identify two complementary mechanisms: hierarchical latent spaces let higher-level planners score progress at the goal’s own level of abstraction, while temporal decomposition breaks long-horizon tasks into easier subproblems (§4.2).
With an inverse-dynamics term, the method extends to DROID, a real-robot manipulation corpus whose scene, lighting and objects change every episode (§4.3).
Related work. Action-conditioned JEPA world models plan in learned latent spaces without pixel reconstruction or reward supervision. DINO-WM [93] learns dynamics over frozen pretrained features, V-JEPA 2 [2] adds action-conditioned prediction after video pretraining, and PLDM [73] and LeWM [53] train the encoder and predictor jointly from pixels. H-JEPA extends this end-to-end setting to a hierarchy of latent spaces and prediction timescales, using LeJEPA’s SIGReg regularizer [6] at every level.
Hierarchical video models learn representations at several timescales but reconstruct observations and do not plan [68, 44, 55], while hierarchical model-based control methods are reward-driven and task-specific [32, 36, 27]. Closest to our setting, HWM [92] trains latent world models at multiple temporal horizons from reward-free offline data and plans hierarchically, but within a shared latent space. H-JEPA additionally learns a distinct representation at each level, so that higher-level objectives can discard detail that lower-level dynamics retain. §M gives the extended survey and Table 13 compares hierarchical control methods.
Method. H-JEPA learns a hierarchy of latent predictive models from trajectories of observations and actions. The hierarchy is temporal: level 1 operates on the finest observation stream, while each higher level consumes the latent states produced by the level below over a coarser stride. This gives a stack of JEPA world models in which each level has its own observation encoder, action encoder, and latent predictor. All levels predict future latent states from past latent states and actions.
Let ot denote an observation and at the action block for the transition from ot to ot+1. The first level encodes the observation, optionally with the proprioceptive state, into a latent state z(1) t = E(1)(ot), and the action block into a(1) t = A(1)(at). Higher levels are built compositionally: both state and action encoders pool temporal windows of latents from the level below.
Each level l> 1 has two temporal hyperparameters. The stride slis the subsampling factor: upperlevel time t maps to lower-level time t · sl, so consecutive upper-level states are sllower-level steps apart. The window size wlis the number of lower-level steps each upper-level state summarizes. The two are independent: slsets how far each window advances and wlhow far it spans.
The state encoder E(l) pools the wl-step lower-level window into one abstract state. The action encoder A(l) instead aggregates sllower-level action embeddings, independent of wl, so a(l) t covers the full transition from the state anchoring z(l) t toward the next upper-level state z(l) t+1. All experiments use wl=1, so upper-level state encoders are pointwise (Fig. 2).
Each level is trained as a JEPA in its own latent space. Given a context of cllatent states (z(l) t−cl+1, . . . , z(l) t ) and the associated action embeddings, a predictor F (l) predicts future latent states. One-step training uses the teacher-forced latent prediction loss The per-level loss also includes SIGReg [6], a sketched normality regularizer that prevents collapse by encouraging an isotropic Gaussian embedding distribution (§K). Let Z(l) ∈Rn×D collect n level-l embeddings of width D, flattened over batch and time. The overall level-lobjective is H-JEPA plans top-down through the learned hierarchy (Fig. 3). The current and goal observations are encoded at every level l, yielding initial latent states {z(l) 0 }L l=1 and goal latent states {g(l)}L l=1, where L is the number of levels. The top level proposes a coarse plan to the goal. Each lower level plans toward subgoals supplied by the predicted trajectory of the level above, refining the plan into finer-scale transitions until level 1 produces primitive actions. The top-level planner optimizes macro-actions to minimize the distance between its final predicted state and the goal:
At each level l, the predicted states (ˆz(l) 1 , . . . , ˆz(l) Hl) = Roll(l)(z(l) 0 , a(l) 0:Hl−1) are obtained by autore- For each lower level l< L, the optimized rollout from the level above, (ˆz(l+1),∗ 1 , . . . , ˆz(l+1),∗ Hl+1 ), supplies subgoals. The lower level optimizes its actions to match these subgoals with its own predicted states, encoded into the upper-level latent space by E(l+1).
We first choose the number of subgoals to match, 1 ≤Kl+1 ≤Hl+1, then set the lower-level horizon to cover the corresponding encoder windows. With upper-level stride sl+1 and window size wl+1, this requires Hl= Kl+1sl+1 + wl+1 −1. The encoded predictions are Each ̃z(l+1) i is aligned with its upper-level subgoal in time and latent space. The lower level solves We optimize action sequences at every level by gradient descent. Level 1 returns primitive actions for execution in the environment. In closed-loop control, we execute an action prefix, encode the new observation, and repeat hierarchical planning until the evaluation budget of environment steps is exhausted. Execution schedules and solver settings are in §E.
Discussion. Hierarchical planning changes two things at once: it decomposes the search in time, and it scores candidate futures in a more abstract latent space. We isolate the second factor, asking how much an objective in a higher, more abstract level helps on its own, independent of temporal decomposition. We focus on Visual AntMaze, where H-JEPA’s substantial gains over HWM suggest benefits from hierarchical representations beyond temporal decomposition alone (Fig. 6, bottom). planning: the upper-level encoders define the cost, but no upper-level planner generates subgoals. Learning a hierarchy of representations can thus help planning simply by providing a more abstract space in which to measure distance to the goal, before any temporal decomposition. We extend this ablation to FourRoom, Push-T and Cube in §H (Table 9). On FourRoom, abstract goal costs improve flat planning even though H-JEPA and HWM perform similarly. On Push-T and Cube, where we observe no selective abstraction, projected costs provide no clear benefit.
Temporal decomposition. Independently of shaping the planning objective through higher-level representations, hierarchy helps by decomposing a long-horizon task into subgoal problems. When each lower-level planner tracks only a prefix of the upper-level plan (Kl+1 < Hl+1), these subproblems require fewer sequential rollout steps per candidate and limit the action search to shorter horizons. This reduces compute, consistent with the improved success–compute tradeoffs in Fig. 6 (top). Second, a flat planner’s full-horizon goal cost can stall or increase even along expert trajectories, providing an inconsistent measure of progress toward the goal (Fig. 8). Across all four simulated environments, the level-1 planner’s costs for tracking level-2 subgoals decrease more consistently than its full-horizon goal costs: for H-JEPA, monotonicity (negative Spearman correlation between cost and time) ranges from 0.89–1.00 for subgoal tracking, compared with 0.48–0.75 for the fullhorizon goal cost (Table 10 in §I). Both HWM and H-JEPA exhibit this improvement, consistent with temporal decomposition contributing to HWM’s planning gains over LeWM even without distinct abstract representations (Fig. 6, bottom).
Conclusion. We presented H-JEPA, an end-to-end method for training a hierarchy of action-conditioned JEPAs, each predicting in its own latent space over progressively longer horizons. Higher levels discard fast features that are unpredictable over their horizon and retain slow, predictable ones. Hierarchical planning with H-JEPA outperforms flat planning at lower planner compute by scoring goals in a more abstract latent space and decomposing long tasks into shorter subgoal problems. With inverse-dynamics supervision, H-JEPA also improves offline, open-loop planning fidelity on diverse real-robot videos from DROID. A next step is to test whether these gains translate to closed-loop control on a physical robot.
Broader directions include learning hierarchies from action-free video and other high-dimensional temporal streams; aligning upper-level latents with language so that goals can be specified in words rather than as target observations; and training upper levels on variable-duration segments instead of fixed strides, which may yield better representations or subgoals for hierarchical planning.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How does decomposing tasks into separate stages affect reasoning quality and safety?- How does temporal decomposition into subgoals improve long-horizon planning stability?
- Does hierarchical abstraction help equally across different manipulation tasks?
- How does planning-before-execution compare to iterative reasoning and action loops?
- What physical structure does a Gaussian-regularized latent space actually encode?
- What makes regularization an implicit factor in embedding geometry?
- What prevents representation collapse in latent-prediction world models like JEPA?
- Can generative reconstruction preserve latent manifold structure better than geometric compression?
- How can diffusion models predict future tokens without completing prior blocks?
- Can architecture changes and early stopping combine to close the diffusion inference gap?