Does breaking a task into 'big-picture plan then small steps' actually help every task equally, or do some tasks gain way more?
Does hierarchical abstraction help equally across different manipulation tasks?
This explores whether breaking a task into layers of abstraction (high-level plans over low-level actions) pays off the same way regardless of the task, or whether some tasks benefit far more than others. Note that the corpus has almost nothing on physical robot manipulation, so this answer draws on navigation, web agents, and reasoning instead.
This explores whether hierarchical abstraction (planning at a coarse level, then filling in fine-grained steps) helps every task equally, with physical manipulation as the motivating case. One caveat up front: the collection doesn't contain studies that compare hierarchy across different robotic manipulation tasks. What it does have is evidence from neighboring areas, and that evidence points clearly to one answer: no. The benefit is uneven, and it depends on how far a task is from its goal and how well the abstraction fits the task.
The clearest number comes from planning in learned world models. When each level of a hierarchy gets its own, more abstract representation instead of sharing one, success on a long-horizon maze navigation task jumps from 18% to 73% Does each hierarchy level need its own latent space?. The reason matters more than the number. Higher levels help because they can judge candidate futures against goal-like objectives that the low-level detail would hide. That suggests hierarchy pays off most when the goal is far away and the intermediate steps are numerous. For a short, contact-heavy task where the hard part is the fine motor control itself, there's less for the upper level to do.
A similar pattern shows up in software agents. Agents that pull reusable sub-routines out of past experience and stack them into larger workflows gain 24% on one web benchmark and 51% on another, and the gains grow as the test tasks differ more from the training tasks Can agents learn reusable sub-task routines from past experience?. So abstraction does more for generalization than for repeating familiar tasks. A related asymmetry appears in reasoning systems that split a planner from an executor. The ability to break a problem into steps transfers across domains, while the ability to carry out those steps does not Does separating planning from execution improve reasoning accuracy?. GUI agents show the same split: separating planning from grounding (working out exactly where on screen to act) improves results because the two fail in different ways Can structured interfaces help language models control GUIs better?. Translated to manipulation, this suggests the high-level plan ('grasp, lift, place') carries over between tasks, but the low-level skill doesn't. Hierarchy helps exactly as much as a task's difficulty sits in the planning rather than the execution.
There's also a failure mode that's easy to overlook: hierarchy only helps if the task is cut into pieces of the right size. In skill-retrieval systems, standard LLM decomposition recovers only 34% of the correct steps, and simply fixing the step count recovers 75% of the gains What blocks skill retrieval in task decomposition?. So a hierarchy that is a good fit for one task can actively hurt another if the step size is wrong. On the reasoning side, abstractions help mainly by pushing exploration toward broader, more diverse strategies when the compute budget is large Can abstractions guide exploration better than depth alone?. That's another sign the payoff depends on the setting rather than being constant.
One thing you might not expect: neural networks seem to build this kind of modular structure on their own, splitting compositional tasks into separate subnetworks, and pretraining makes that more reliable Do neural networks naturally learn modular compositional structure?. So the practical question for manipulation may be less 'should we add a hierarchy?' and more 'does the task have enough long-horizon, reusable structure for an explicit hierarchy to beat the one the model already learns?' To answer that properly for robots, you'd need sources the collection doesn't currently have.
Sources 7 notes
H-JEPA improves Visual AntMaze planning from 18% to 73% success by giving each hierarchy level a distinct, more abstract latent space rather than sharing one. This per-level abstraction lets higher levels score candidate futures in abstract space better matched to goal-like objectives.
Agent Workflow Memory induces sub-task routines at finer granularity than full tasks, abstracts example-specific values, and compounds them hierarchically. This produces 24.6% relative gain on Mind2Web and 51.1% on WebArena, with larger gains as train-test gaps widen.
Modular architectures with separate decomposer and solver models outperform monolithic LLMs, with decomposition ability transferring across domains while solving ability does not. The separation prevents planning-execution interference and produces more generalizable skills.
Agent S's dual-input design—visual input for environmental understanding plus image-augmented accessibility trees for grounding—achieved 9.37% improvement over baseline by factoring planning and grounding into separate optimization paths rather than forcing end-to-end prediction.
Standard LLM decomposition reaches only 34% step-level recall, gating retrieval success. Correcting step count recovers 75% of gains in iterative methods, shifting the bottleneck to representation-level reranking rather than vocabulary alignment.
Show all 7 sources
RLAD jointly trains abstraction and solution generators, showing that allocating test-time compute to diverse abstractions outperforms parallel solution sampling at large budgets. Abstractions create structured breadth-first exploration that prevents the underthinking failure mode of depth-only reasoning chains.
Pruning experiments reveal that neural networks implement compositional subroutines in isolated subnetworks, with ablations affecting only their corresponding function. Pretraining substantially increases the consistency and reliability of this modular structure across architectures and domains.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Agent S: An Open Agentic Framework that Uses Computers Like a Human
- Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models
- Reasoning LLMs are Wandering Solution Explorers
- Divide-or-Conquer? Which Part Should You Distill Your LLM?
- Distilling LLMs' Decomposition Abilities into Compact Language Models
- Break It Down: Evidence for Structural Compositionality in Neural Networks
- H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning
- Scaling can lead to compositional generalization