INQUIRING LINE

Why might letting an AI explore, fail, and backtrack teach it more than handing it a tidy, pre-sorted set of correct examples?

Why does open-ended search outperform direct curriculum building approaches?

This explores why letting a model learn by exploring (trying, failing, backtracking) can beat handing it a carefully built sequence of training examples. The corpus has no head-to-head study proving that open-ended search always wins, but it does explain when and why exploration-based approaches beat tidy, pre-built learning paths.


This explores why learning by exploring, with its wrong turns and backtracking, can beat learning from a clean, pre-built path of examples. One caveat first: the corpus has no paper that directly pits open-ended search against curriculum design and declares a winner. What it does have is several findings that point to the same reason why exploration often comes out ahead: a model learns more from seeing how problems get solved, mistakes included, than from seeing only the right answers.

The clearest evidence comes from Stream of Search. When models were trained on the full, messy record of a search (dead ends, backtracking and recovery), they solved 25% more problems than models trained only on the optimal solution path Does training on messy search processes improve reasoning?. A curated path of 'correct moves' hides the very thing the model most needs: a sense of how to tell it's lost and how to recover. Without that, today's reasoning models act as 'wandering explorers'. They retry and drift rather than search systematically, so their success rate falls off exponentially as problems get deeper Why do reasoning LLMs fail at deeper problem solving?. Search also produces its own grading signal. In AlphaLLM, tree search ranks solution paths by whether they succeed, and that ranking replaces the step-by-step human labels that a hand-built curriculum would need Can tree search replace human feedback in LLM training?.

The surprising part is that the curricula that work best behave a lot like search. Reverse curriculum learning doesn't decide in advance what the model should study. It starts the model near the end of a solution and moves the starting point backward step by step. Wherever the model starts failing, that's where it needs work, and this gets close to the value of step-level supervision using only right/wrong feedback on final answers Can curriculum learning approximate expensive process supervision?. The flip side shows up in teacher-refined data. 'Better' examples written by a stronger model can actually hurt a student when they go beyond what the student is ready to learn from Does teacher-refined data always improve student model performance?. That's the main weakness of building a curriculum directly: the designer has to guess where the learner's limits are, while search finds them by running into them.

There's also a question of what gets learned. Reasoning ability seems to come from broad, reusable know-how about how to do things, picked up from many different sources, rather than from memorizing specific examples Does procedural knowledge drive reasoning more than factual retrieval?. Exploration produces exactly that kind of know-how. Training data that mixes different strategies and includes critique-and-retry exchanges, built as branching trees of attempts instead of single answers, has produced strong open-model reasoners What alignment data structure best trains reasoning generalists?. There's a limit, though: search inside a fixed problem space isn't creativity. Work on creative reasoning argues that combining ideas, exploring a space and reshaping the space itself are separate abilities that current methods mostly ignore Can LLMs reason creatively beyond conventional problem-solving?.

The takeaway you may not have expected: curriculum versus search is less a contest than a question of who finds the learner's limits. A designer who guesses ahead of time can guess wrong. The model's own failures during search, or a curriculum that adjusts based on those failures, tell you where the limits really are.


Sources 8 notes

Does training on messy search processes improve reasoning?

Stream of Search pretraining, which represents exploration and backtracking as serialized strings, achieves 25% higher accuracy than optimal-trajectory-only training. Models learn internal world models for search and adaptive strategies rather than fixed external methods.

Why do reasoning LLMs fail at deeper problem solving?

Current reasoning models lack the three properties of systematic exploration: validity, effectiveness, and necessity. This causes success probability to drop exponentially with problem depth, making medium problems solvable but deep problems catastrophically harder.

Can tree search replace human feedback in LLM training?

AlphaLLM uses tree search outcomes and three critic models to derive dense reward signals equivalent to human-labeled feedback. Tree structure naturally ranks solution paths by success, replacing the annotation oracle that standard RLHF requires.

Can curriculum learning approximate expensive process supervision?

R3 progressively slides the reasoning start state backward from near-completion, creating a curriculum that reveals step-level failure modes using only outcome feedback. This achieves process supervision granularity without expensive human step annotations.

Does teacher-refined data always improve student model performance?

Teacher-refined data degrades performance when it exceeds the student's learning frontier, even if objectively higher quality. Students should filter refinements using their own statistical profile to retain only compatible improvements.

Show all 8 sources
Does procedural knowledge drive reasoning more than factual retrieval?

Analysis of 5 million pretraining documents shows reasoning relies on broad, transferable procedural knowledge from diverse sources, unlike factual recall which depends on narrow, document-specific memorization of target facts.

What alignment data structure best trains reasoning generalists?

Eurus achieved state-of-the-art open-model reasoning by training on ULTRAINTERACT, an alignment dataset structured as preference trees per instruction. The tree format unified diverse planning strategies, interaction-and-critique trajectories, and pairwise data for both SFT and preference learning.

Can LLMs reason creatively beyond conventional problem-solving?

Research identifies combinational, exploratory, and transformational reasoning as distinct creative modes grounded in cognitive science. Existing LLM reasoning methods address only conventional problem-solving, leaving creative paradigms unaddressed and potentially explaining diversity collapse in ideation.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.