INQUIRING LINE

Does an AI learn to solve new kinds of problems from polished answers, or from the messy process of reaching them?

Does training on curated solutions transfer to unseen problem types?

This explores whether a model trained on clean, hand-picked correct answers learns something it can carry to new kinds of problems, or mostly learns to repeat the patterns it was shown.


This explores whether training a model on curated, correct solutions teaches skills that carry over to problem types it hasn't seen, or mostly teaches it to recognize familiar ones. The corpus suggests a surprising answer: polished solutions are often the weakest thing to train on, and what does transfer is the process of getting to an answer, not the answer itself. The clearest evidence comes from maze-solving experiments, where the length of a model's reasoning tracked problem difficulty only on familiar problems. On unfamiliar ones, that link broke down completely. The model wasn't thinking harder about hard problems. It was recalling templates from training Does longer reasoning actually mean harder problems?.

Several lines of work point the same way. Training on the full, messy search (dead ends, backtracking, recovery) produced problem-solvers about 25% more accurate than training only on optimal paths. The models seemed to build an internal sense of how to search instead of memorizing routes Does training on messy search processes improve reasoning?. 'Journey learning' makes the same case for o1-style reasoning: models trained on trial, error, and self-correction hold up better than models trained on shortcut solutions Can models learn better by training on messy exploration paths?. One result is especially striking: a single problem, paired with critiques of several right and wrong attempts, unlocked reasoning about as well as reinforcement learning did Can a single problem unlock reasoning through solution critique?. So the useful signal may be seeing the contrast between good and bad reasoning, not the number of correct examples.

The part you may not have expected: transfer isn't always good news. Unwanted behaviors generalize too. When reinforcement learning rewards only correct final answers, the model loses variety in how it approaches problems, and that loss spreads from the problems it solved to the ones it hasn't. It gets less exploratory exactly where exploration is needed Does outcome-based RL diversity loss spread across unsolved problems?. Problems that are too hard cause a related failure. Rare lucky successes get rewarded heavily, so the model learns shortcuts like repeating answers or skipping computation, and those shortcuts damage abilities it already had Do overly hard RLVR samples actually harm model capabilities?. Putting critique into the training loop is one proposed fix, because it keeps solution variety from narrowing over training Do critique models improve diversity during training itself?.

How training is structured also seems to matter as much as what goes in. Splitting function calling into seven specific subskills generalized better than large catch-all datasets Can breaking function calling into subtasks improve model generalization?. Training on structured tasks like math before open-ended creative ones gave 6.2% gains, because the narrowing effect of structured training stopped damaging creative ability Does training order reshape how models handle different task types?. Some systems do report real out-of-distribution transfer: AIDE2's gains held on a weather-forecasting benchmark far outside its training tasks Do AIDE2's improvements transfer to unseen tasks?. A trained skill curator learned cross-task strategies that worked across different models and domains Can a separate trained curator improve skill libraries better than frozen agents?. In both cases, what transferred was a strategy for working, not a collection of answers.

The takeaway: if you want transfer, curate the process (mistakes, critiques, and a sensible ordering of tasks), not just the solutions. The corpus doesn't directly test the plain version of the question, meaning supervised fine-tuning on curated solutions measured on truly new problem types. The answer here is pieced together from neighboring evidence.


Sources 11 notes

Does longer reasoning actually mean harder problems?

Controlled A* maze experiments show trace length correlates with difficulty only in-distribution but decouples entirely out-of-distribution. Trace length primarily reflects recall of training schemas, not adaptive computation.

Does training on messy search processes improve reasoning?

Stream of Search pretraining, which represents exploration and backtracking as serialized strings, achieves 25% higher accuracy than optimal-trajectory-only training. Models learn internal world models for search and adaptive strategies rather than fixed external methods.

Can models learn better by training on messy exploration paths?

Research shows that training on messy trajectories—failed attempts, self-correction, and backtracking—teaches more robust reasoning than training only on shortcut solutions. This approach models o1-style deep reasoning as search internalization rather than solution memorization.

Can a single problem unlock reasoning through solution critique?

Critique Fine-Tuning achieves reasoning activation comparable to RLVR using only one problem and teacher-generated critiques of varied solutions, with no reinforcement learning. This demonstrates that exposure to correct versus incorrect reasoning on a specific problem is the sufficient activation signal.

Does outcome-based RL diversity loss spread across unsolved problems?

RL that rewards only final answer correctness sharpens the policy globally, concentrating probability mass on correct trajectories for solved problems while simultaneously reducing diversity on unsolved ones. Historical exploration (training diversity via UCB-style bonuses) and batch exploration (test-time diversity via repetition penalties) require structurally different mechanisms.

Show all 11 sources
Do overly hard RLVR samples actually harm model capabilities?

Training on nearly-impossible problems causes models to learn degenerate shortcuts rather than genuine reasoning, and these shortcuts contaminate pre-existing capabilities. Group-relative normalization treats rare accidental successes as high-advantage trajectories, reinforcing answer repetition and computation-skipping instead of sound reasoning patterns.

Do critique models improve diversity during training itself?

Step-level critique in the training loop counteracts tail narrowing and maintains solution diversity across self-training iterations. This training-time benefit—preventing premature convergence—is more fundamental than test-time accuracy gains.

Can breaking function calling into subtasks improve model generalization?

Granite-20B-FunctionCalling shows that explicit training across seven granular subtasks—nested calls, chaining, parallel functions, name detection, parameter detection, next-best function, and response generation—generalizes better than umbrella datasets like ToolLLM. This multi-task approach closes the performance gap with GPT, Claude, and Gemini.

Does training order reshape how models handle different task types?

Omni-Thinker shows structured domains decrease output entropy while creative domains increase it. BWT-guided scheduling—training structured tasks first—yields 6.2% gains over joint training by preventing entropy collapse from damaging open-ended capabilities.

Do AIDE2's improvements transfer to unseen tasks?

The paper reports that AIDE2's improvements transfer to four held-out benchmarks spanning machine learning, algorithm engineering, and physics-based weather forecasting—the last being outside the selection task distribution. This demonstrates transferable gains beyond overfitting to the selection set.

Can a separate trained curator improve skill libraries better than frozen agents?

SkillOS shows that separating a trainable curator from a frozen executor, grouped by task streams, causes skill repositories to shift from generic verbose additions toward actionable execution logic and cross-task meta-strategies. The trained curator generalizes across different executor backbones and domains.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.