SYNTHESIS NOTE
Topics›Tool Computer Use›this note

Can small models match large models on function calling?

Explores whether small language models fine-tuned with the right training method can achieve comparable performance to large models on structured reasoning tasks requiring precise function calls, and what training approach makes this possible.

Synthesis note · 2026-05-03 · sourced from Tool Computer Use

The insight in this paper is methodological: function-calling for reasoning tasks is a domain where DPO outperforms SFT for small models, because the failure modes are more about preferring the right format and call sequence than about generating any plausible text. The proposed framework uses an agent that, given a problem and a callable function set, queries a large LLM by injecting function descriptions and examples and managing calls in a step-by-step reasoning chain. The byproduct is a dataset of correct AND incorrect chat completions — preference pairs ready for DPO.

Why DPO rather than SFT or PPO. SFT teaches the model to imitate good examples but provides no signal about what to avoid — and rigid output formats (precise variable names, JSON, argument values) punish near-misses harshly, so explicit negative examples matter. PPO would work but requires extensive human feedback to train a reward model, making it resource-intensive. DPO removes the reward-model step by incorporating preferences directly into the training objective, with demonstrated stability advantages over PPO.

The structural move is that a large LLM does double duty: it generates the candidate reasoning chains AND its successes/failures provide the preference labels for the small model's training. This is a teacher-distillation pattern but with both polarities — the small model learns what the large model gets right and what it gets wrong, not just to imitate the large model's right answers. The pattern fits the broader case for Can small language models handle most agent tasks?: function-calling is exactly the kind of repetitive, scoped, format-rigid work where a fine-tuned small model can replace a large general-purpose one.

The practical implication: when output format is rigid and small-model deployment is the goal, the question is not "can SFT close the gap" but "what's the cheapest source of preference signal." Self-generated preference pairs from a strong teacher are essentially free relative to human feedback.

Inquiring lines that read this note 121

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does preference optimization undermine conversational grounding in language models? Can iterative DPO substitute for online RL in studying misalignment? How does model capacity affect learning performance on diverse downstream tasks? Can inference-time computation adaptively substitute for static model capacity? What are the fundamental limits of prompting for language models? Can smaller specialized models match frontier models on key metrics? When do simpler collaborative filtering approaches outperform complex LLM recommenders? Should models ask for clarification when facing ambiguous or under-specified information? Why does AI verification capability persistently exceed generation capability? Can base models hide emergent misalignment through alignment training? How susceptible are language models to conversational persuasion and belief change? What prevents language models from performing systematic logical reasoning? How do training data quality and composition affect downstream model performance? How does fine-tuning trade off accuracy against reasoning quality? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? What explains the gap between benchmark scores and true reasoning capability? How does diversity prevent model convergence on superficial patterns? Does training data format shape model reasoning more than domain content? What limits language model accuracy in evaluating ideas? Can minimal training unlock latent reasoning already present in base models? Why do training associations persist despite contradictory contextual information? Do accumulated memories help or hurt continual learning in models? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? Does intelligent routing among smaller models outperform training larger models? How do knowledge graph structures enable efficient multi-hop reasoning and retrieval? Does augmenting symbolic reasoning improve LLM logical reasoning ability? Do AI coding tools measurably improve developer productivity and code quality? How do reward models systematically fail to represent diverse human preferences? When should retrieval systems decide to fetch new information? What causes coordination failures in multi-agent language model systems? How should systems validate code that agents generate? What prediction granularity best trains models to generate reliable reasoning? Does pretraining establish the ceiling for what reward learning can improve? How do curriculum design and feedback approaches affect model learning? Can recurrent computation unlock reasoning capabilities that fixed-depth models cannot? How can persistent memory architectures preserve information across ultra-long contexts? How should agents coordinate through shared persistent code artifacts? When do multi-agent systems improve over single frontier models? What makes agent memory systems durable and reusable across sessions?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 156 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

DPO-trained small models can match large models on function-calling reasoning chains — preference data from a teacher beats SFT for the rigid output format