SYNTHESIS NOTE
Topics›Routers›this note

Can routing beat building one better model?

Does directing queries to specialized models via semantic clustering outperform investing in a single frontier model? This challenges whether model improvement or model selection drives performance gains.

Synthesis note · 2026-02-23 · sourced from Routers

Avengers-Pro demonstrates that routing queries to different models based on semantic clustering can exceed the performance of any individual model in the pool — including frontier models. The mechanism: embed incoming queries, cluster by semantic similarity, evaluate per-cluster model performance-efficiency scores, and route each query to the highest-scoring model for its cluster.

Three results establish the claim:

The earlier Avengers work made an even more striking claim: ten models of ~7B parameters each, with routing, surpassed GPT-4.1 and 4.5 across 15 datasets. This suggests the performance gain from optimal model selection can be comparable to the gap between model generations.

The architecture is lightweight: three operations at inference time (embedding, nearest-cluster lookup, score aggregation). The heavy work — fitting the clustering model and estimating per-cluster performance statistics — happens offline on a calibration set (70% for fitting, 30% for evaluation). This makes the approach deployable as a thin routing layer atop any model API ecosystem.

Since Can we allocate inference compute based on prompt difficulty?, Avengers-Pro adds a complementary optimization axis. Compute-optimal scaling asks "how much inference budget per query?" Routing asks "which model per query?" These are independent — a routing layer could be composed with per-query compute allocation for a two-dimensional Pareto optimization. Since Can inference compute replace scaling up model size?, routing extends this: you don't need a bigger model OR more compute — you need the right model for this specific query type.

The implication challenges the frontier model race: rather than building one model that dominates on everything, assembling a diverse pool of specialized-ish models with good routing may be both cheaper and more effective. This aligns with the heterogeneous architecture thesis in Can small language models handle most agent tasks? — routing makes the heterogeneous approach practical.

Inquiring lines that read this note 82

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

When do simpler collaborative filtering approaches outperform complex LLM recommenders? Does intelligent routing among smaller models outperform training larger models? Can recurrent computation unlock reasoning capabilities that fixed-depth models cannot? Why do LLM research ideation systems generate novelty but lack diversity? Can AI systems discover fundamental improvements to their own architectures? What capabilities differentiate diffusion from autoregressive language models? Can smaller specialized models match frontier models on key metrics? When do multi-agent systems improve over single frontier models? How does model capacity affect learning performance on diverse downstream tasks? What explains the gap between benchmark scores and true reasoning capability? How should retrieval strategies adapt to multi-step reasoning demands? Why does self-revision amplify confidence in wrong model answers? How does diversity prevent model convergence on superficial patterns? How should recommendation systems balance individual preference and diversity? How do reward models systematically fail to represent diverse human preferences? Do single-axis benchmarks accurately measure agent capability for real deployment? How do thinking tokens exhibit diminishing returns in reasoning? How can persistent memory architectures preserve information across ultra-long contexts? How does fine-tuning trade off accuracy against reasoning quality? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? Why do abstract preferences outperform episodic memories in personalization? How does decomposing tasks into separate stages affect reasoning quality and safety? How do multi-agent systems fail when coordination breaks down? Can code harness improvements rival direct model scaling for capability? How do individually-safe actions create collectively-unsafe outcomes? Can AI research automation sustain progress through accelerating feedback loops?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 140 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

test-time model ensembling via embedding-cluster routing surpasses any individual frontier model — model selection is a stronger lever than model improvement