Could an AI agent figure out which skill to use just by trying things, instead of trusting whichever instructions sound closest to the task?
Can empirical testing replace semantic retrieval for finding useful skills?
This explores whether an agent could find the right skill (a reusable procedure or instruction document) by trying candidates and checking what actually improves results, instead of picking whichever skill's description sounds closest to the task.
This explores whether an agent could find useful skills by testing them and keeping what works, rather than choosing whichever skill description sounds most similar to the task. The collection has no head-to-head study of the two approaches. It does make a strong case that semantic similarity is the wrong measure of usefulness, and it shows empirical testing working well one step over: deciding how a skill should be written rather than which skill to pick.
Start with why similarity falls short. One recurring diagnosis in retrieval research is that embeddings measure association, not relevance. Two texts can sit close together in embedding space without one being useful for the other's task, and there are mathematical limits on what a fixed-size embedding can represent at all Where do retrieval systems fail and why?. The memory literature gives a sharper warning. Memories that were recorded accurately and are genuinely relevant to the task still made LLM reasoning worse, and every memory framework tested scored below a no-memory baseline Can relevant memories actually harm LLM reasoning?. If relevant context can hurt, then "this skill looks relevant" doesn't tell you "this skill will help." Only trying it does.
The clearest case for testing is SkillOpt, which treats a skill document like model weights. An optimizer proposes edits to the text, and each edit is accepted only if it improves performance on held-out validation tasks Can skill documents be optimized like neural network weights?. That is empirical testing used as a filter on skill content, and it adds no cost when the agent runs. A similar logic shows up in training: tree search can stand in for human judgment by simply recording which solution paths succeed Can tree search replace human feedback in LLM training?. Across these examples, outcome signals keep replacing judgments about what looks right.
The less obvious finding is that retrieval may not be the main bottleneck. In compositional skill retrieval, the bigger problem is how the task gets broken into steps. Standard LLM decomposition matches only about a third of the steps it should, and fixing the step count recovers most of the lost performance What blocks skill retrieval in task decomposition?. If the agent splits a task into the wrong pieces, neither similarity search nor testing will find the right skill, because it is searching for the wrong things. Other work suggests the fix may be to change how you search rather than drop search entirely: agents that grep raw text beat embedding-based retrieval on precise queries Can direct corpus search beat embedding-based retrieval?, and routing a query to the right kind of knowledge structure beats one uniform retrieval method Can routing queries to task-matched structures improve RAG reasoning?.
The likely answer, then, is that testing doesn't replace retrieval but checks it. Retrieval proposes a few candidates cheaply, and testing confirms which ones actually help, the way SkillOpt's validation step gates every edit. Testing every skill in a large library would be expensive, and the collection doesn't yet show anyone doing it for selection. The part you might not expect is that the biggest gain may come before either step, from breaking the task into the right pieces.
Sources 7 notes
RAG systems fail at three structural levels: adaptive triggering (fixed intervals waste context), semantic-task mismatch (embeddings measure association, not relevance), and mathematical limits (embedding dimension constrains representable document sets). These require fundamentally different retrieval approaches, not tuning.
MemTrapBench shows that all five tested memory frameworks underperform a no-memory baseline, with drops exceeding 10%, despite memories being accurately stored and task-relevant. This reveals a failure mode at the point of use that standard memory benchmarks miss.
SkillOpt treats skill documents as trainable external state of frozen agents, using a text-space optimizer with held-out validation gating to accept only edits that improve performance. Across 52 benchmark cells and seven models, the approach matches or exceeds baselines while adding zero inference cost and enabling transfer across models.
AlphaLLM uses tree search outcomes and three critic models to derive dense reward signals equivalent to human-labeled feedback. Tree structure naturally ranks solution paths by success, replacing the annotation oracle that standard RLHF requires.
Standard LLM decomposition reaches only 34% step-level recall, gating retrieval success. Correcting step count recovers 75% of gains in iterative methods, shifting the bottleneck to representation-level reranking rather than vocabulary alignment.
Show all 7 sources
GrepSeek trains agents to retrieve via executable shell commands over raw text, achieving better multi-hop performance on entity-constrained queries than dense embeddings. The approach scaffolds unstable search mechanics with supervised trajectories, then refines task-oriented behavior through reinforcement learning.
StructRAG demonstrates that selecting knowledge structure type based on query demands—via DPO-trained router choosing among tables, graphs, algorithms, catalogues, and chunks—improves knowledge-intensive reasoning over standard retrieval. The approach grounds this in cognitive load and cognitive fit theory from cognitive science.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Chain-of-Retrieval Augmented Generation
- You Don't Need Pre-built Graphs for RAG: Retrieval Augmented Generation with Adaptive Reasoning Structures
- Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs
- Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose
- UR2: Unify RAG and Reasoning through Reinforcement Learning
- BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms
- On the Theoretical Limitations of Embedding-Based Retrieval
- GrepSeek: Training Search Agents for Direct Corpus Interaction