← All clusters

Knowledge Retrieval and Reasoning

Research on how AI systems acquire, store, and retrieve structured and unstructured knowledge to answer questions and generate informed responses. Covers RAG architectures, knowledge graphs, domain specialization, and the integration of external information into language model reasoning.

76 notes (primary) · 414 papers · 6 sub-topics
View as

Retrieval-Augmented Generation (RAG)

22 notes

Does lexical search outperform agentic navigation as corpus size grows?

As document collections scale, does a simple indexed search like BM25 become more reliable than an intelligent agent that explores the corpus sequentially? This matters because it shapes how to architect production RAG systems.

Explore related Read →

When should retrieval happen during model generation?

Explores whether retrieval should occur continuously, at fixed intervals, or only when the model signals uncertainty. Standard RAG retrieves once; long-form generation requires dynamic triggering based on confidence signals.

Explore related Read →

Can question features alone predict when to retrieve?

Can lightweight external features of a question—rather than expensive model uncertainty checks—reliably decide whether retrieval is needed? This matters because uncertainty-based methods promise efficiency but add computation.

Explore related Read →

Can you adapt retrieval models without accessing target data?

Explores whether dense retrieval systems can adapt to new domains using only a textual description, rather than actual target documents—especially relevant for privacy-restricted or competitive scenarios.

Explore related Read →

What do enterprise RAG systems need beyond accuracy?

Academic RAG benchmarks focus on question-answering accuracy, but enterprise deployments in regulated industries face five distinct requirements—compliance, security, scalability, integration, and domain expertise—that standard architectures don't address.

Explore related Read →

Can fine-tuning replace query augmentation for retrieval?

Query augmentation helps retrievers handle ambiguous queries but increases input cost. Does fine-tuning the retrieval model achieve comparable performance without this overhead?

Explore related Read →

Can long-context models resolve retriever-reader imbalance?

Traditional RAG systems force retrievers to find precise passages because readers had small context windows. Do modern long-context LLMs change what architecture makes sense?

Explore related Read →

Can query-time graph construction replace pre-built knowledge graphs?

Does building dependency graphs from individual queries at inference time offer a more flexible and cost-effective alternative to constructing knowledge graphs over entire document collections upfront?

Explore related Read →

Can retrieval learn what actually helps answer questions?

Standard RAG trains retrievers to find similar documents and generators to produce answers separately. But does surface similarity match what genuinely helps generate correct responses? This explores whether retrieval can receive feedback from answer quality.

Explore related Read →

Can knowledge graphs enable multi-hop reasoning in one retrieval step?

Standard RAG retrieves once but misses chains; iterative RAG follows chains but costs more. Can we encode multi-hop paths in a knowledge graph so one retrieval pass discovers them all?

Explore related Read →

Can long-context LLMs replace retrieval-augmented generation systems?

Explores whether loading entire corpora into LLM context windows can eliminate the need for separate retrieval systems, and what task types this approach handles well or poorly.

Explore related Read →

Can a model's partial response guide what to retrieve next?

Does using the model's in-progress output as a retrieval signal reveal information needs better than the original query alone? This explores whether generation itself can diagnose what documents are missing.

Explore related Read →

Does question type determine the right retrieval strategy?

Explores whether different non-factoid question types require distinct retrieval and decomposition approaches. Matters because standard RAG fails when applied uniformly to debate, comparison, and experience questions despite being effective for factoid queries.

Explore related Read →

Why do queries and documents occupy different embedding spaces?

Queries and documents express the same information in fundamentally different ways—short and interrogative versus long and declarative. Understanding this mismatch is crucial for why direct embedding retrieval often fails.

Explore related Read →

Can rationale-driven selection beat similarity re-ranking for evidence?

Can LLMs generate search guidance that outperforms traditional similarity-based evidence ranking? This matters because current re-ranking lacks interpretability and fails against adversarial attacks.

Explore related Read →

Can relevance guide grep search beyond just selecting documents?

Most retrieval systems use relevance only to pick top-k documents for the model. Could relevance instead shape how an agent explores the corpus—ordering traversal, setting entry points, and filtering matches—to find evidence faster?

Explore related Read →

Does synthetic content in search results hide ecosystem decay?

As AI-generated content dominates search rankings, do traditional accuracy metrics mask a silent loss of source diversity and ecosystem health? This matters because hidden fragility could make systems vulnerable to future corruption.

Explore related Read →

Can document count be learned instead of fixed in RAG?

Standard RAG systems use a fixed number of documents regardless of query complexity. Can an RL agent learn to dynamically select both how many documents and their order based on what helps the generator produce correct answers?

Explore related Read →

Can retrieval systems ground answers in the right time?

Explores whether document retrieval for language models can distinguish between multiple versions of the same content from different time points, and whether adding temporal awareness to retrieval scoring helps answer time-sensitive questions accurately.

Explore related Read →

Can smaller language models outperform larger ones at graph extraction?

GraphRAG pipelines typically use large models for knowledge graph construction, but does extraction actually require world knowledge or just language skills? This explores whether compact specialized models can match larger general-purpose alternatives.

Explore related Read →

Why does retrieval-augmented generation fail in production?

RAG systems work in controlled demos but break in real-world deployment, especially for high-stakes domains like medicine and finance. Understanding the three structural failure modes reveals why.

Explore related Read →

Do vector embeddings actually measure task relevance?

Vector embeddings rank semantic similarity, but RAG systems need topical relevance. When these diverge—as with king/queen versus king/ruler—does similarity-based retrieval fail in production?

Explore related Read →

Domain Specialization in LLMs

9 notes

Does LLM assistance help clinicians build better differentials?

A randomized study tested whether giving clinicians access to an LLM improved their diagnostic reasoning on challenging cases. Understanding this matters for evaluating AI's role in clinical decision support beyond standalone performance.

Explore related Read →

Do benchmark gains in medical AI reflect real-world progress?

Med-Gemini achieves 91.1% on MedQA, but clinician review found ~7% of questions have missing information or labeling errors. The question is whether such benchmark improvements actually signal meaningful advances in clinical capability.

Explore related Read →

Does medical model architecture or training data drive performance gains?

MedGemma claims its medical improvements come from domain-specific training data rather than architectural changes. But the evidence comes only from the developers' own benchmarks, without independent validation or real-world clinical testing.

Explore related Read →

Can machines learn to predict which research ideas will work?

Can a fine-tuned language model with access to published papers predict which unimplemented AI ideas will succeed empirically, and would it outperform human researchers making the same judgment?

Explore related Read →

How much peer review text shows signs of LLM modification?

Researchers analyzed AI conference reviews to estimate what fraction might have been substantially altered by large language models. Understanding this helps clarify how AI tools are entering academic peer review.

Explore related Read →

How much scientific writing has LLMs actually modified?

Researchers estimated the fraction of LLM-modified content across nearly one million scientific papers from 2020 to 2024. Understanding these population-level trends matters for gauging AI's real impact on scientific publishing.

Explore related Read →

When do graph databases outperform vector embeddings for retrieval?

Vector similarity struggles with aggregate and relational queries that require traversing multiple entity connections. Can graph-oriented databases with deterministic queries solve this failure mode in enterprise domain applications?

Explore related Read →

Does medical pretraining actually improve model performance?

When medical models are fairly compared to their base counterparts with per-model prompt tuning and statistical significance testing, do they consistently outperform on medical question answering tasks?

Explore related Read →

Do LLM benchmarks actually measure what they claim to measure?

A systematic review examined whether 445 LLM benchmarks have sound construct validity—whether their tasks and metrics truly capture the phenomena they're designed to test. This matters because flawed benchmarks can mask model failures and mislead research.

Explore related Read →

Knowledge Graphs

2 notes

Can community detection enable RAG systems to answer global corpus questions?

Standard RAG struggles with corpus-wide questions that require understanding overall themes rather than retrieving specific passages. Can graph community detection overcome this limitation at scale?

Explore related Read →

How vulnerable is GraphRAG to tiny text manipulations?

GraphRAG converts raw text into knowledge graphs for question answering. This explores whether adversaries can degrade accuracy with minimal edits to source documents, and what makes the system susceptible.

Explore related Read →