The Insanity of Relying on Vector Embeddings: Why RAG Fails

Paper · Source
Retrieval-Augmented Generation (RAG)LLM Failure Modes

Wrong Tool for the Job

RAG fails in production because vector embeddings are the wrong choice for determining percentage of sameness. This is easily demonstrated. Consider the following three words:

King

Queen

Ruler

King and ruler can refer to the same person (and are thus considered synonyms). But king and queen are distinctly different people. From the perspective of percentage of sameness, king/ruler should have a high score and king/queen should be literally zero.

In other words, if the query is asking something about a “king” then chunks discussing a “queen” would be irrelevant; but chunks discussing a “ruler” might be relevant. Yet, vector embeddings consider “queen” to be more relevant to a search on “king” than “ruler.” Here are the vector similarity scores for queen and ruler when compared to king using OpenAI’s ADA-002 embeddings:

King

Queen: 92%

Ruler: 83%

Reranking

Reranking is perhaps the most recommended Advanced RAG strategy. However, as the RankRAG study shows, even using a fine-tuned model for reranking only results in a 54.2 score on EN.QA. Using general reranking models had an even worse score.

GraphRAG and Knowledge Graphs

A recent study on KG-RAG (RAG enhanced with Knowledge Graphs) showed an F1 score of 25% and an accuracy of 32% for CWQ dataset. Interestingly, Knowledge Graph RAG had a lower accuracy than regular embedding RAG (which had a 46% accuracy).

Real-World vs Hype

There simply is no study showing that vector embeddings, combined with dozens of Advanced RAG techniques, results in a reliable chatbot in production environments containing numerous documents. Moreover, the added latency of many Advanced RAG techniques makes them impractical for real-world chatbots—irrespective of the accuracy issue.

But even if larger models could overcome the problem, they would be slower and more expensive. In other words, they’d be too slow and too expensive for any practical purpose. Would companies pay more for a chatbot than for a person, when the chatbot would require up to a minute for each unreliable answer?

That’s the actual state of RAG. That’s the actual outcome of relying on vector embeddings.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

When do simpler collaborative filtering approaches outperform complex LLM recommenders? Why do vector embeddings fail at capturing task-relevant relationships? How should recommendation systems balance individual preference and diversity? What representations best capture screen understanding for task execution? When should retrieval systems decide to fetch new information? What limits language model accuracy in evaluating ideas? How do knowledge graph structures enable efficient multi-hop reasoning and retrieval?