Contextual Retrieval cuts RAG failure rates by prepending chunk-specific context before embedding, so your knowledge base actually knows what it’s talking about. (Anthropic)
- Anthropic's Contextual Retrieval prepends chunk-specific context to each piece of text before embedding, fixing the core RAG problem where chunks lose meaning when separated from their source documents.
- Combined with BM25 and a reranking step, the method cuts retrieval failure rates by 67%, tested across codebases, fiction, and academic papers.
- Prompt caching makes this practical at scale: generating context for a million document tokens costs around $1.02.