RAG
RAG
Retrieval-Augmented Generation. The dominant pattern for "give an LLM a pile of documents": chunk the docs, embed them, retrieve top-k chunks per query, stuff them in the prompt, generate.
NotebookLM, ChatGPT file uploads, and most "chat with your docs" products are RAG.
Why it works
Raw documents stay raw. Retrieval is cheap. Adding new docs is just re-indexing. Good enough for many tasks.
Why Karpathy contrasts it with the LLM Wiki Pattern
RAG re-derives knowledge from scratch on every query. No accumulation between questions. A subtle synthesis question forces the LLM to find and re-stitch the same fragments every time. Nothing compounds.
The LLM Wiki Pattern inverts this: the LLM compiles a synthesis once into persistent markdown, then queries read from that. Cross-references and contradictions are pre-resolved.
The RAG family (per Context Engineering and GraphRAG (IBM Technology))
- Vanilla RAG — vector similarity over chunked + embedded docs; one-shot retrieve-then-answer. Best for simple lookups.
- Agentic RAG — iterative; the agent does a first retrieval pass, decides if it has enough, fetches more if not. Step up from one-shot. Implements the Agentic Loop over retrieval.
- GraphRAG — graph navigation: what entities are connected to this client, what docs relate to those entities? Vector search fills detail within graph-defined scope.
- Context compression — summarize and rank what reaches the model. Even with large context windows, more noise = worse results.
These are complementary, not alternatives. A real system mixes them: graph to find scope → vector to find detail → compression to fit the window.
Where RAG is just one pillar
Context Engineering and GraphRAG (IBM Technology) frames RAG as one tool inside the broader Context Engineering discipline. The four pillars of a contextual system are connected access, knowledge layer, precision retrieval, and runtime governance — RAG is a precision-retrieval mechanism, not the whole picture.
Where RAG still wins
Karpathy doesn't claim the wiki replaces RAG everywhere. RAG is right when:
- Sources are too numerous or too dynamic to maintain a curated wiki
- Queries are mostly lookup-shaped, not synthesis-shaped
- You don't have a human curator in the loop
The two patterns can coexist — e.g. use a wiki for the curated layer and RAG for the long tail of raw sources.
Where RAG breaks — and why Agent Memory is not just RAG-with-updates
The arxiv 2606.24775 benchmark (via Agent-Native Memory Systems (Akshay Pachaar)) gives empirical shape to what Karpathy argued conceptually. Flat, chunk-and-embed stores fail on two workloads that any long-running agent hits:
- Fact updates. Append-only vector stores keep handing back old versions of updated facts — Pachaar's "hallucinations of the past". Nothing at retrieval time can tell "the current version" from "an old chunk that happens to be semantically similar".
- Cross-session reasoning. Nothing compounds between sessions — the same synthesis is re-stitched from scratch every query, which is precisely Karpathy's original complaint.
The fix is not "better RAG"; it's a different substrate — typed, updating, maintained. See Agent Memory.
Still the enterprise default workload (2026-09-26)
Cedric Clyburn in Essential Skills for Becoming an AI Engineer (IBM Technology) puts RAG at the centre of tier 2 of the AI engineer skill stack and calls it the #1 production use case he sees — HR services, hospitals, customer chatbots, internal Q&A: "Almost every company experimenting with AI wants some version of RAG, even if you're not using embeddings or working with another form of storage." His pipeline is the vanilla one above (ingest → fixed-size chunks → embed → vector store; retrieve + question → context window → grounded answer).
note Tension with The 7 Skills You Need to Build AI Agents (IBM Technology) and Agent-Native Memory Systems (Akshay Pachaar) Clyburn presents fixed-size chunking as the default entry-level recipe; Kopecki treats chunk sizing, embedding semantics, and re-ranking as a deep discipline that sets the ceiling on agent quality, and Pachaar shows flat chunk-and-embed stores fail on fact updates. Not a contradiction so much as altitude: Clyburn's is the on-ramp version, the others describe where it breaks in production.
Sources
- LLM Wiki (Karpathy gist)
- Context Engineering and GraphRAG (IBM Technology)
- Agent-Native Memory Systems (Akshay Pachaar) — 2026-07-02; empirical evidence for flat-store failure modes on fact updates and cross-session reasoning
- Essential Skills for Becoming an AI Engineer (IBM Technology) — 2026-09-26; RAG as tier-2 core skill and #1 enterprise production use case