Every few months the retrieval world mints a new acronym and declares the previous one dead. RAG replaced fine-tuning as the default answer. Now CAG and KAG are circling RAG the way RAG circled fine-tuning. If you believe the discourse, you are perpetually one architecture behind.
Here is the uncomfortable truth: these three are not generations of progress. They are three answers to one question we have been asking since before LLMs existed — where does the knowledge live, and who pays to keep it fresh? That is a database conversation. The acronyms just dressed it up.
I run retrieval systems in production. Here is what each one actually is, where each one honestly wins, and the decision framework I use when a team asks which flavor to build.
What They Actually Are, Without the Branding
RAG — retrieval-augmented generation. Your knowledge lives in an index (usually vectors, increasingly hybrid). At query time you search it, take the best chunks, and hand them to the model. The model answers from what you retrieved. You pay at query time: retrieval latency, and the permanent engineering tax of keeping retrieval quality high. I have written about where RAG is heading — the short version is that the retrieval step is getting smarter, not disappearing.
CAG — cache-augmented generation. Your knowledge lives in the context window. You load the entire corpus into the model up front, cache the computed state, and answer every subsequent query against that cache. No retriever, no index, no “wrong chunk” failure mode — the model sees everything. You pay up front and at refresh: the corpus must fit in context, and every document change means recomputing the cache.
KAG — knowledge-augmented generation. Your knowledge lives in a structure — a knowledge graph of entities and relationships that the generation step reasons over. Multi-hop questions (“which suppliers feed the product line that failed audit?”) stop being retrieval roulette and become graph traversal. You pay in construction and upkeep: someone has to build that graph, and someone has to keep it true.
Cache, index, or schema. We have been making this exact tradeoff since the first time someone asked whether to denormalize a table. The vocabulary is new. The engineering is not.
Where Each One Honestly Wins
CAG wins when the corpus is small and stable. A policy manual, a product catalog, a contract set measured in hundreds of pages, not millions. If it fits comfortably in context and changes monthly rather than hourly, skipping the retriever is not lazy — it is correct. The failure modes you delete (bad chunking, embedding drift, retrieval misses) are the ones that quietly kill RAG systems. The ceiling is hard, though: corpora grow, and the day yours outgrows the window, you are rebuilding under deadline.
RAG wins when the corpus is large or fast-moving. Millions of documents, constant updates, per-user permissions on what can be retrieved. The index pays for itself precisely because you cannot afford to show the model everything. This is still the default answer for most production systems, and the default is not a dirty word.
KAG wins when the questions are structural. If your users ask things that are secretly joins — across systems, across entities, across time — no amount of better chunk retrieval saves you. The relationships are the knowledge. Engineering and manufacturing data is full of this: parts that belong to assemblies that belong to programs that have suppliers. But a knowledge graph is a product you now own forever, and a stale graph lies with more confidence than a stale index.
The Decision Framework That Ignores the Acronyms
Ask three questions, in order:
1. Does the whole corpus fit in context, and does it change less than weekly? Yes to both: start with CAG, or honestly just long context without the ceremony. You can graduate to RAG later; teams almost never need to graduate in the other direction.
2. Are the hard questions lookups or joins? Lookups — “what does the warranty clause say” — are RAG’s home turf. Joins — “which clauses differ across these forty contracts” — push you toward structure: a graph, or more often, honestly, a relational database the agent can query. KAG is sometimes just SQL wearing a lanyard.
3. Who maintains the knowledge layer in month twelve? This kills more architectures than any benchmark. An index needs re-embedding discipline. A cache needs refresh discipline. A graph needs a librarian. Whichever discipline your team will actually sustain — knowledge layers rot the same way data lakes do — is the architecture you should pick. The best retrieval system is the one that is still true a year after launch.
And if the real problem is that the model does not behave the way you want — tone, format, task patterns — none of these are your answer, and neither is training: that road is mapped in Fine-Tuning is a Trap.
Rapid-Fire: The Questions Behind the Searches
Is CAG better than RAG?
For small, stable corpora — yes, and it is simpler too. Past the context window, or with fast-changing data, CAG is not an option, so the comparison dissolves. They are tools for different corpus shapes, not competitors on one leaderboard.
Does long context make RAG obsolete?
It shrinks RAG’s territory from below. Every context-window increase moves the “just load it all” boundary up. But cost, latency, and permissioning keep indexes alive for large corpora — showing the model everything on every query is a bill, not an architecture.
What is KAG actually for?
Domains where relationships carry the meaning: engineering bills of materials, compliance chains, org and supplier networks. If your questions are joins, structure wins. If they are lookups, a graph is an expensive way to feel sophisticated.
Can you combine them?
Production systems already do: structured store for the join-shaped questions, index for the lookup-shaped ones, long context as working memory. The combination nobody needs is all three built on day one. Start boring. Earn complexity.
The acronyms will keep coming. The question underneath them has not changed in forty years. Answer the question, not the acronym.