This is the first installment of Failure Modes, a monthly series. One production failure pattern, why it happens, and the fix, every first Tuesday. If you want it in your inbox, the Substack gets each one the day it ships.

The failure pattern I see most often in production LLM systems is also the one no alert ever fires for. It goes like this. The system launched strong. Users were happy. Nobody shipped anything for three weeks. And somewhere in that quiet stretch, answer quality started sinking, a point or two at a time, until someone important forwarded a bad output with the subject line “is this thing getting worse?”

It was. It had been for a while. And the uncomfortable part is that every dashboard was green the entire time, because the dashboards were watching the wrong layer. Latency was fine. Error rate was zero. Token spend was normal. The system was failing in the only dimension that matters, the quality of what it says, and that dimension was not instrumented.

Traditional software does not fail this way. Code that passed its tests keeps doing what it did; when it breaks, it breaks loudly, with a stack trace and a timestamp. LLM systems fail like produce, not like code. They rot. Quietly, gradually, from the inside, while looking fine on the shelf.

The Five Places the Rot Comes From

Silent degradation is not one failure. It is five independent ones that happen to share a symptom, and diagnosing which one you have is most of the fix.

1. The model moved under you. You call an API. The provider updates the model behind it, retunes a safety layer, or shifts traffic between deployments. Your prompt was tuned against last month's behavior and nobody told you this month's is different, because from the provider's side nothing is wrong. Pin model versions where the platform allows it, and treat any unpinnable alias as a dependency that changes without notice, because it is one.

2. Your corpus drifted away from your index. Retrieval quality decays the moment documents change faster than embeddings are refreshed. Deleted pages keep answering from the index. New policy supersedes old policy, but the old policy has six months of accumulated vector similarity. The system confidently cites things that are no longer true, which is worse than failing, because it fails with citations. I measured how much the chunking layer alone moves retrieval in this month's data report, and that is with a frozen corpus; a moving one compounds it.

3. Prompt rot. Every incident adds a sentence to the system prompt. None of them are ever removed. Eighteen months in, the prompt is four thousand tokens of contradictory special cases, and the model obeys whichever instructions sit closest to the end. Each addition made sense in its week; the accumulation is incoherent. This one is self-inflicted, which also makes it the most fixable: prompts need the same refactoring discipline as code, and a token budget per section that someone owns.

4. Your users changed the question. The query distribution you launched against is not the one you serve six months later. New user segments arrive, power users develop shorthand, the product grows a feature that funnels in a question shape nobody evaluated. The system did not get worse at its old job; it inherited a new job silently, and nobody re-ran the interview.

5. Stale cache, fresh world. Cached context and precomputed summaries are a cost optimization right up until they outlive the truth they summarized. A cache without an expiry policy tied to source-document changes is a slow-release lie.

The Observability Stack That Catches It

None of the five announce themselves, so detection has to be built as its own layer. The stack that works is small, and almost every team builds it only after the forwarded-email incident. Build it before.

Canary prompts, daily. A fixed set of thirty to fifty representative inputs with known-good outputs, run against production on a schedule and scored automatically. This is the single highest-leverage piece: when canary scores move and you shipped nothing, something upstream did, and you find out Tuesday morning instead of quarter end. Canaries are also the only clean way to separate failure one (model moved) from failure four (users moved), because the canary inputs never change.

Sampled judge scores on live traffic. Score a few percent of real responses with an LLM judge against a calibrated rubric, and trend the result per criterion, not as one blended number. A blended score sinking tells you to worry; a per-criterion view tells you where to dig. Groundedness falling while tone holds points at retrieval; the reverse points at the prompt or the model. The Rubric Builder outputs the judge prompt and schema this needs.

Retrieval metrics against a living golden set. Recall against a fixed test set measures your January corpus forever. The golden set has to grow, feeding every incident and every flagged output back in as a labeled case, or your eval suite becomes a museum of problems you used to have.

Input drift alarms. Embed a sample of incoming queries and track the distribution against your eval set's distribution. When the distance trends up, failure four is underway, and the fix is new eval cases before new prompts.

The boring dashboards, kept honest. Latency, cost per request, token counts per section of the window. Not because they catch quality rot, but because they catch its common causes: a context assembly bug shows up as a token-count step change weeks before anyone notices the answers.

The Fix Protocol

When the alarms do fire, resist the universal first instinct, which is to edit the prompt. Prompt edits against an undiagnosed failure are how you get failure three while chasing failure two. The order that works: reproduce on canaries, diff the layer that moved (model version, index freshness, query distribution, prompt history, cache age), fix that layer, and then add the case to the golden set so the same rot cannot return silently. Diagnosis first. The prompt is the last thing you touch, not the first.

The teams that handle this well do not have more dashboards than everyone else. They have the five failure sources written down, an alarm pointed at each one, and the discipline to diagnose before they edit. That is the entire difference between finding out from a canary and finding out from a customer.

FAQ

What is LLM observability?
Instrumenting an LLM system so quality changes are visible: canary prompts on a schedule, judge-scored samples of live traffic, retrieval metrics against a maintained golden set, input drift alarms, and the standard latency and cost dashboards underneath.

Why do LLM systems degrade without a deploy?
Because their behavior depends on things outside the repo: the provider's model, the corpus and its index, the user query distribution, accumulated prompt edits, and cached context. Any of them can move while your code stays still.

What are canary prompts?
A fixed set of representative inputs with known-good outputs, run against production daily and scored automatically. When canary scores move and nothing shipped, something upstream changed. Cheapest early warning an LLM product can have.

How do I tell provider model drift from user drift?
Canaries never change, users do. Canary scores falling means the model or your stack moved. Live scores falling while canaries hold means the questions moved. Instrument both and the diagnosis is one glance.