Every LLM guardrails demo looks the same. The prompt injection gets caught, the toxic output gets blocked, the PII gets masked, and the room nods. Then ask the same team six months after launch how the guardrails are doing and you get one of three answers: we loosened them, we turned most of them off, or nobody is sure anymore what they actually catch. The demo was real. The guardrails just did not survive contact with live traffic.
That is not an argument against guardrails. It is an argument against how they usually get built: as a checklist bolted on in the last sprint before launch, with no latency budget, no false positive target, and no decision about what should happen when one fires. Guardrails built that way die. Here is how they die, and what the ones that survive look like.
Guardrails Are Not Evals
Start with the distinction most teams blur. Evals run before you ship: a fixed dataset, scored offline, deciding whether this version is good enough to deploy. Guardrails run during the request: live traffic, one response at a time, deciding whether this specific output reaches this specific user. Observability runs after the fact and tells you how the whole system is trending. Three layers, three jobs. Evals gate the deploy, guardrails gate the response, observability watches the trend line.
The reason the distinction matters is that the two ends have opposite economics. An eval can take a minute and cost real money per case, because it runs a few hundred times before a release. A guardrail runs on every single request, inside the user's patience and your margin. A check that is a rounding error in an eval suite is a tax on every interaction in production. Most dead guardrails died because someone designed them with eval economics and deployed them into request economics.
The Three Ways Guardrails Die
Death by latency. A guardrail is code in the request path. A model-based input check that takes 400ms in front of a response that takes 2.5 seconds is a 16 percent latency tax before the model has read a word. Stack an injection classifier, a topic filter, and an output scan sequentially and the tax compounds, and it compounds worst at p95, which is the latency your angriest users live at. Teams respond exactly the way you would predict: under pressure to make the product feel fast, the checks get loosened, sampled, or quietly removed. The latency budget was never written down, so the guardrails lose the argument one standup at a time.
Death by false positives. Everyone measures what the guardrail catches. Almost nobody measures what it wrongly blocks, and the second number is the one that kills it. A PII filter that regexes anything digit-shaped starts eating order numbers. A self-harm filter starts blocking a nurse's documentation workflow. Each false positive is a support ticket at best. At worst it is silent: a blocked user does not file a bug, they rephrase and retry until something goes through, which means your most motivated users are now running a small red-team against your filters as a daily workflow. A guardrail with an unmeasured false positive rate is not a safety control. It is a superstition with latency.
Death by quiet bypass. This one is organizational. An enterprise customer hits a false positive during a pilot, so an allowlist entry goes in, marked temporary. An oncall disables a check at 2 a.m. to stop a page, and the ticket to re-enable it ages out. Eighteen months later the guardrail config is an archaeology site, and the only honest answer to “what do our guardrails block?” is “nobody knows.” There is a technical flavor of the same death: if your guardrail is itself an LLM judge, it is itself injectable, and text that talks its way past the judge is not a hypothetical, it is the standard second move of anyone probing your system. A guardrail that nobody re-tests after launch decays exactly like the system it guards.
Block, Flag, or Log
The design decision that separates surviving guardrails from dead ones is not which framework you pick. It is deciding, per check, what happens when it fires. There are only three honest options, and the ranking is the opposite of what most teams ship.
Log is the default. Most checks should observe and record without touching the response at all. Logged signals cost nothing at request time, have no false positive blast radius, and feed the trend lines your observability layer actually needs. A spike in injection-shaped inputs is worth knowing about even when every individual attempt failed.
Flag is the middle tier: let the response through, or route it to a degraded path, and queue it for review. Flagging is where model-based checks belong, because it converts their false positives from blocked users into review items.
Block is reserved for the irreversible. Credentials or PII leaving the boundary. A tool call that moves money, sends an email, deletes a record. Anything where no apology fixes it. Blocking is the only tier where a false positive is an acceptable price, and that is precisely why it should be the rarest tier.
The pattern underneath: gate actions, not words. A model that writes something embarrassing is a flag. A model that does something irreversible is the thing the block tier exists for. This site is built on that exact split: an AI agent writes most of the commits, and a hook refuses the push, so a human runs every deploy to main. The agent's words are reviewed; the agent's irreversible action is simply not available to it. That one blocking gate does more than any output filter on the page, and you can read the full setup.
The Stack That Survives
Deterministic checks first. Schema validation on structured outputs, allowlists on tool arguments, length caps on inputs, regex for the PII that actually has a shape (card numbers, SSNs, API key prefixes). These run in microseconds, cost nothing, and fail loudly and reproducibly. Every check you can make deterministic is a check that will never lose the latency argument. Most teams invert this and reach for a model first, which is how you end up paying inference prices for what a regular expression does for free.
Then one model-based check, where it earns the spot. If semantic filtering is genuinely needed, run one small, fast classifier, not a chain of them. Moderation endpoints and the small safety-classifier models exist for exactly this seat. Run it in parallel with generation where the flow allows, so its latency hides behind the model's. And know the streaming conflict up front: you cannot fully scan an output you are already streaming to the user. Either buffer and pay the latency, scan chunk-by-chunk and accept weaker guarantees, or stream freely and reserve blocking for the action layer. Pretending this tradeoff does not exist is how output filters end up silently disabled the week streaming ships.
The system prompt is a guardrail layer too. Scope restrictions and refusal instructions in the prompt reduce how often the output layer has to intervene. But prompt instructions accumulate the same rot as everything else in the context window, so they need the same budget discipline: a known token cost, an owner, and a periodic prune.
Capability scoping over input filtering. For prompt injection specifically, filters catch known patterns and miss novel ones, and the miss rate is not a solvable problem, it is the nature of the attack. The defense that holds is boring: the model cannot leak a document it was never given, and cannot take an action it has no tool for. Scope the context per request, gate the tools per action, and let the injection filter be a logged signal rather than the wall.
Finally, your guardrails need evals. A guardrail is a classifier, and a classifier has a precision and a recall whether you measure them or not. Keep a labeled set of real blocked and real legitimate traffic, re-run it on a schedule, and trend both numbers. A calibrated rubric works for the judge-based checks. This closes the loop with the bypass problem too: the allowlist archaeology stops when every exception has to survive the next scheduled re-test. The teams whose guardrails survive are not the ones with the most checks. They are the ones who can tell you, with numbers, what each check catches, what it costs, and what it breaks.
FAQ
What are LLM guardrails?
Runtime checks in the request path of an LLM application. They validate input before the model and output before the user, and each check either blocks, modifies, flags, or logs. Unlike evals, they run on every live request and pay for themselves in latency.
What is the difference between guardrails and evals?
Evals run before deployment against a fixed dataset and decide whether to ship. Guardrails run at request time and decide whether one response reaches one user. Observability runs after the fact and watches the trend. Production systems need all three, and each has different economics.
Do LLM guardrails add latency?
Every synchronous one does. Deterministic checks are effectively free; model-based checks add real time per call, and stacking them sequentially compounds at p95. Write the latency budget down, and move every check that does not need to block out of the synchronous path.
How do you defend against prompt injection?
Not with filters alone. They catch known patterns and miss novel ones. The defense that holds is capability scoping: the model cannot leak what it was never given and cannot do what it has no tool for. Gate the actions, scope the context, log the filter hits.