Here is the uncomfortable truth about LLM evaluation: most teams either have no evals at all, or they have evals that pass while production quality quietly rots. The second group is worse off. The first group at least knows they are flying blind.
I have watched this movie from both seats — building eval suites for our own products and inheriting other people’s. The pattern is consistent. A team ships an AI feature on vibes. Something embarrassing reaches a customer. Leadership demands “testing.” Someone stands up an eval framework in an afternoon, scores the model against fifty hand-picked examples, posts an 87% in Slack, and everyone relaxes.
Nothing about that number predicts what happens in production.
This essay is the guide I wish someone had handed me: what LLM evaluation is actually for, which metrics mean something, what the framework landscape looks like without the vendor gloss, and the build order that gets you evals that catch regressions before your customers do.
Why Most LLM Evaluation Is Theater
Benchmark worship. MMLU, HumanEval, leaderboard deltas — useful for choosing a base model, nearly useless for your product. A model’s score on graduate-level reasoning questions tells you nothing about whether it extracts payment terms from your contracts correctly. Public benchmarks measure the model. Your evals must measure your system — the model plus your prompts, your retrieval, your tools, your data.
The fifty-example vibe check. Fifty hand-picked examples — usually written by the same person who wrote the prompt, usually covering the happy path — is not an eval suite. It is a demo with a scoreboard. The failures that hurt you live in the long tail: the malformed input, the ambiguous request, the document in a format nobody anticipated.
Goodhart’s law, LLM edition. The moment a metric becomes the target, teams optimize the metric. I have seen prompt changes that moved a judge score four points while making real outputs worse — longer, more hedged, more confidently padded. The metric went up because the judge model likes length. Nobody checked.
I keep watching the same scene play out, and it has not changed in two years. A team ships a document-summarization feature. The eval suite is green — 87%, the number everyone remembers from the Slack thread. Then a customer forwards back a summary in which the model has confidently inverted a clause — the contract says the agreement survives termination, the summary says it doesn’t. The suite never had a prayer of catching it, because the person who wrote the eval examples also wrote the prompt, and nobody writes tests for the mistakes they don’t know they make. The eval wasn’t wrong. It was answering a smaller question than everyone in that Slack thread believed it was.
The Three Layers That Actually Predict Production
Evaluation that works is not one score. It is three layers, each catching what the previous one cannot.
Layer 1: Assertions — evals as unit tests. Deterministic checks on things that are objectively right or wrong: the output parses as JSON, the schema validates, the extracted date is a date, the cited document actually exists in the corpus, the SQL runs. No judge model, no ambiguity, near-zero cost. For extraction and structured tasks, assertions cover more ground than most teams expect — and they never lie.
# Assertions: the layer that never lies (pytest-style)
def test_contract_extraction(golden_case):
out = extract_terms(golden_case.document)
assert out.validates(ContractTerms) # schema, not vibes
assert out.payment_days in range(0, 181) # sane business bounds
assert out.governing_law in KNOWN_JURISDICTIONS
# every cited clause must exist in the source document
for cite in out.citations:
assert cite.text in golden_case.document.text
def test_no_regression_on_past_failures():
# the golden set IS the spec: every past production
# failure, replayed on every pull request
for case in load_golden_set("contract-extraction"):
result = extract_terms(case.document)
assert result.passes(case.assertions), case.incident_id
Two things to notice. The bounds are business logic, not ML logic — a payment term of 4,000 days is not a “low-confidence prediction,” it is a bug. And the second test is the flywheel made executable: the golden set replays every past incident on every change. When this fails, it fails with the incident ID attached. Nobody debates it.
Layer 2: Judged quality — rubrics, not ratings. For subjective qualities — is the summary faithful, is the tone right, did it answer the actual question — you use a model as judge. Done naively this is where evals start lying (more on that below). Done well, it means a written rubric with concrete criteria, scored per-criterion, calibrated against human labels on a sample you re-check every so often.
Layer 3: Online signals — what users do. Regenerate rates, edit distance between the draft and what the human actually sent, thumbs-down clusters, task abandonment. This is the ground truth the other layers approximate. The teams that win wire these signals back into their golden set: every production failure becomes tomorrow’s regression test.
That last sentence is the whole game. An eval suite is not a test you write once. It is a flywheel fed by production.
LLM Evaluation Metrics That Mean Something
Metric names matter less than what task they are attached to. By task type:
Extraction and classification: precision, recall, and exact-match against labeled examples. Boring, decades old, still correct. If your task has a right answer, measure it like it has a right answer.
RAG and retrieval systems: two families, measured separately — retrieval quality (did the right chunks surface: hit rate, recall@k) and generation faithfulness (is every claim in the answer grounded in what was retrieved). Conflating them is why teams “fix the prompt” for a month when the retriever was the problem all along.
Generation and rewriting: rubric-scored criteria — faithfulness, completeness, tone, length discipline — each scored separately. A single “quality: 8/10” is a judge’s mood, not a measurement.
Agents and multi-step systems: task completion rate first, then trajectory checks — did it call the right tools in a sensible order, did it stop when it should, what did it cost. Agent systems fail in ways single calls cannot, and end-to-end success is the only metric users feel.
Notice what is absent: BLEU, ROUGE, and their cousins. Word-overlap metrics predate instruction-following models and mostly measure whether two texts share vocabulary. If you are reporting ROUGE on a summarization product in 2026, you are measuring the wrong thing with precision.
The Framework Landscape, Honestly
Here is the take that will annoy the vendors: the framework is the least important decision in your eval stack. The labeled examples are the asset. The rubrics are the asset. The production-failure flywheel is the asset. The framework is plumbing around them, and good plumbing is table stakes now.
That said, the landscape in one honest paragraph each:
DeepEval — open-source, pytest-style, batteries included for the common metrics (faithfulness, relevance, RAG triad). The fastest path from zero to CI-gated evals if your team lives in Python. Ragas — the de facto standard vocabulary for RAG metrics; use it for retrieval systems even if you adopt nothing else from it. promptfoo — config-driven, great for side-by-side prompt and model comparisons; the quickest honest answer to “is the new model actually better for us.” LangSmith and Langfuse — tracing platforms that grew eval features; their strength is that your eval examples come straight from logged production traces, which is exactly where they should come from. Braintrust — polished eval-first platform with strong diffing and review workflows; the one non-engineers can actually use. OpenAI Evals — historically important, fine inside that ecosystem.
| Framework | Shape | Strongest for | Watch out |
|---|---|---|---|
| DeepEval | Open source, pytest-style | Zero-to-CI fastest for Python teams | Built-in metrics still need calibration |
| Ragas | Open source, RAG-focused | The standard RAG metric vocabulary | Narrow outside retrieval systems |
| promptfoo | Open source, config-driven | Side-by-side prompt & model comparisons | Less suited to deep custom pipelines |
| LangSmith / Langfuse | Tracing platforms + evals | Eval cases drawn from production traces | Gravity pulls you into their stack |
| Braintrust | Hosted, eval-first | Review workflows non-engineers can run | Commercial; data leaves your walls |
| OpenAI Evals | Open source registry | OpenAI-ecosystem benchmarking | Less momentum than the newer tools |
Pick whichever one your team will actually run on every pull request. The eval suite that runs is infinitely better than the sophisticated one that does not.
LLM-as-Judge: Useful, Biased, Not a Substitute
Using a strong model to grade outputs is the only way subjective evaluation scales. It is also where eval suites quietly go wrong, because judge models have documented, reproducible biases:
Verbosity bias — longer answers score higher, independent of quality. Position bias — in pairwise comparisons, the first option wins more than it should; always score both orderings. Self-preference — models rate their own family’s style above others’. Rubric drift — a judge without concrete criteria invents its own, differently each run.
The mitigations are unglamorous and they work: write rubrics with specific, checkable criteria; score criteria separately; randomize answer order and average both directions; pin the judge model version; and — this is the one everyone skips — calibrate against human labels. Take a hundred outputs, label them yourself, and measure how often the judge agrees with you. If agreement is low, your judge is a random-number generator with good grammar, and every dashboard built on it is fiction.
What a working rubric actually looks like — concrete, checkable, scored one criterion at a time:
| Criterion | Scoring rule (0–2, scored separately) |
|---|---|
| faithfulness | Every factual claim appears in the ticket thread or a linked KB article. One unsupported claim = 0. |
| completeness | Addresses every question the customer asked — count them. Missing one = 1. Missing more = 0. |
| actionability | The customer can act on the reply without writing back to ask “how?” |
| tone | Matches the support voice guide: direct, no blame, no “I apologize for any inconvenience.” |
Not a rubric: “Rate the overall quality from 1 to 10.”
The same discipline applies to any evaluation data you trust: I have watched expensive projects fail on evaluation sets nobody had actually verified.
A Build Order That Works
You do not need a platform decision to start. You need two weeks of discipline:
First: collect twenty real failures. Not hypotheticals — actual bad outputs from production or honest dogfooding. These are your first golden set, and they are worth more than five hundred synthetic examples, because they encode how your system actually fails.
Second: write assertions for everything objective. Format, schema, grounding, length, required fields. Wire them into CI so a prompt change that breaks structure fails the build like any other regression.
Third: add a rubric judge for the subjective residue — and calibrate it against your own labels before you trust a single number it produces.
Fourth: close the loop. Every production failure that reaches a human gets added to the golden set the same week. Review the suite monthly; delete tests that no longer discriminate. Treat the eval suite like the product asset it is — because by month six, it is the most defensible thing you own: a specification of what “good” means for your product, written in examples. And once agents act on real systems, that specification is also what your monitoring and governance hang off.
Two questions that decide whether any of this survives contact with your roadmap. First, cost: an eval run is not free — judge calls on a thousand-case suite add up — but it is rounding error next to one bad release reaching customers. Run assertions on every commit, the judged suite on merges and model changes, the full calibration check monthly. Second, ownership: evals die when they are everyone’s job. Give the suite an owner the way you give the build an owner — one person accountable that it runs, that it grows, and that a red result blocks a ship. If nobody’s name is on it, you are six weeks from green-dashboard theater.
The flywheel sounds abstract until the first time it saves you. Here is what it looks like when it works: a model upgrade lands, and the debate that used to run on adjectives — “the new one feels smarter” — runs on evidence instead. Someone kicks off the golden set over lunch; by the afternoon you are reading a diff, not a vibe. And every so often that diff catches the thing that would have burned you: the shiny new model that aces the demo and quietly breaks the one output format a downstream system depends on. The teams that have this don’t argue in Slack about model upgrades anymore. The suite argues for them, in twenty minutes, with receipts.
The Bottom Line
LLM evaluation is not a leaderboard, a framework, or a dashboard. It is the discipline of writing down what good means for your product — in assertions where you can, in calibrated rubrics where you must — and feeding it every failure production hands you.
The teams shipping reliable AI in 2026 are not the ones with the smartest models. They are the ones who can change a prompt on Tuesday and know by Wednesday whether it made things worse.
If your evals cannot tell you that, they are not lying to you yet. They are not saying anything at all.
LLM Evaluation: The Questions Behind the Searches
What is LLM evaluation?
The practice of measuring an LLM-powered system’s output quality against defined criteria — through deterministic assertions, rubric-based judging, and production signals — so changes can be shipped with evidence instead of vibes. The unit under test is your whole system, not the model.
Which LLM evaluation metrics should I use?
Match the metric to the task: precision and recall for extraction, retrieval hit rate plus faithfulness for RAG, per-criterion rubric scores for generation, task completion for agents. Skip word-overlap metrics like BLEU and ROUGE for modern instruction-following systems.
What is the best LLM evaluation framework?
The one your team runs on every pull request. DeepEval for pytest-style open source, Ragas for RAG metrics, promptfoo for prompt and model comparisons, LangSmith or Langfuse when evals should flow from traces, Braintrust for review workflows. The labeled examples matter more than the tool.
How many eval examples do I need?
Twenty real production failures beat five hundred synthetic cases. Start there, grow the golden set with every incident, and prune monthly. Statistical confidence matters less than coverage of how your system actually fails.
Is LLM-as-judge reliable?
Only when engineered: concrete rubrics, per-criterion scores, randomized ordering, pinned judge version, and calibration against human labels. Uncalibrated judge scores drift with verbosity, position, and the judge’s own style preferences.