Eval Sample Size Calculator
The most common failure in LLM evaluation is not a bad metric. It is declaring victory off 50 cases, where a 6-point “improvement” is indistinguishable from noise. This calculator answers both directions of the question: how many cases you need to detect a given difference, and what difference your current suite can actually see. It knows the two things generic power calculators do not: most LLM evals are paired (both systems run the same set), and your judge is not a perfect oracle.
Required cases vs. detectable difference
The math, stated plainly. Two-sided α with zα/2 ∈ {1.645, 1.960, 2.576} and power zβ ∈ {0.842, 1.282, 1.645}.
Paired design (McNemar approximation): n = (zα + zβ)² · p_disc / δ², where p_disc is the share of cases the two systems resolve differently; the same eval set needs far fewer cases than two independent ones, which is one more reason to keep a fixed golden set.
Independent arms: n/arm = (zα + zβ)² · (p₁q₁ + p₂q₂) / δ².
Single rate: n = zα² · pq / E².
An imperfect judge that agrees with humans a% of the time (symmetric errors) attenuates the observable difference by (2a − 1), which inflates required n by 1/(2a − 1)²; at 85% agreement that is a 2.04x penalty, which is why judge calibration is not optional. These are normal approximations, fine for n > 30 and rates away from 0 or 100; for tiny suites or extreme rates use exact tests. Full argument in Your LLM Evals Are Lying to You. The engine is open source, tested, with a CLI: github.com/AlexRyan92/eval-sample-size.