LLM Judge Rubric Builder
Concrete, checkable criteria with 0-2 score anchors, then a ready-to-paste judge prompt and JSON schema. The method from Your LLM Evals Are Lying to You. Runs entirely in your browser.
Pick criteria, then switch tabs. Everything regenerates as you edit.
Before you trust a single number it produces
- Calibrate: label ~100 outputs yourself; measure judge-human agreement per criterion. Below ~80%, fix the rubric before trusting the dashboard.
- Pairwise: when comparing two outputs, score both orderings and average; judges favor whichever answer comes first.
- Pin the judge: a judge model upgrade is a rubric change. Re-calibrate when you swap it.
- Separately means separately: one JSON field per criterion. A single overall score gives you nothing to act on.
The full method, including the biases this protects you from, is in the essay.