Teaching machines the human definition of quality.
The signal to evaluate your AI agents - and to benchmark and train the models themselves.
Six ways a naive judge fails - and the research we've done to fix them.
Predefined rubrics don't learn and miss case-specific nuance.
Tune it for agreement and it becomes reliably wrong.
Make it accurate and it costs too much to run on everything.
arXiv:2604.13717Our judge paper is accepted at two ICML 2026 workshops - RewardBench 2, 71.7 to 85.8.
Four papers on arXiv · two open benchmarks · the failure-mode taxonomy published
A generic judge is a weekend's work. One your experts would stand behind is three years of ours.
Where this fits, and what it gives you.
Five capabilities, one integrationGuardrails on quality
A calibrated judge checks every output before it reaches a user, and holds back the ones that fail.
Monitoring that learns
A score on everything production runs, sharpening every time one of your experts corrects it.
Failure-mode discovery
Failures group themselves into named modes, surfacing the risks nobody thought to look for.
Expert escalation
The cases worth an expert's time are routed to one - never the whole stream.
The audit trail
Every judgement on the record, with its reasoning and the evidence it cited.
Start with one. The others arrive on the same integration.
Building an AI judge is easy. Building one that works isn't.
We measure ours against human ground truth and publish the numbers, the techniques that failed included.
Decomposing LLM-judge uncertainty to target expert labels
A judge's uncertainty splits into your experts' genuine disagreement and its own fixable ignorance. Targeting the fixable half removes 83% more error per expert label.
arXiv:2609.06444On Cost-Effective LLM-as-a-Judge Improvement Techniques
Two drop-in techniques take an LLM judge to 85.8% on RewardBench 2, up 13.5 points from a 71.7% baseline, at 1.3x the cost. The ones that didn't win are in the paper too.
arXiv:2604.13717PrimeBench
400 graded response pairs across finance, medicine, news and technical QA, built to catch evaluators that can't tell nuanced differences apart. Over 6,000 downloads on Hugging Face.
Hugging FaceAn ontology of LLM failure modes
Eight families, about seventy named failures, synthesised from the literature and our own production work. Published for anyone to use.
Read itAlso on the shelf: the two clinical papers behind OmissionBench - a verified census of three deployed AI scribes, where one note in three carries a verified failure, and what recovers a judge's blindness to omissions.
All researchExpert signal for the models themselves.
Expert-graded data, independent benchmarks, and evaluation on the tasks where the right answer takes an expert to recognise - built with clinicians and domain experts across fields, on published methods and open benchmarks.
See what we find in a real clinical AI output
See how Composo evaluates a real clinical AI output - with analysis, source citations, and expert corrections that compound over time.
See it on your data.
Send us a handful of production traces. We'll deliver scored results with a failure report - what's going wrong, how often, and how severe. Takes under a week. Your team reviews it, and if it doesn't match their judgement, you've lost nothing.