Skip to content
New on arXiv: a judge's uncertainty decomposes, so expert labels go exactly where they remove error
Expert judgement, turned into signal

Teaching machines the human definition of quality.

The signal to evaluate your AI agents - and to benchmark and train the models themselves.

Used by teams at
  • Accenture
  • Instrumentl
  • MIT
  • SentiSum
  • Palantir
  • Leena AI
  • ETH Zurich
  • DigitalGenius
  • Qapitol QA
The numbers 2026
90% Agreement with domain experts on flagged failures.
1M+ Production outputs evaluated - the largest customers run 10K+ a day.
30+ AI teams across healthcare, fintech, CX, legal and multi-agent systems.
2-4 weeks To production, against a 3-6 month internal build.

Six ways a naive judge fails - and the research we've done to fix them.

01
Written in advance
The bar for this case

Predefined rubrics don't learn and miss case-specific nuance.

02
Checked
Never there

It verifies presence, not absence.

arXiv:2608.31016
03
Named modes
Not yet named

It only checks for failures you wrote a test for.

The failure ontology
04
Reliability
Validity

Tune it for agreement and it becomes reliably wrong.

05
Needs an expert
Reviewed

Its own confidence is the wrong escalation signal.

arXiv:2609.06444
06
Sampled
Never looked at

Make it accurate and it costs too much to run on everything.

arXiv:2604.13717

Our judge paper is accepted at two ICML 2026 workshops - RewardBench 2, 71.7 to 85.8.

Four papers on arXiv · two open benchmarks · the failure-mode taxonomy published

A generic judge is a weekend's work. One your experts would stand behind is three years of ours.

Where this fits, and what it gives you.

Five capabilities, one integration
01

Guardrails on quality

A calibrated judge checks every output before it reaches a user, and holds back the ones that fail.

02

Monitoring that learns

A score on everything production runs, sharpening every time one of your experts corrects it.

03

Failure-mode discovery

Failures group themselves into named modes, surfacing the risks nobody thought to look for.

04

Expert escalation

The cases worth an expert's time are routed to one - never the whole stream.

05

The audit trail

Every judgement on the record, with its reasoning and the evidence it cited.

Start with one. The others arrive on the same integration.

Research

Building an AI judge is easy. Building one that works isn't.

We measure ours against human ground truth and publish the numbers, the techniques that failed included.

Papers and open benchmarks
2026 · preprint

Decomposing LLM-judge uncertainty to target expert labels

A judge's uncertainty splits into your experts' genuine disagreement and its own fixable ignorance. Targeting the fixable half removes 83% more error per expert label.

arXiv:2609.06444
2026 · ICML workshop paper

On Cost-Effective LLM-as-a-Judge Improvement Techniques

Two drop-in techniques take an LLM judge to 85.8% on RewardBench 2, up 13.5 points from a 71.7% baseline, at 1.3x the cost. The ones that didn't win are in the paper too.

arXiv:2604.13717
2026 · open benchmark

PrimeBench

400 graded response pairs across finance, medicine, news and technical QA, built to catch evaluators that can't tell nuanced differences apart. Over 6,000 downloads on Hugging Face.

Hugging Face
2026 · taxonomy

An ontology of LLM failure modes

Eight families, about seventy named failures, synthesised from the literature and our own production work. Published for anyone to use.

Read it

Also on the shelf: the two clinical papers behind OmissionBench - a verified census of three deployed AI scribes, where one note in three carries a verified failure, and what recovers a judge's blindness to omissions.

All research
For frontier AI

Expert signal for the models themselves.

Expert-graded data, independent benchmarks, and evaluation on the tasks where the right answer takes an expert to recognise - built with clinicians and domain experts across fields, on published methods and open benchmarks.

See it in action

See what we find in a real clinical AI output

See how Composo evaluates a real clinical AI output - with analysis, source citations, and expert corrections that compound over time.

See it on your data.

Send us a handful of production traces. We'll deliver scored results with a failure report - what's going wrong, how often, and how severe. Takes under a week. Your team reviews it, and if it doesn't match their judgement, you've lost nothing.