Preprint

Evaluation without ground truth

M. Costa

·

·

16 pages

A practical method for scoring agent output when no labelled dataset exists. We describe a rubric-based approach used across six client engagements, in which domain experts write graded criteria once and a judge model applies them continuously. We report where rubric scores agreed with later human review, where they diverged, and the two failure patterns that produced most of the divergence. The method is cheap enough to run on every output in production.

Most production agents launch before a labelled evaluation set exists, because the labels can only come from the work the agent is about to do. This paper describes the rubric method we use to close that gap.

Key findings:

  • Rubric scores matched later human review in 84 percent of sampled outputs across six engagements

  • Most divergence came from two patterns: rubrics rewarding surface completeness, and experts disagreeing with each other

  • Writing the rubric surfaced requirement conflicts that the project brief had hidden

Method: for each engagement, two domain experts wrote graded criteria for a sample of real tasks. A judge model applied the rubric to every output for ninety days, and a blind human panel re-scored a weekly sample.

The practical point: the rubric is not a substitute for ground truth. It is a way of making the absence of ground truth visible and cheap to argue about.

CITATION

Costa, M. (2026). Evaluation without ground truth. Preprint.

CONFRONTO

More research

All

Strategy

Engineering

Process

Opinion

Field notes

FORJA

MENU

Create a free website with Framer, the website builder loved by startups, designers and agencies.