Preprint
Evaluation without ground truth
M. Costa
·
·
16 pages
A practical method for scoring agent output when no labelled dataset exists. We describe a rubric-based approach used across six client engagements, in which domain experts write graded criteria once and a judge model applies them continuously. We report where rubric scores agreed with later human review, where they diverged, and the two failure patterns that produced most of the divergence. The method is cheap enough to run on every output in production.
Most production agents launch before a labelled evaluation set exists, because the labels can only come from the work the agent is about to do. This paper describes the rubric method we use to close that gap.
Key findings:
Rubric scores matched later human review in 84 percent of sampled outputs across six engagements
Most divergence came from two patterns: rubrics rewarding surface completeness, and experts disagreeing with each other
Writing the rubric surfaced requirement conflicts that the project brief had hidden
Method: for each engagement, two domain experts wrote graded criteria for a sample of real tasks. A judge model applied the rubric to every output for ninety days, and a blind human panel re-scored a weekly sample.
The practical point: the rubric is not a substitute for ground truth. It is a way of making the absence of ground truth visible and cheap to argue about.
CITATION
Costa, M. (2026). Evaluation without ground truth. Preprint.
CONFRONTO