All posts

"LLM-as-judge: useful, noisy, and how to keep it honest"

When output is free text, teams reach for a model to grade the model. Judges are useful and noisy: they drift, they carry biases, and an uncalibrated 8/10 means nothing. How to calibrate a judge against human labels — and when to skip it entirely.

Evals4 min read7 August 2026by Ahmed
"LLM-as-judge: useful, noisy, and how to keep it honest"

A judge score of 8.2 lands in the dashboard and nobody in the room can say what it means. Is 8.2 good? Was last month's 8.5 meaningfully better? Would a human agree with any of it? Teams building AI features with free-text output — summaries, drafts, answers — end up here quickly, because you can't string-compare a summary, so you ask another model to grade it and hope the number is telling the truth.

Have an AI feature stuck between demo and production?

The gap — reliability, evals, cost control, the plumbing that keeps it running unattended — is exactly the work I do. If that sounds familiar, a short conversation is usually enough to point you the right way.

Book a free consultation

© 2026 Ahmed Fareed. All rights reserved.

LOADING