# Why evaluating AI output isn't like testing software

- Published: 2026-09-30
- Reading time: 8 min
- Tags: AI Evaluation, Quality, Engineering Process

Traditional tests have one correct answer and an assert that checks it. Evaluating a language model's output means judging a range of acceptable answers that keeps shifting underneath you — and that difference changes the entire quality process.

A team building its first AI-driven features usually reaches for the same testing framework by instinct: give it an input, compare the output to an expected value, and turn it green if they match. That framework works for deterministic code because deterministic code is supposed to do exactly the same thing every time. A language model's output isn't like that. Two identical runs on the same input can produce two different sentences that are both correct; a silent update to the underlying model can shift behavior without your team touching a single line of code. The real problem in AI evaluation isn't how to write better asserts. It's how to build a process that lives with uncertainty instead of pretending it doesn't exist.

### Why exact equality doesn't work

The first thing every team learns is that string-equality tests on language model output almost always either fail constantly or become meaningless. If your assert expects one exact sentence, the test either fails routinely (because the model phrases it slightly differently each time) or gets loosened so much it accepts anything (because it only checks length or the presence of one word). Neither gives you useful information. The alternative is measuring quality against semantic criteria, not literal ones: does the answer contain the correct facts? Is the tone and structure appropriate? Did it say something it shouldn't have? These are questions that either a carefully built rule can check, or a human can, or — carefully — another model playing the role of judge.

### The golden set: a foundation that has to be real

Every serious evaluation process starts with a golden set: a collection of real inputs, paired with acceptable answers or scoring criteria, that every change to the system — a prompt edit, a model swap, a change in surrounding logic — gets run against. The quality of this set matters more than any other part of the evaluation process, because a set that covers easy, safe inputs but skips the edge cases ignores exactly where failures actually happen. The golden set needs to be built from real — or at least realistic — data, not from cases an engineer imagines a user might ask. And it needs to keep growing: every real failure discovered in production should be added to this set, because the best source of edge cases is things that have already actually happened.

### Rubric grading

When the correct answer isn't a fixed string, the practical way to measure quality is defining a rubric: an explicit list of criteria a good answer needs to satisfy, each with a defined weight. For example: are the claimed facts accurate? Is the answer actually relevant to the question asked? Does it avoid stating incorrect information with false confidence? Is its length and tone appropriate to the context? The rubric has to be written before seeing the actual outputs, not after — otherwise the temptation to tune the criteria to justify an answer that's already been produced makes the evaluation worthless. And the rubric needs to be precise enough that two independent graders, human or model, land on similar scores for the same answer; if two graders disagree sharply on one rubric, the rubric itself is the thing that's ambiguous. A small, generic rubric usually looks something like this — not as a final template, but as a starting point for any team that doesn't have one yet:

```
factual_accuracy: 0-3   # each incorrect claim deducts a point
relevance:        0-2   # does it actually answer the question
tone_fit:         0-2   # does the tone match the context
overconfidence:   0-3   # penalty for stating uncertain things flatly
```

The point isn't that these particular weights are correct; it's that each one is explicit, observable and arguable — the exact opposite of a vague, implicit judgment like "it looked good."

### The trap of a model judging itself

Using a language model as a judge to score another model's output is tempting because it's fast and cheap — but it has a specific trap that needs explicit attention: models tend to score outputs that are stylistically closer to their own writing more favorably, even when there's no real difference in accuracy or quality. If the model you use for generation is the same one you use for judging, this bias gets worse: the model is effectively grading its own work. The full fix isn't to never use a model as a judge — for evaluation at scale, that's sometimes the only practical option. The fix is to separate judge from generator (a different model for judging, or at least a distinct version), and to periodically validate the judge model's scores against a sample of human scores. If the judge model's scoring systematically drifts from human scoring — in either direction — that judge is no longer trustworthy, no matter how fast or cheap it is.

### Human review as a stage, not an exception

Human review shouldn't be something that only activates when an output "looks suspicious." It needs to be a fixed stage in the process, with a defined sample drawn from every category of output — not just cases the system flagged as suspicious, because the system can't flag exactly the things it can't see. A regular random sample of outputs that "look fine" needs just as much review as flagged cases, because silent failures — an answer that looks correct but isn't — are exactly the ones no automated rule catches. Human review also needs to feed its findings back into the golden set. A reviewer who finds a failure and only writes it up in a report leaves that same failure reproducible a second time. That case needs to be added to the evaluation set so every future automated run checks for it.

### Regression when the model changes underneath you

One of the strangest differences between AI evaluation and software testing is that your code can stay exactly the same while its behavior changes, because the provider updated the underlying model. This kind of regression doesn't really exist in traditional testing — fixed code produces fixed behavior unless a dependency changes, and that dependency is usually versioned and pinned. Hosted models aren't like that; the version sitting behind an API name can change without explicit notice. The practical solution is running the golden set regularly — not just when code changes, but on a schedule — and comparing scores against the previous baseline. If the average rubric score on the same evaluation set suddenly drops without your team having changed anything, that's itself a warning that deserves to be taken as seriously as any other regression.

### What to measure before calling an AI feature done

Before an AI-driven feature gets declared "ready," at least four things need to be measured, not just felt: success rate against the golden set with a clear definition of success; the distribution of failures — not just what percentage fail, but what kind, because a harmless failure and a misleading one are fundamentally different; cost and latency per request under realistic load, not just in a development environment; and an ongoing monitoring plan after launch — because no amount of pre-launch evaluation substitutes for watching real behavior in production. These four measurements need to be documented before launch, not as a one-time checklist but as something re-measured after launch too. A team that only gathers these numbers once, on launch day, and then forgets about them is effectively repeating the same mistake described in the regression section: assuming that what's correct today stays correct tomorrow, when the very thing this feature is built on — the underlying model — sits outside the team's direct control. What's never enough is a handful of manual spot checks and "it looked good." That sentence is exactly what separates AI evaluation from software testing: in traditional testing, "it's green" is a precise claim. In AI evaluation, only a continuous, documented, repeatable process can make that same claim honestly.


---

Source: https://larsima.com/sr/insights/evaluating-ai-output-is-not-testing
Organisation: LARSIMA (شرکت فن‌آوران توسعه لار سیما), registered in Iran, no. 318515, since 2007-12-02.
