ALL ARTICLES

We Made Software Probabilistic. Then Forgot How to Test It.

Why AI quality starts with defining good, not picking a model.

Imagine changing a prompt, trying five examples, and liking every answer. The responses are fluent. The demo works. You ship it.

What have you actually learned?

That the system can produce five convincing responses. Not how often it succeeds, which requests it fails on, or whether the change made something else worse.

That gap is the starting point for our view at BENCH: probabilistic software needs an explicit quality discipline. A good-looking answer is an observation. It is not yet an evaluation.

01 / The output is no longer a single promise

The familiar mental model for a deterministic function is straightforward: provide an input, execute the code, check the expected output. Hold the relevant state fixed, and the same input gives the same result.

An LLM works differently. At each generation step, it assigns probabilities to possible next tokens. A decoding strategy then selects one. Sampling draws from that distribution; greedy decoding chooses the highest-scoring token. Repeat until the response is complete. [1]

Lowering temperature during sampling concentrates probability on more likely tokens. It can reduce variation. It does not make a wrong answer correct, and a probabilistic model can still be used with deterministic decoding. [2]

THE FAMILIAR MENTAL MODEL

Input
Code
Expected
output

THE GENERATIVE SYSTEM

Input
AI system
Output A
Output B
Output C
01 / Test behavior across cases and runs, not just one convincing output. A, B, and C are illustrative possibilities, not equally likely outcomes.

Then there is the application around the model. Different retrieved documents, a new tool response, reordered context, or a model update can change the result. That is not necessarily random behavior under identical conditions. The effective inputs or system have changed.

Infrastructure adds another consideration: numerical frameworks do not guarantee reproducibility across every release, platform, and execution setting. Repeatability needs its own controls and tests. [3]

Software was never entirely deterministic. Distributed systems and traditional ML already taught us that. LLMs have made this challenge part of everyday application development.

02 / Variation is not the problem. Unmeasured variation is.

A support assistant does not need to use exactly the same words every time. Two different explanations can both be accurate, useful, and appropriate.

But inventing a refund policy is not an acceptable variation. Neither is disclosing information it should not share or confidently skipping a necessary escalation.

The engineering question is not simply, “Can we force the same answer?” It is, “How often does this system meet our requirements, and what happens when it does not?”

Exact-match tests still belong in the toolkit. Use them for schemas, required fields, and deterministic business rules. Add behavioral evaluation where multiple different answers can legitimately succeed.

03 / Machine learning already gave us a starting point

Established supervised ML practice separates the examples used to fit or tune a system from those used to estimate its performance. Test data is held aside because doing well on familiar examples is not enough. Repeatedly tuning against that test set undermines the separation. [4]

That discipline asks useful questions: what data represents the task? Which metric matches the objective? Which failure classes matter? How does the candidate compare with the baseline?

Generative applications do not make those questions obsolete. They make the answers more task-specific.

A fluent paragraph is easy to judge by feel. That makes it tempting to mistake a quick read for a reliable test. But a persuasive demo cannot tell you how the system behaves across the requests your users actually make.

For an LLM application, tuning might mean changing a prompt or retrieval setting rather than training model weights. The need for unseen test cases remains.

04 / Before evaluation comes definition

Our starting point for BENCH is simple: before asking how to measure quality, define what quality means for the task.

“Make the answer better” is not a usable specification. Better for whom? More concise? More complete? More accurate? What must never happen?

Consider a support assistant. A useful rubric could ask whether it:

  • Answers using accurate information and the current company policy.
  • Addresses the customer’s actual problem and explains the next step.
  • Avoids promises or refunds it has no authority to offer.
  • Uses an appropriate tone and escalates when necessary.

These dimensions should not automatically cancel each other out. A warm tone does not compensate for an unauthorized refund. Define critical failures separately from an average quality score.

Research offers a useful parallel. In the EvalGen study, practitioners refined their evaluation criteria after inspecting model outputs. The authors call this criteria drift. The study involved nine experienced practitioners, so it is evidence of a pattern, not proof of what every team does. [5]

Our practical recommendation is to treat examples and criteria as things you develop together. Record rubric changes so a new score is not silently compared with an old definition of success.

There is no universal “good AI.” There is a system doing a particular job, for particular people, under particular constraints.

05 / Optimize the system, not just the model

Which model should we use? It is a reasonable question. It is rarely the only one.

An application combines a prompt, a model, context, tools, retrieval, data, and workflow decisions. Berkeley researchers describe this broader shift as the move toward compound AI systems: interacting components that must work well together. [6]

THE MODEL IS ONE PART OF THE SYSTEM

Prompt
Model
Context
Tools
Retrieval
Data
Workflow

Interacting parts. Shared consequences.

What the user experiences
02 / Evaluate the application end to end. Inspect its parts to understand why it fails.

Imagine an assistant answering from an outdated policy document. A stronger model may still produce the wrong answer. The first thing to investigate is which document retrieval supplied.

Or imagine the right tool exists, but the workflow never calls it. A model leaderboard will not diagnose that failure.

These are illustrative cases, not BENCH experiment results. Their point is practical: evaluate the experience end to end, then inspect the intermediate steps to locate the problem.

Change one factor where practical. Compare against a baseline on the same cases. When outputs vary, include repeated runs. Look at failure categories and critical errors, not only an overall average. Then weigh quality alongside latency and cost.

06 / Quality is a feedback loop

A benchmark before launch gives you a snapshot. It cannot stand in for every future model update, new user request, changed document, or tool failure.

The workflow we advocate is continuous: define good, evaluate, investigate failures, change the system, evaluate again, and monitor what happens in production.

QUALITY IS A LOOP, NOT A LAUNCH GATE

  1. 01Define goodCriteria + examples
  2. 02EvaluateA versioned test set
  3. 03Find failuresGroup + investigate
  4. 04Change the systemA focused hypothesis
  5. 05Evaluate againCheck gains + regressions
  6. 06Monitor productionLearn from real use
Feed what you learn back into the next cycle.
03 / New failures become new test cases. Changed criteria become a new rubric version.

Production failures should inform future test cases, with sensitive data handled appropriately. Keep a separate held-out set for final comparisons. Version the configuration and rubric. If automated judges help score subjective answers, calibrate them against human-reviewed examples and examine disagreements.

This is our proposed engineering discipline, not a claim that one test suite can guarantee reliability. The amount of evidence and human oversight should match the consequences of failure.

For us, that is the point of BENCH: making AI quality explicit enough to measure, compare, and improve.

Not “Does this answer look good?”

“What does good mean, how often do we achieve it, and did this change make us better?”

THE LATEST FROM BENCH

Keep learning with us.

GET LATEST UPDATES