ALL ARTICLES

Your Agent Finished. Did the Job Get Done?

How business context, expert judgment and production feedback give AI agent evaluation a useful standard.

A buyer asks your sales assistant whether your product fits their business. The agent gives a confident answer, qualifies the lead, and records a next step in the CRM. Every call succeeds.

Did it help the buyer make a good decision?

Perhaps it understood the requirement. Perhaps it promised an integration you do not offer and booked a meeting someone will spend apologizing for. The same completed workflow can contain either outcome.

AI agent evaluation starts with understanding that difference. At BENCH, our starting point is the business and the people using its AI. Before choosing a better setup, discover what a better outcome would mean.

01 / Discover what the business is promising

Before launch, specify the business context: who the customers are, what they buy, how the service delivers value, and which rules govern it. For this illustrative workflow, sit with the people who sell, onboard and support the product. Follow an apparently successful lead into the next stage of work.

The best outcome might be a suitable buyer reaching the right person with accurate expectations. A useful alternative might be explaining that the product does not fit. An uncertain requirement might need clarification.

Then describe the worst credible outcomes: an unsupported commitment, information from the wrong account, or an existing customer sent through a new-business funnel. Ask who bears the cost of each mistake. Meeting count alone misses these distinctions.

A first task specification should capture:

  • The buyer’s need and the evidence available to the agent.
  • What the product can deliver and what the agent may promise.
  • Acceptable outcomes, including clarification and a useful handoff.
  • Required record changes and failures that must block a pass.

Ask a domain expert to judge concrete examples and explain borderline decisions. Those judgments can reveal a requirement that nobody thought to write down. The EvalGen study observed nine practitioners refining criteria as they reviewed outputs, a useful research parallel for this discovery process. [5]

DEFINE WHAT COUNTS AS DONE

TASK ACCEPTANCE

Right accountMet
Qualification supportedMet
Next step recordedMissing

One missing condition. Still unfinished.

01 / Illustrative task checklist. Finishing the run does not establish that every business condition was met.

Customer context gives “good” and “bad” their meaning. Record changes to that meaning so improvements in the score remain interpretable.

02 / Let production challenge the definition

Use the initial specification to build cases before launch. Generated variations should exercise different business conditions, including cases where the right outcome changes. Check their expected answers. Rephrasing the same easy request hundreds of times adds little coverage.

Once the assistant is live, production traces, feedback and resulting records challenge and extend that initial standard. Look for failures, unusually good interactions and routine requests. OpenAI recommends combining domain expertise with historical and production examples. [2]

A trace supplies evidence of behavior. Someone still needs to establish whether that behavior served the customer. A positive reaction can accompany an incorrect promise; silence can mean the buyer gave up.

Turn reviewed examples into cases with known starting conditions and expected outcomes. Include two companies with similar names, an existing customer using an old domain, and a buyer whose required integration is unavailable.

Keep good examples too. If the assistant asked exactly the clarification needed to prevent a bad commitment, preserve that behavior when testing the next change.

Separate routine traffic samples from deliberately difficult cases. A suite assembled from the worst incidents is useful for finding weaknesses. Its pass rate does not estimate how often ordinary customer requests succeed.

03 / Check the outcome and inspect the path

Start with checks that have definite answers. Was the note attached to the intended account? Did the proposed integration exist in the supplied product information? Was the existing owner preserved?

Then assess whether the answer addressed the buyer’s requirement, represented uncertainty and explained the next step. Calibrate automated judgments against examples reviewed by people who understand the business. Anthropic describes combining code, model and human graders. [1]

Keep severe failures separate from writing quality. A polished answer that makes an unauthorized commitment should fail the task. Preserve the individual checks so the team can see what caused that verdict.

The first thing an evaluation may need to improve is the definition of success.

A correctly handled conversation is a measurable task outcome. Increased revenue is a later business outcome influenced by many other factors. An offline score cannot establish that the agent caused more sales.

Inspect the trace to form a hypothesis about a failure: outdated retrieval, an instruction to always book a meeting, or a missing identity check. A score can identify a problem without identifying its cause.

04 / Turn the diagnosis into an experiment

Suppose the assistant makes unsupported promises because its product lookup omits availability constraints. Test a lookup that includes them. Keep the other components fixed so the result can speak to that hypothesis.

Model, prompt, context, tool selection and architecture all remain candidates for change. The agent harness is the application code that executes tools and decides whether to continue or stop. Berkeley’s work on compound AI systems provides a useful foundation for considering interacting components together. [6]

Record the setup, cases, rubric and grader version for each experiment. A model leaderboard can help shortlist candidates. It cannot establish which complete setup meets this buyer’s requirements.

CHANGE ONE FACTOR AT A TIME

BASELINECANDIDATETest casesSame setSame setPromptSame textSame textToolsSame toolsSame toolsModelModel AModel B

Compare outcomes, failures, latency and cost.

02 / An illustrative model comparison. Freeze the cases, instructions and tools so the changed factor is clear.

Reconstruct the starting state in an isolated test environment. Reset it between candidates so one cannot inherit another’s changes. Test fixtures provide an established mechanism for setup and cleanup. [4]

Repeat cases when behavior varies and retain every attempt. Changes can interact: a prompt that helps one model may hurt another. Test a combined candidate before shipping it, and distinguish evidence about that combination from evidence about any one component.

05 / Learn from a case without overstating it

A production failure used to design a fix is now a development case. Keeping it as a regression test is useful. Calling it unseen evidence afterward would misrepresent what the experiment established.

Evaluation contamination can happen through repeated prompt edits and candidate selection, even without training model weights. Established evaluation practice separates development from final testing because repeated access can make scores overly optimistic. [3]

Reserve fresh cases for a later assessment. Variants generated from the same seed can share its assumptions and mistakes; putting them in different splits does not make them independent. Look for overlap across seeds, conversations and accounts. If a reserved case guides a fix, record that change in status and replenish the evidence.

Freeze a dataset version within each comparison; expand it between experiments. When the rubric changes, assess both candidates under the revised rubric. A higher score on an easier definition of success is not a system improvement.

06 / Take the experiment back to production

Define the quality conditions for release, then compare latency and cost per successful task among eligible candidates. A change that fixes one severe failure while introducing another is unfinished work.

QUALITY SETS THE CONDITIONS

Meets the
task requirements?

NO

Investigate failures

YES

Compare cost + latency
03 / Quality checks determine which candidates can proceed. Compare their cost and latency, then verify the selected change in production.

With the criteria established, the experiment cycle can be automated: propose a fix, compare it with the baseline, and deploy changes that satisfy the team’s release policy. Define acceptance and regression checks and a way to stop or roll back. The rubric itself should not quietly change to let a candidate pass.

After release, inspect actual outcomes again. A controlled comparison on live traffic, when volume permits, provides stronger evidence of impact than comparing two different weeks of requests. Monitor failures as well as gains. [1]

A failed experiment can still improve the next decision. Keep its evidence. Add newly understood failures and good examples to the next test version, while retaining fresh cases for assessment. The aim is improvement over time; individual releases can still regress.

At BENCH, this is the decision we want evaluation to support: which change helps this system do this business’s work? Our public comparisons examine specific tasks; the sales scenario here is illustrative. The tool calling guide follows the effects of actions in more detail.

What did the last real customer interaction teach you, and which experiment will test what to change?

THE LATEST FROM BENCH

Keep learning with us.

GET LATEST UPDATES