ALL ARTICLES

The Tool Call Worked. The Task Didn’t.

Why tool calling evaluation starts with the customer’s intended change and follows its consequences into production.

The tool call is valid. The request succeeds. The agent updates the wrong customer record.

An argument check might pass. An API status check might pass too. The customer now has to discover the mistake and get someone to repair it.

Tool calling evaluation needs to follow that consequence. A valid request is one part of the evidence. The requested change, the affected account and the customer’s next step determine whether the task worked.

Consider an account assistant that lets customers maintain their company’s contact details through conversation. We will use an illustrative job-title update throughout. Every tool action should be evaluated against the business promise it is supposed to fulfill.

01 / Discover what the customer meant to change

“Update Alex’s role” sounds simple until you ask what role means. A job title? A billing contact? Permission to administer the account? Those changes have different consequences.

Work through examples with customers and the team maintaining the records. In our example, the request concerns a job title. The assistant must identify the intended Alex within the account the customer is authorized to manage.

An accurate update and a verifiable confirmation fulfill the request. Modifying another company’s contact or changing permissions would create a different problem. Company boundaries must be enforced by the application, including its read tools.

Write those distinctions into the test. Specify the intended record and new title, the records that must remain untouched, and when the assistant should ask a clarifying question.

SAME NAME. DIFFERENT RECORDS.

Update Alex’s job title

Alex

Other company

Wrong account

Alex

Authorized account

Context matches

Check identity before changing state.

01 / Illustrative identity check. A valid identifier is insufficient: the intended person and authorized account must both match.

This approach has a research precedent. The τ-bench benchmark evaluates tool-using agents partly by comparing the database state after an interaction with the intended goal state. That checks whether the environment reflects the requested result. [2]

The verifier should inspect resulting records independently and check the confirmation sent to the customer. A correct write followed by the wrong confirmation can still leave the customer acting on false information.

02 / Separate valid arguments from correct arguments

A function schema can require a contact identifier and a new job title. It can constrain their types and reject unexpected fields. OpenAI’s strict function-calling mode is designed to make calls adhere to the supplied schema. [1]

That is useful. But a well-formed identifier can refer to the wrong Alex. A title can be a valid string and still misrepresent the customer’s request.

For function calling evaluation, we would check several things separately:

  • Was the selected tool appropriate for the request?
  • Did the arguments satisfy the tool’s schema?
  • Did the identifier and field values match the task evidence?
  • Was the operation authorized, and did it produce the intended change?

These checks expose different fixes. Invalid fields might point to schema design. The wrong identity might point to missing context or a poor search result. A rejected write might reveal an access-control problem. Changing the model before locating the failure is a guess.

03 / Make room for a useful pause

If the business measures only how often the assistant completes an update without help, it can reward the wrong behavior. A clarification adds a turn. It may also prevent an incorrect write and a much longer support exchange.

Add a question about the contact’s current title, a request that does not identify which Alex, and a request from someone who can read the record but cannot edit it. Include an update that has already been completed.

In each case, specify the acceptable behavior and verify the absence of unintended writes. Asking for clarification, answering from an existing record or explaining an access restriction may complete the task.

The cost of a wrong action continues after the tool call ends.

Authorization belongs in application enforcement as well as evaluation. A test can reveal that an agent attempted a forbidden operation; the tool implementation should still prevent it. Passing a collection of examples is not a substitute for that control.

Preserve examples of good restraint as well as failures. A change that writes more often can look more capable while making the customer’s experience worse.

04 / Follow the call through to its effects

Now suppose the CRM accepts the update, but the connection closes before the agent receives a response. The agent retries. Whether that is safe depends on the operation and how retries are handled.

Writing the same title twice might leave the field unchanged while still producing duplicate activity entries or notifications. A customer who receives conflicting messages can reasonably doubt that the task finished.

A LOST RESPONSE LEAVES AN OPEN QUESTION

1Write accepted
2Reply lost
?Retry safely?

INSPECT WHAT ACTUALLY HAPPENED

RecordActivity logSide effects
02 / A timeout does not establish whether a write happened. Inspect the record, activity log and side effects before choosing a safe recovery.

Stripe’s API documents one concrete retry mechanism: idempotency keys let clients repeat supported requests without creating an additional operation. That behavior depends on the API’s contract; it is not a property every tool automatically has. [4]

For the account assistant, test a timeout before a write, a lost response after a write, and a clear rejection. The expected recovery may differ. Check the resulting state, relevant side effects and the explanation the customer receives.

In production, relate the trace to activity records and later support requests where those links are available. A successful API response does not tell you whether someone had to repair the result afterward.

05 / Turn an incident into a testable change

Suppose a live trace shows the assistant choosing between two contacts without enough context. Read the customer’s request and the tool response together. Did search omit a useful identifier? Did the prompt encourage guessing? Was the request itself ambiguous?

Anthropic’s tool-design guidance recommends meaningful evaluation tasks with verifiable outcomes. It also cautions against requiring one exact tool sequence when several valid ways of completing a task exist. [3]

One experiment could add distinguishing context to results from an account-scoped search. Another could replace a broad “save” tool with a tool that can only update job titles. Compare each with the baseline before considering a model change.

Make the expected improvement explicit: fewer wrong-person selections, with correct updates and useful clarifications preserved. Inspect failures by category so better schema validity cannot conceal worse identity resolution.

That is the BENCH perspective on tool optimization: discover what the customer needs the action to accomplish, then investigate which part of the complete setup prevents it. Prompt, tool choice, tool calling and harness behavior are all experiment candidates.

06 / Close the loop on real customer work

Turn reviewed production examples into a suite covering ordinary updates, ambiguous names, read-only requests, existing changes, forbidden writes and lost responses. OpenAI’s evaluation guidance calls for ongoing testing and new cases informed by application behavior. [5]

Reconstruct starting records in an isolated environment and restore them between attempts. Replaying a customer trace must not resend its writes or notifications to the live account.

CHECK THE CHANGE AND WHAT STAYED THE SAME

BEFORE

Intended Alex

Analyst

Other Alex

Director

AFTER

Intended Alex

Manager

Other Alex

Director

03 / Illustrative record diff. Only the intended job title changes. Restore the isolated test environment before the next attempt.

Use the known incident to develop the fix and fresh cases to assess whether it generalizes. Our agent-evaluation guide explains this separation. A fix that passes its motivating example still has something to prove.

Automation can carry candidate changes through evaluation and into deployment when the team’s release conditions are met. Retain the baseline, record what changed, and define which production failures stop or reverse the rollout.

A shadow run can compare proposed actions on live inputs with writes suppressed. It cannot demonstrate that those writes would complete correctly. Check actual outcomes after a controlled rollout, including corrections and customer follow-up.

A new failure might call for another tool experiment, or reveal that the original definition of success was incomplete. Good examples tell you what the next change must preserve.

Which promise did this action make to the customer, and what evidence shows that it kept it?

THE LATEST FROM BENCH

Keep learning with us.

GET LATEST UPDATES