The Cheapest Model Can Cost You More.
How to test whether a cheaper AI setup saves work or passes the cost to your customers.
You switch a customer-facing AI assistant to a cheaper model. The token bill falls. A week later, support is handling more repeat contacts and customers are explaining the same problem twice.
Did the system become cheaper?
The invoice answers part of that question. Some of the work may have moved to people whose time never appears in the model bill. Some may have moved to the customer.
LLM cost optimization starts by understanding the service being delivered. Define an acceptable customer outcome, then measure what the complete system spends to achieve it. A saving is a hypothesis until the work and its consequences have been counted.
01 / Understand the service behind the bill
Consider an assistant that lets customers change an order through conversation. Before testing models, work with operations and support to understand the promise: which changes are allowed, when fulfillment makes them impossible, and what the customer needs confirmed.
Use concrete examples to establish good and bad outcomes with those domain experts. The best result might be the intended change completed once and clearly confirmed. A useful alternative is a correct explanation of why it cannot happen. Changing the wrong order is a critical failure, however cheaply it was done.
A correct escalation can be a successful agent decision while still leaving human work to finish the customer’s task. Follow that handoff. Where does the cost end up, and who remains responsible for the result?
Choose which boundary you are measuring. For automated completion, count tasks that pass without intervention. For the full business process, include the human work and count the tasks completed afterward. Report both when the distinction matters.
WHAT GOES INTO THE TASK COST?
Model invoice
Full task cost
Include failed attempts and retries.
Within that boundary, divide the total cost of all attempted tasks by the number that meet the acceptance criteria. Keep the spend from failed attempts in the numerator. If no task succeeds, there is no finite cost per successful task to report.
Show completion rate and critical failures alongside cost. Otherwise a system serving only easy requests can appear efficient while leaving much of the workload unresolved. Do not let an average saving conceal a failure the business cannot accept.
02 / Put the correction work back in the calculation
Here is a deliberately simplified example, using invented numbers rather than model prices or BENCH measurements. Two setups each receive the same 1,000 requests. Assume every request is ultimately completed after any necessary correction.
Setup A spends $20 on automated processing. One hundred requests need a human correction. At an assumed $1 per correction, its total is $120, or $0.12 per completed request.
Setup B spends $40 on automated processing. Twenty requests need the same correction effort. Its total is $60, or $0.06 per completed request.
SAME 1,000 COMPLETED TASKS
Without review costs, A wins this comparison. With the stated review assumptions, B wins. Neither conclusion transfers to your workload until you measure its correction rate and effort.
The difference could shrink or reverse if reviewers correct A’s mistakes quickly, if B needs expensive specialist review, or if errors go undetected. Record the assumptions instead of hiding them inside a single efficiency score.
There is another boundary here: neither total prices the customer’s effort. Track repeat contacts, abandonment and time to resolution separately where you can. Assigning an invented dollar value to frustration would make the arithmetic look more complete than the evidence allows.
03 / Follow the work across the handoff
A task may involve several model calls before the answer appears. It may retrieve documents, invoke a paid service, retry after a timeout and ask a second model to check the result.
Attribute those operations to the task that caused them. Distinguish uncached input, cached input and output where pricing differs. Include any separately billed execution or tool usage rather than treating the visible answer as the entire cost.
Keep evaluation costs separate from runtime costs, unless evaluation is part of the live workflow. A development-only judge should not appear as a production call in every request. For the overall investment decision, also account for maintaining datasets, running experiments and reviewing proposed changes.
In production, connect traces to order state and support work where possible. Choose a follow-up window that captures delayed corrections. Mark cases still awaiting resolution as pending; the absence of a complaint is not verified completion.
Track typical completion time and slower cases. A cheap answer can still keep a customer waiting for a tool, a second conversation or a person. These are useful operational measures even when you cannot reliably convert them into money.
A saving that moves unfinished work to the customer needs a different name.
04 / Make each saving an experiment
Inspect a few complete executions. Does the assistant retrieve the same order twice? Generate a long explanation that no one reads? Ask a model to perform a calculation the application could compute directly?
OpenAI’s latency guidance identifies fewer requests, shorter outputs and avoiding unnecessary LLM work as useful optimization directions. These are changes to investigate, with quality checked alongside their operational effect. [2]
Write a hypothesis before changing the workflow: “Reusing the verified order lookup will remove a call without changing which requests complete correctly.” Compare the baseline and candidate on the same case and rubric versions. Include difficult cases and successful behavior that must survive.
Prompt caching is another possible saving when requests share reusable content. Its behavior depends on the provider, model and configuration. OpenAI documents matching prompt prefixes and eligible cache boundaries; shared text alone is not a guarantee that a particular request will receive a cache hit. [3]
Measure usage and charges against the chosen quality conditions. A shorter prompt can remove useful instructions. Cached content can be wrong. Each saving has to survive a check of what customers actually receive.
05 / Compare model, prompt and routing choices
Now test whether a different model can handle the same work. Keep the criteria fixed. A candidate that saves money by dropping required checks is performing a different service. A critical failure remains disqualifying even when it is rare in the sample.
A prompt change may help a smaller model follow the workflow. A tool response with less irrelevant information may reduce both input size and confusion. A simpler harness may eliminate a loop that rarely changes the decision.
Routing is another option: direct straightforward cases to one setup and difficult cases to another. Anthropic describes this pattern while also recommending that teams add complexity only when the task warrants it. [1]
The router becomes part of the experiment. How often does it send a difficult order change down the cheap path? What do classification, fallback and repeated processing cost? Its savings need to survive those additions.
THE ROUTER IS PART OF THE BILL
ROUTINE
COMPLEX
Measure the whole boundary, including fallback.
Keep routine and exception cases visible. A cheaper average can result from receiving easier requests, even if neither setup improved. Compare equivalent workloads and investigate changes in the mix before claiming a saving.
This is why BENCH starts with the business context and quality standard. They determine which costs can be reduced and which apparent savings undermine the service. The agent-evaluation process makes that standard testable across complete setups.
06 / Keep testing the saving after release
Assess the chosen setup on cases that were not used to tune it. Repeated selection against the same examples can make performance look overly optimistic. Generated variants of a familiar case may carry the same assumptions, so provenance matters too. [4]
Define release conditions for task quality, critical errors, latency and cost. An automated improvement pipeline can propose changes, evaluate them and deploy candidates that satisfy the team’s policy. Its objective must include those quality conditions; optimizing the invoice alone would automate the wrong decision.
Then use live production traces and completed customer outcomes to revisit the estimate. Did correction work increase? Did a new kind of order appear? Did the expected cache usage materialize? Keep rollout comparisons controlled where possible, and stop or reverse a change when it fails the release conditions.
Reviewed failures, edge cases and good examples extend the next evaluation suite. OpenAI describes this ongoing combination of application monitoring and evaluation updates. [5] The business context establishes the initial standard; production keeps testing whether the standard and the system remain useful.
With a probabilistic system, a good result on one batch leaves uncertainty about the next. Repeat important comparisons, retain failed experiments, and measure the full operating cost of the improvement process. A cost estimate belongs to a particular workload and setup; it is not a permanent property of the model.
Did the customer’s task become cheaper to complete, or did part of the work simply disappear from your bill?