- Handled right
- 448
- Got it wrong
- 59
- Pass rate
- 88%
Best from each maker · 13 models tested
- Claude Opus 4.5$0.01939/39
- GPT-OSS 120B$0.00138/39
- Qwen3 235B$0.00138/39
- Gemini 2.5 Flash$0.00438/39
- DeepSeek V3$0.00121/39
BENCH COMPARISONS
BENCH takes a real piece of business work, runs it through every model and every harness worth trying, and publishes what came out. Which setup actually does the job, what each one gets wrong, and the cheapest way to keep it done.
Best from each maker · 13 models tested
Can an agent apply your refund policy without promising money you do not owe?
Can it catch the invoice that breaks your approval limits, without flagging the ones that do not?
Can it quote and discount inside the limits your team actually works to?
A contract reviewer, a support agent, a research assistant: the work changes, the process does not. It starts at your codebase and does not stop once you ship.
STEP01
BENCH plugs into your system.Legal review, a support agent, anything you run. BENCH reads it and lists every prompt, every tool it calls and every harness setting around them.
STEP02
BENCH writes down what good means.Your own rules, in your words, next to the ones every system needs: right facts, right tone, knowing when to refuse.
STEP03
BENCH builds the cases.Real work from your business, plus the hardest ones: cases sitting exactly on the line, or missing it by a day. Those never make it into a spec.
STEP04
BENCH runs the real thing.Your live system on every case, harness and all, not a mock of it. One judge scores every answer the same way.
STEP05
BENCH tries the fixes.Reword the prompt, change the harness and the tools it can call, swap the model. One change at a time, scored on the same cases.
STEP06
BENCH weighs quality against cost.Public benchmarks only narrow the shortlist. What decides it is what each option gets right on your work, and what it costs to run.
STEP07
Testing once tells you about one day. BENCH watches what your system actually answers in production, catches failures nobody wrote a test for, and adds them to your cases so the next run covers them. The good answers get kept too, as the examples that hold the bar.
A cheap model that scores badly is rarely making many different mistakes. It is usually making one, over and over. These are the four things BENCH changes to close that gap, in the order it reaches for them.
01
The promptMost failures are one sentence that reads two ways. BENCH finds the sentence, rewrites it, and scores the rewrite on the same cases to prove it helped.
Usually the cheapest fix there is.
02
The harness and its toolsEverything wrapped around the model: what it can look up, how much context it is given, how many turns it gets, what it does when a tool fails or times out.
Fixes the failures no prompt can reach.
03
The modelPublic leaderboards narrow the field. Your cases decide it. BENCH runs the shortlist on your work and shows which one is actually better at it.
Often not the model you expected.
04
The costEvery score comes with what it costs to run and how long it takes. A model that ties the leader at a sixth of the price is a result, not a footnote.
Quality you can afford to keep running.

YOUR AI SYSTEM. YOUR STANDARD.