BENCH COMPARISONS

Which AI actually
does your work?

BENCH takes a real piece of business work, runs it through every model and every harness worth trying, and publishes what came out. Which setup actually does the job, what each one gets wrong, and the cheapest way to keep it done.

LEGAL2026-10-07
Handled right
448
Got it wrong
59
Pass rate
88%
CAUGHT 543INVENTED 46

Best from each maker · 13 models tested

  1. Claude Opus 4.5$0.01939/39
  2. GPT-OSS 120B$0.00138/39
  3. Qwen3 235B$0.00138/39
  4. Gemini 2.5 Flash$0.00438/39
  5. DeepSeek V3$0.00121/39
VIEW RESULTS
CUSTOMER SUPPORTCOMING SOON

Refund and policy replies

Can an agent apply your refund policy without promising money you do not owe?

ACCOUNTINGCOMING SOON

Invoice and expense checks

Can it catch the invoice that breaks your approval limits, without flagging the ones that do not?

SALESCOMING SOON

Outbound and discount rules

Can it quote and discount inside the limits your team actually works to?

How BENCH tests, whatever your system does.

A contract reviewer, a support agent, a research assistant: the work changes, the process does not. It starts at your codebase and does not stop once you ship.

  1. YOUR SYSTEMYOUR AGENT, LIVELEGAL · SUPPORT · ANY WORKFLOWWHAT WE FOUND INSIDEPROMPTS14TOOLS THEY CALL9HARNESS SETTINGS6

    STEP01

    BENCH plugs into your system.

    Legal review, a support agent, anything you run. BENCH reads it and lists every prompt, every tool it calls and every harness setting around them.

  2. WHAT GOOD MEANSWRITTEN DOWN WITH YOU, BEFORE ANYTHING RUNSYOUR OWN RULESGETS THE FACTS RIGHTSTAYS ON TONEREFUSES WHEN IT SHOULDYOURS+ THE ONES EVERY SYSTEM NEEDS

    STEP02

    BENCH writes down what good means.

    Your own rules, in your words, next to the ones every system needs: right facts, right tone, knowing when to refuse.

  3. THE CASESREAL WORK, PLUS THE EDGES NOBODY WRITES DOWNEVERYDAYRIGHT ON THE LINE

    STEP03

    BENCH builds the cases.

    Real work from your business, plus the hardest ones: cases sitting exactly on the line, or missing it by a day. Those never make it into a spec.

  4. RUN IT FOR REALEVERY CASEYOUR SYSTEMPROMPTHARNESSTOOLSMODELNOT A MOCKONE JUDGE

    STEP04

    BENCH runs the real thing.

    Your live system on every case, harness and all, not a mock of it. One judge scores every answer the same way.

  5. TRY THE FIXESSAME CASES, ONE CHANGE AT A TIMEREWORD THE PROMPTKEEP THIS ONESWAP THE MODELCHANGE THE HARNESSLEAVE IT ALONE

    STEP05

    BENCH tries the fixes.

    Reword the prompt, change the harness and the tools it can call, swap the model. One change at a time, scored on the same cases.

  6. QUALITY vs COSTQUALITYCOST PER RUNPICKGOOD ENOUGH, CHEAP ENOUGH, EVERY TIME

    STEP06

    BENCH weighs quality against cost.

    Public benchmarks only narrow the shortlist. What decides it is what each option gets right on your work, and what it costs to run.

  7. IN PRODUCTIONRUNNING CONTINUOUSLYEVERY REAL ANSWER YOUR SYSTEM GIVES, WATCHEDNEW FAILURE, NEVER TESTED FORADDED TO YOUR CASESA GOOD ANSWER WORTH KEEPINGSAVED AS THE EXAMPLE TO BEATBACK TO STEP 01, EVERY TIME

    STEP07

    BENCH keeps going after you ship.

    Testing once tells you about one day. BENCH watches what your system actually answers in production, catches failures nobody wrote a test for, and adds them to your cases so the next run covers them. The good answers get kept too, as the examples that hold the bar.

How a cheaper option catches up.

A cheap model that scores badly is rarely making many different mistakes. It is usually making one, over and over. These are the four things BENCH changes to close that gap, in the order it reaches for them.

  1. 01

    The prompt

    Most failures are one sentence that reads two ways. BENCH finds the sentence, rewrites it, and scores the rewrite on the same cases to prove it helped.

    Usually the cheapest fix there is.

  2. 02

    The harness and its tools

    Everything wrapped around the model: what it can look up, how much context it is given, how many turns it gets, what it does when a tool fails or times out.

    Fixes the failures no prompt can reach.

  3. 03

    The model

    Public leaderboards narrow the field. Your cases decide it. BENCH runs the shortlist on your work and shows which one is actually better at it.

    Often not the model you expected.

  4. 04

    The cost

    Every score comes with what it costs to run and how long it takes. A model that ties the leader at a sixth of the price is a result, not a footnote.

    Quality you can afford to keep running.

YOUR AI SYSTEM. YOUR STANDARD.

WHAT WOULD
YOUR AI MISS?

BENCH YOUR AI SYSTEM (opens calendar in a new tab)