All comparisons

LEGAL · EMPLOYMENT CONTRACT REVIEW

Which model can
check your contracts?

A company wrote down the six rules every employment contract it signs has to meet. BENCH drafted 39 contracts against those rules and put 13 models through them.

They find the problems. Then they add some.

Spotting a real breach turned out to be the easy part. What separates one model from another is how often it flags a clause that was already fine.

Real problems found

543 of 559

Across every model. Spotting a breach is the part they are good at, including the ones buried in an annex that overrides the clause above it.

Problems invented

46

Clauses that already met the rule, reported as breaches anyway. This is where the models actually differ, and every one costs somebody a re-read.

What fixes it

One sentence

Claude Haiku 4.5 reads the non-compete rule as stricter than it is written. That one misreading is all 13 of its wrong calls. Spell the rule out and 13 more contracts come out right.

Every model BENCH tried, on the same contracts.

Nothing changes between these cards except the model. A contract counts only when all six rules were called correctly on it, so a model that finds the breach and invents a second one does not score.

TOP SCORE

Claude Opus 4.5

100% CORRECT

39 of 39 contracts handled correctly

Rule checks correct
234 of 234
Breaches invented
0
Cost per contract
$0.019
Time per contract
10s
BEST VALUE

GPT-OSS 120B

97% CORRECT

38 of 39 contracts handled correctly

Rule checks correct
233 of 234
Breaches invented
0
Cost per contract
$0.001
Time per contract
8s

Qwen3 235B

97% CORRECT

38 of 39 contracts handled correctly

Rule checks correct
233 of 234
Breaches invented
0
Cost per contract
$0.001
Time per contract
8s

Gemini 2.5 Flash

97% CORRECT

38 of 39 contracts handled correctly

Rule checks correct
233 of 234
Breaches invented
0
Cost per contract
$0.004
Time per contract
7s

Gemini 2.5 Pro

97% CORRECT

38 of 39 contracts handled correctly

Rule checks correct
233 of 234
Breaches invented
1
Cost per contract
$0.020
Time per contract
22s

GPT-5.6 Terra

95% CORRECT

37 of 39 contracts handled correctly

Rule checks correct
232 of 234
Breaches invented
1
Cost per contract
$0.004
Time per contract
4s

Gemini 3.5 Flash

95% CORRECT

37 of 39 contracts handled correctly

Rule checks correct
232 of 234
Breaches invented
2
Cost per contract
$0.004
Time per contract
11s

Claude Sonnet 5

95% CORRECT

37 of 39 contracts handled correctly

Rule checks correct
232 of 234
Breaches invented
2
Cost per contract
$0.018
Time per contract
15s

Claude Opus 5

95% CORRECT

37 of 39 contracts handled correctly

Rule checks correct
232 of 234
Breaches invented
2
Cost per contract
$0.036
Time per contract
16s

GPT-4.1 mini

85% CORRECT

33 of 39 contracts handled correctly

Rule checks correct
228 of 234
Breaches invented
6
Cost per contract
$0.001
Time per contract
4s

GPT-5.4 mini

74% CORRECT

29 of 39 contracts handled correctly

Rule checks correct
224 of 234
Breaches invented
3
Cost per contract
$0.002
Time per contract
3s

Claude Haiku 4.5

67% CORRECT

26 of 39 contracts handled correctly

Rule checks correct
221 of 234
Breaches invented
12
Cost per contract
$0.003
Time per contract
4s

DeepSeek V3

54% CORRECT

21 of 39 contracts handled correctly

Rule checks correct
213 of 234
Breaches invented
17
Cost per contract
$0.001
Time per contract
15s

Which one to run, and what it costs.

Two questions, in this order. Can you get the outcome at all, and what is the least you can pay for it? Best value is the cheapest model within five points of the top score. The top scorer is marked separately so you can see what the extra money buys.

BEST VALUE

GPT-OSS 120B

Contracts right
38 of 39
Breaches invented
0
Per 1,000 contracts
$1

97% correct against 100% for Claude Opus 4.5, the top scorer, at 32 times the price. That is 1 more contract out of 39 for $18 more per thousand. Best value is the cheapest model still doing the job, not the highest number on the page.

ModelCorrectWhat it costs to runPer 1,000 contracts
Claude Opus 4.5100%$0.019 eachPer 1,000 contracts$19TOP SCORE
GPT-OSS 120B97%$0.001 eachPer 1,000 contracts$1BEST VALUE
Qwen3 235B97%$0.001 eachPer 1,000 contracts$1
Gemini 2.5 Flash97%$0.004 eachPer 1,000 contracts$4
Gemini 2.5 Pro97%$0.020 eachPer 1,000 contracts$20
GPT-5.6 Terra95%$0.004 eachPer 1,000 contracts$4
Gemini 3.5 Flash95%$0.004 eachPer 1,000 contracts$4
Claude Sonnet 595%$0.018 eachPer 1,000 contracts$18
Claude Opus 595%$0.036 eachPer 1,000 contracts$36
GPT-4.1 mini85%$0.001 eachPer 1,000 contracts$1
GPT-5.4 mini74%$0.002 eachPer 1,000 contracts$2
Claude Haiku 4.567%$0.003 eachPer 1,000 contracts$3
DeepSeek V354%$0.001 eachPer 1,000 contracts$1

How BENCH makes a cheap model good enough.

This is the part a leaderboard cannot do for you. A cheap model that looks hopeless on score usually is not making many different mistakes. It is making one mistake repeatedly, and that is something you can fix.

Where it starts

97%

GPT-OSS 120B is the cheapest model here at $1 per thousand contracts, and on raw score it looks unusable.

What is actually wrong

1 rule

Not 1 separate problems. One misreading of the non-compete rule, repeated on every contract where it applies.

What BENCH changes

1 sentence

Spell out in the prompt what the rule does and does not require, then score the change on the same 39 contracts.

Where it lands

100%

The same outcome as Claude Opus 4.5 at 100%, for 32 times less. Projected from the verdicts already recorded, and BENCH re-runs it to prove it.

One rule carries almost every failure.

It is not only the cheap one. Each model here has a single rule it reads as stricter than it is written, and that one rule accounts for nearly everything it gets wrong.

Claude Haiku 4.5$0.003 per case

What goes wrong

All 13 of its wrong calls are on one rule: Non-compete. It reads that rule as stricter than it is written.

RIGHT 26WRONG 13

What BENCH would fix

Spell out in the prompt what the non-compete rule actually requires, then run the same 39 contracts again.

67%100%+13 contracts

A projection, not a second run: every one of the 13 contracts it got wrong failed only on that rule. BENCH proves it by running them again.

DeepSeek V3$0.001 per case

What goes wrong

16 of its 21 wrong calls are on one rule: Non-compete. It reads that rule as stricter than it is written.

RIGHT 21WRONG 18

What BENCH would fix

Spell out in the prompt what the non-compete rule actually requires, then run the same 39 contracts again.

54%87%+13 contracts

A projection, not a second run: 13 of the 18 contracts it got wrong failed only on that rule. BENCH proves it by running them again.

GPT-4.1 mini$0.001 per case

What goes wrong

5 of its 6 wrong calls are on one rule: Non-compete. It reads that rule as stricter than it is written.

RIGHT 33WRONG 6

What BENCH would fix

Spell out in the prompt what the non-compete rule actually requires, then run the same 39 contracts again.

85%97%+5 contracts

A projection, not a second run: 5 of the 6 contracts it got wrong failed only on that rule. BENCH proves it by running them again.

The same pattern holds for 9 more models: one rule, misread the same way, accounting for most of what it gets wrong.

Every wrong call, in plain words.

Grouped by model, with repeats counted rather than listed. This is the evidence behind the numbers above.

GPT-OSS 120B1 wrong call
Got it wrongNon-compete

The 13-month duration is a real breach and is correctly flagged. But the review also treats the 50% pay as a breach, claiming the non-compete rule requires full salary, while the non-compete rule only requires 'compensation for that period'. That invents a payment breach, and the replacement wording imposes full pay.

Qwen3 235B1 wrong call
Got it wrongInvention ownership

Clause 7 only claims outside-hours inventions that 'could reasonably be of interest to Northstar's business', so unrelated inventions stay with the employee and the draft complies with the invention ownership rule. The review invents a breach and proposes replacement wording, and it also contradicts its own opening line that there are no breaches.

Gemini 2.5 Flash1 wrong call
Got it wrongInvention ownership

Clause 7 claims only inventions made in employment or of reasonable interest to the business, and so does not take unrelated out-of-hours inventions. The review nonetheless reports a Clause 7 breach, saying a positive statement of employee ownership is missing. That is an invented breach in a compliant clause.

Gemini 2.5 Pro1 wrong call
Made it upNotice period

Clause 3 'exactly 3 months' written notice' satisfies the 'at least 3 months' rule, so the draft is compliant. The review invents a breach ('exactly' is more restrictive) and proposes unnecessary replacement wording.

GPT-5.6 Terra2 wrong calls
Made it upNotice period

The draft's 'ninety days' written notice appears intended as a paraphrase of three months in an otherwise compliant draft, and the review flagged it as a breach. Ninety days can fall short of three months, so this is a judgment call, and I resolved it against the review.

Got it wrongProbation length

The 9-month probation is a breach and the response eventually says so ('it sets probation at 9 months, exceeding the 6-month maximum'). It first contradicts itself ('All other reviewed clauses comply'), and it gives no replacement wording, which the criterion requires.

Gemini 3.5 Flash2 wrong calls
Made it upNotice period2×

The review flags 13 weeks' notice as a breach of the 3-month rule. 13 weeks (91 days) is essentially 3 months, and the draft has no annex changing it, so this reads as a false positive. The 'less than three months in most parts of the year' reasoning is also dubious.

Claude Sonnet 52 wrong calls
Made it upNotice period

The body and table correctly mark Clause 3 (3 months' notice) as the notice period rule-compliant. The closing line then says the 9-month probation 'violates the notice period rule the probation length rule', which wrongly attributes a violation to the notice period rule and contradicts the earlier finding, so it reports a notice problem in a compliant draft.

Made it upNon-compete

The 6-month non-compete is paid EUR 1/month, which literally satisfies the non-compete rule (12 months or less, with compensation paid). The response adds a 'genuine compensation' test and says the rule 'explicitly treats' token pay as a breach, which the playbook does not say. It therefore invents a breach in a draft that satisfies the rule as written.

Claude Opus 52 wrong calls
Made it upNotice period2×

The review says clause 3 'Complies with the notice period rule in substance', but then flags 'ninety days' as possibly 'marginally short of the notice period rule' and proposes replacement wording. That reports a notice problem in a draft the playbook most likely treats as compliant, a false positive on a paraphrase of three months.

GPT-4.1 mini6 wrong calls
Made it upNon-compete5×

Clause 6 is 12 months and pays 50% of salary, which is compensation and so satisfies the non-compete rule. The review invents a breach, saying 'unpaid for half the time', and rewrites the clause to 100% pay.

Made it upNotice period

The review says "Ninety days equals three months", but 90 days is generally shorter than 3 months (most 3-month spans are 91–92 days). Clause 3 therefore likely breaches the notice period rule, and the review missed it and offered no corrected wording.

GPT-5.4 mini10 wrong calls
Got it wrongInvention ownership3×

Clause 7 assigns all inventions to Northstar 'whether or not they were made during working hours and whether or not they relate to Northstar's business', which breaches the invention ownership rule. The response says there are no breaches and offers no replacement wording.

Made it upNon-compete2×

The response opens with 'I found one breach' and says the clause 'breaches the non-compete rule', then says it complies. This contradiction reports a non-compete problem in a compliant draft and leaves the finding unclear.

Made it upOvertime

Clause 4 is compliant, but the response opens with 'I found one breach' and a 'Breach: Clause 4' / 'Rule broken: the overtime rule' heading before retracting it. This reports an overtime problem in a compliant draft and contradicts itself.

Got it wrongNon-compete

Clause 6 sets a 13-month non-compete, which exceeds the 12-month limit in the non-compete rule. The review says the draft is compliant and gives no correction, so it misses the breach.

Got it wrongOvertime

Clause 4 says 'The first 10 hours of overtime each month are covered by the salary'. The overtime rule treats wording that makes overtime covered by salary as a breach, but the review calls it compliant, reads the rule as breached only when all overtime is covered, and offers no replacement wording.

Got it wrongHoliday

Clause 5 gives only 20 days of holiday, below the 28-day minimum. The review said 'I do not see any other breaches', so it missed this breach and proposed no correction.

Got it wrongProbation length

Clause 2 sets a 9-month probation, which breaches the probation length rule's 6-month limit. The review says 'I do not see any other breaches' and misses it, with no corrective wording.

Claude Haiku 4.513 wrong calls
Made it upNon-compete12×

The 12-month non-compete is paid at 50%, which satisfies the non-compete rule. The review wrongly reports a breach and invents a full-salary requirement.

Got it wrongNon-compete

Breach 1 (13 months over the 12-month limit) is correct. Breach 2 is invented: the non-compete rule only requires that compensation be paid, and the clause pays 50% of salary. The review claims the non-compete rule 'requires full compensation', which the playbook does not say, and proposes unnecessary replacement wording.

DeepSeek V321 wrong calls
Made it upNon-compete14×

The non-compete is 12 months and paid (50%), so it is compliant under the non-compete rule; the review wrongly reports a breach and proposes 100% pay.

Got it wrongNon-compete2×

The review calls the 24-month clause 'compliant' while admitting it exceeds the 12-month maximum. It relies on an annex with 100% pay that does not appear in the supplied document, so the the non-compete rule conclusion is unsupported and self-contradictory.

Made it upNotice period2×

Flagged clause 3 'thirteen weeks' as breaching the notice period rule with an incorrect rationale ('3 months and 1 week'); thirteen weeks is about 3 months, so this invents a breach in a compliant draft.

Made it upInvention ownership

Clause 7 already says inventions made outside working hours and unrelated to the business remain the employee's property, so the invention ownership rule is satisfied. The review invents a breach and ignores that second sentence.

Got it wrongInvention ownership

Clause 7 only claims out-of-hours inventions that 'could reasonably be of interest to Northstar's business', so unrelated inventions stay with the employee and the draft satisfies the invention ownership rule. The review nonetheless reports a Breach 1 and proposes replacement wording, which invents a breach in a compliant draft.

Got it wrongHoliday

The supplied contract has 20 days of holiday and no annex text. The review declares the holiday rule compliant on the strength of an annex text that is not in the provided document, so the compliance finding is unsupported.

The harness BENCH ran.

A model is only part of what you ship. What you actually run is a model, a prompt and a harness around it: what it can see, what it can call, how many turns it gets, what happens when something fails. Every model here got this one, unchanged.

The instruction every model was given, word for word

Review the supplied contract against the supplied company rules. Report every breach with a clause reference, explain why it breaches the rule, and propose replacement wording. If the contract complies, say so. Do not give legal advice or invent requirements.

  • ShapeSingle pass. One model call per contract, no agent loop, no follow-up question, no second look at its own answer. The simplest harness anyone would build first.
  • What goes inTwo messages. The instruction above, then the company's six rules and the contract text together in one payload. Nothing else reaches the model: no case law, no past decisions, no company wiki.
  • ToolsNone wired up. No search, no retrieval, no document store. Every answer comes from what was handed over in that one request.
  • Output limit4,096 tokens. An empty answer aborts the run rather than being scored as a failure, because a truncated reply is a harness bug and not the model's.
  • SamplingProvider defaults, no temperature tuning, one review per contract. No best-of, no retries, no voting between runs.
  • Transport180 second timeout, transport retries off, and a cost ceiling reserved before every call so a runaway run stops instead of spending.
  • Held constantEvery item above is identical on every row. The only thing that changes between models is which model id is called.

Where the contracts came from.

Written for this test, not taken from anyone. A benchmark is only as good as its documents, so here is how these were made and what is in them.

  • Where they come fromWritten for this test against a fictional employer, Northstar Software. No client repository, no real employment agreement and no identifiable person is involved.
  • How they were builtOne compliant base contract, then 38 variations of it. Changing one clause at a time means a difference in a verdict traces to that clause rather than to a different writing style.
  • What is in the set12 contracts fully comply, including ones sitting exactly on the limit at six months' probation or 28 days' holiday. The other 27 carry at least one planted breach, up to one contract that breaks all six rules.
  • The hard onesBreaches split between a numbered clause and an annex that replaces it, so neither half reads wrong on its own. The rules tell the model that annexes govern, so this is a reading test rather than a trick.
  • Held out30 of the 39 were never looked at while the setup was being written, so the instruction could not be tuned against them.

How it was scored: one judge, rule by rule.

A score is only worth something if the same thing scored everybody. This is what did the scoring and what it did when it went wrong.

  • The judgeOne fixed model scores every review on every row. Holding it still is what makes a row difference the model rather than the scoring.
  • How it scoresTwo passes per review: general answer quality, then the six company rules one at a time. Rule by rule is what turns a score into a sentence you can go and fix.
  • When the judge failsAn unparseable judge reply is retried once, then recorded as that single case rather than throwing away a finished suite.
  • What counts as rightA contract scores only when all six rules were called correctly. Finding the real breach and inventing a second one does not count, because somebody still has to re-read it.
  • CostProjected from the tokens each run reported, at published list prices. It is not a provider invoice, and it covers the review rather than the judging.

What BENCH would change next.

None of these has been run on this suite yet. Each would be scored on the same frozen contracts before anybody claims it worked.

  1. PROMPTNOT RUN YET

    Say what the rule does not require.

    The dominant failure is a model reading “the employee must be paid for it” as “must be paid in full”, so compliant contracts paying half get flagged. One added sentence should remove it.

    Would address: almost every invented breach in the table

  2. HARNESSNOT RUN YET

    Ask one rule at a time.

    Today one call covers all six rules at once. Splitting it into six smaller checks costs more calls but gives each rule the model's full attention, and makes a wrong call easier to localise.

    Would address: rules missed when a breach hides in an annex

  3. TOOLSNOT RUN YET

    Let it look up past decisions.

    Give the review a lookup over contracts the company has already approved, so a borderline clause can be checked against how the same clause was treated before rather than judged cold.

    Would address: contracts sitting exactly on the limit

The rules, written down before anything ran.

  • Notice periodAt least three months' notice once probation has ended.
  • Probation lengthProbation may not run longer than six months.
  • Non-competeTwelve months at most, and the employee must be paid for it.
  • OvertimePaid, or given back as time off. Never just covered by salary.
  • HolidayAt least 28 paid days a year, on top of public holidays.
  • Invention ownershipWork an employee does in their own time, unrelated to the business, stays theirs.

Methodology: how this was done.

  • The rules came firstThe six rules were written and frozen before any contract was drafted, so nothing could be tuned to a result.
  • 39 contracts, hardest includedTwelve sit exactly on the limit: probation at precisely six months, holiday at exactly 28 days, a non-compete paying a token amount that still satisfies “must be paid”. The hardest split a breach across a clause and an annex, so neither reads wrong on its own.
  • Same work, one judgeEvery model got the identical contract, rules and instructions. One judge scored all 507 reviews rule by rule, giving 3042 separate verdicts.
  • What counts as rightA contract scores only when all six rules were called correctly. Finding a real breach and inventing a second one does not count, because somebody still has to re-read it.
  • Cost and timingCost is projected from reported tokens at list prices, not an invoice, and covers the review rather than the judging. Timings are wall-clock on a single run.
  • Projections are labelledWhere the page shows what a fix would be worth, it is arithmetic over verdicts already recorded, never a second run, and it says so on the card.
  • No client dataEvery contract was drafted for this test. No real agreement and no identifiable person appears in the suite.

YOUR AI SYSTEM. YOUR STANDARD.

WHAT WOULD
YOUR AI MISS?

BENCH YOUR AI SYSTEM (opens calendar in a new tab)