Claude Opus 4.5
100% CORRECT
39 of 39 contracts handled correctly
- Rule checks correct
- 234 of 234
- Breaches invented
- 0
- Cost per contract
- $0.019
- Time per contract
- 10s
LEGAL · EMPLOYMENT CONTRACT REVIEW
A company wrote down the six rules every employment contract it signs has to meet. BENCH drafted 39 contracts against those rules and put 13 models through them.
Spotting a real breach turned out to be the easy part. What separates one model from another is how often it flags a clause that was already fine.
Real problems found
543 of 559
Across every model. Spotting a breach is the part they are good at, including the ones buried in an annex that overrides the clause above it.
Problems invented
46
Clauses that already met the rule, reported as breaches anyway. This is where the models actually differ, and every one costs somebody a re-read.
What fixes it
One sentence
Claude Haiku 4.5 reads the non-compete rule as stricter than it is written. That one misreading is all 13 of its wrong calls. Spell the rule out and 13 more contracts come out right.
Nothing changes between these cards except the model. A contract counts only when all six rules were called correctly on it, so a model that finds the breach and invents a second one does not score.
100% CORRECT
39 of 39 contracts handled correctly
97% CORRECT
38 of 39 contracts handled correctly
97% CORRECT
38 of 39 contracts handled correctly
97% CORRECT
38 of 39 contracts handled correctly
97% CORRECT
38 of 39 contracts handled correctly
95% CORRECT
37 of 39 contracts handled correctly
95% CORRECT
37 of 39 contracts handled correctly
95% CORRECT
37 of 39 contracts handled correctly
95% CORRECT
37 of 39 contracts handled correctly
85% CORRECT
33 of 39 contracts handled correctly
74% CORRECT
29 of 39 contracts handled correctly
67% CORRECT
26 of 39 contracts handled correctly
54% CORRECT
21 of 39 contracts handled correctly
Two questions, in this order. Can you get the outcome at all, and what is the least you can pay for it? Best value is the cheapest model within five points of the top score. The top scorer is marked separately so you can see what the extra money buys.
BEST VALUE
GPT-OSS 120B
97% correct against 100% for Claude Opus 4.5, the top scorer, at 32 times the price. That is 1 more contract out of 39 for $18 more per thousand. Best value is the cheapest model still doing the job, not the highest number on the page.
This is the part a leaderboard cannot do for you. A cheap model that looks hopeless on score usually is not making many different mistakes. It is making one mistake repeatedly, and that is something you can fix.
Where it starts
97%
GPT-OSS 120B is the cheapest model here at $1 per thousand contracts, and on raw score it looks unusable.
What is actually wrong
1 rule
Not 1 separate problems. One misreading of the non-compete rule, repeated on every contract where it applies.
What BENCH changes
1 sentence
Spell out in the prompt what the rule does and does not require, then score the change on the same 39 contracts.
Where it lands
100%
The same outcome as Claude Opus 4.5 at 100%, for 32 times less. Projected from the verdicts already recorded, and BENCH re-runs it to prove it.
It is not only the cheap one. Each model here has a single rule it reads as stricter than it is written, and that one rule accounts for nearly everything it gets wrong.
What goes wrong
All 13 of its wrong calls are on one rule: Non-compete. It reads that rule as stricter than it is written.
What BENCH would fix
Spell out in the prompt what the non-compete rule actually requires, then run the same 39 contracts again.
67%100%+13 contracts
A projection, not a second run: every one of the 13 contracts it got wrong failed only on that rule. BENCH proves it by running them again.
What goes wrong
16 of its 21 wrong calls are on one rule: Non-compete. It reads that rule as stricter than it is written.
What BENCH would fix
Spell out in the prompt what the non-compete rule actually requires, then run the same 39 contracts again.
54%87%+13 contracts
A projection, not a second run: 13 of the 18 contracts it got wrong failed only on that rule. BENCH proves it by running them again.
What goes wrong
5 of its 6 wrong calls are on one rule: Non-compete. It reads that rule as stricter than it is written.
What BENCH would fix
Spell out in the prompt what the non-compete rule actually requires, then run the same 39 contracts again.
85%97%+5 contracts
A projection, not a second run: 5 of the 6 contracts it got wrong failed only on that rule. BENCH proves it by running them again.
The same pattern holds for 9 more models: one rule, misread the same way, accounting for most of what it gets wrong.
Grouped by model, with repeats counted rather than listed. This is the evidence behind the numbers above.
The 13-month duration is a real breach and is correctly flagged. But the review also treats the 50% pay as a breach, claiming the non-compete rule requires full salary, while the non-compete rule only requires 'compensation for that period'. That invents a payment breach, and the replacement wording imposes full pay.
Clause 7 only claims outside-hours inventions that 'could reasonably be of interest to Northstar's business', so unrelated inventions stay with the employee and the draft complies with the invention ownership rule. The review invents a breach and proposes replacement wording, and it also contradicts its own opening line that there are no breaches.
Clause 7 claims only inventions made in employment or of reasonable interest to the business, and so does not take unrelated out-of-hours inventions. The review nonetheless reports a Clause 7 breach, saying a positive statement of employee ownership is missing. That is an invented breach in a compliant clause.
Clause 3 'exactly 3 months' written notice' satisfies the 'at least 3 months' rule, so the draft is compliant. The review invents a breach ('exactly' is more restrictive) and proposes unnecessary replacement wording.
The draft's 'ninety days' written notice appears intended as a paraphrase of three months in an otherwise compliant draft, and the review flagged it as a breach. Ninety days can fall short of three months, so this is a judgment call, and I resolved it against the review.
The 9-month probation is a breach and the response eventually says so ('it sets probation at 9 months, exceeding the 6-month maximum'). It first contradicts itself ('All other reviewed clauses comply'), and it gives no replacement wording, which the criterion requires.
The review flags 13 weeks' notice as a breach of the 3-month rule. 13 weeks (91 days) is essentially 3 months, and the draft has no annex changing it, so this reads as a false positive. The 'less than three months in most parts of the year' reasoning is also dubious.
The body and table correctly mark Clause 3 (3 months' notice) as the notice period rule-compliant. The closing line then says the 9-month probation 'violates the notice period rule the probation length rule', which wrongly attributes a violation to the notice period rule and contradicts the earlier finding, so it reports a notice problem in a compliant draft.
The 6-month non-compete is paid EUR 1/month, which literally satisfies the non-compete rule (12 months or less, with compensation paid). The response adds a 'genuine compensation' test and says the rule 'explicitly treats' token pay as a breach, which the playbook does not say. It therefore invents a breach in a draft that satisfies the rule as written.
The review says clause 3 'Complies with the notice period rule in substance', but then flags 'ninety days' as possibly 'marginally short of the notice period rule' and proposes replacement wording. That reports a notice problem in a draft the playbook most likely treats as compliant, a false positive on a paraphrase of three months.
Clause 6 is 12 months and pays 50% of salary, which is compensation and so satisfies the non-compete rule. The review invents a breach, saying 'unpaid for half the time', and rewrites the clause to 100% pay.
The review says "Ninety days equals three months", but 90 days is generally shorter than 3 months (most 3-month spans are 91–92 days). Clause 3 therefore likely breaches the notice period rule, and the review missed it and offered no corrected wording.
Clause 7 assigns all inventions to Northstar 'whether or not they were made during working hours and whether or not they relate to Northstar's business', which breaches the invention ownership rule. The response says there are no breaches and offers no replacement wording.
The response opens with 'I found one breach' and says the clause 'breaches the non-compete rule', then says it complies. This contradiction reports a non-compete problem in a compliant draft and leaves the finding unclear.
Clause 4 is compliant, but the response opens with 'I found one breach' and a 'Breach: Clause 4' / 'Rule broken: the overtime rule' heading before retracting it. This reports an overtime problem in a compliant draft and contradicts itself.
Clause 6 sets a 13-month non-compete, which exceeds the 12-month limit in the non-compete rule. The review says the draft is compliant and gives no correction, so it misses the breach.
Clause 4 says 'The first 10 hours of overtime each month are covered by the salary'. The overtime rule treats wording that makes overtime covered by salary as a breach, but the review calls it compliant, reads the rule as breached only when all overtime is covered, and offers no replacement wording.
Clause 5 gives only 20 days of holiday, below the 28-day minimum. The review said 'I do not see any other breaches', so it missed this breach and proposed no correction.
Clause 2 sets a 9-month probation, which breaches the probation length rule's 6-month limit. The review says 'I do not see any other breaches' and misses it, with no corrective wording.
The 12-month non-compete is paid at 50%, which satisfies the non-compete rule. The review wrongly reports a breach and invents a full-salary requirement.
Breach 1 (13 months over the 12-month limit) is correct. Breach 2 is invented: the non-compete rule only requires that compensation be paid, and the clause pays 50% of salary. The review claims the non-compete rule 'requires full compensation', which the playbook does not say, and proposes unnecessary replacement wording.
The non-compete is 12 months and paid (50%), so it is compliant under the non-compete rule; the review wrongly reports a breach and proposes 100% pay.
The review calls the 24-month clause 'compliant' while admitting it exceeds the 12-month maximum. It relies on an annex with 100% pay that does not appear in the supplied document, so the the non-compete rule conclusion is unsupported and self-contradictory.
Flagged clause 3 'thirteen weeks' as breaching the notice period rule with an incorrect rationale ('3 months and 1 week'); thirteen weeks is about 3 months, so this invents a breach in a compliant draft.
Clause 7 already says inventions made outside working hours and unrelated to the business remain the employee's property, so the invention ownership rule is satisfied. The review invents a breach and ignores that second sentence.
Clause 7 only claims out-of-hours inventions that 'could reasonably be of interest to Northstar's business', so unrelated inventions stay with the employee and the draft satisfies the invention ownership rule. The review nonetheless reports a Breach 1 and proposes replacement wording, which invents a breach in a compliant draft.
The supplied contract has 20 days of holiday and no annex text. The review declares the holiday rule compliant on the strength of an annex text that is not in the provided document, so the compliance finding is unsupported.
A model is only part of what you ship. What you actually run is a model, a prompt and a harness around it: what it can see, what it can call, how many turns it gets, what happens when something fails. Every model here got this one, unchanged.
The instruction every model was given, word for word
Review the supplied contract against the supplied company rules. Report every breach with a clause reference, explain why it breaches the rule, and propose replacement wording. If the contract complies, say so. Do not give legal advice or invent requirements.
Written for this test, not taken from anyone. A benchmark is only as good as its documents, so here is how these were made and what is in them.
A score is only worth something if the same thing scored everybody. This is what did the scoring and what it did when it went wrong.
None of these has been run on this suite yet. Each would be scored on the same frozen contracts before anybody claims it worked.
PROMPTNOT RUN YET
Say what the rule does not require.The dominant failure is a model reading “the employee must be paid for it” as “must be paid in full”, so compliant contracts paying half get flagged. One added sentence should remove it.
Would address: almost every invented breach in the table
HARNESSNOT RUN YET
Ask one rule at a time.Today one call covers all six rules at once. Splitting it into six smaller checks costs more calls but gives each rule the model's full attention, and makes a wrong call easier to localise.
Would address: rules missed when a breach hides in an annex
TOOLSNOT RUN YET
Let it look up past decisions.Give the review a lookup over contracts the company has already approved, so a borderline clause can be checked against how the same clause was treated before rather than judged cold.
Would address: contracts sitting exactly on the limit

YOUR AI SYSTEM. YOUR STANDARD.