IQdocApplied AI for legal.
← Blog

· IQdoc

Legal AI benchmarks, and who is marking the exam

We went through every legal AI benchmark we could find. Four vendors built their own and all four won. The most useful number in the field is the gap between two ways of scoring the same answer.

Every legal AI vendor quotes a score. We went looking for where those scores come from, who administers them, and whether a buyer can use any of it to choose between products.

Short version: the model benchmarks are real, current and increasingly uninformative. The product benchmarks are mostly written by the companies being tested. And the single most useful number in the field is not a score at all — it is the distance between two ways of marking the same answer.

What exists

BenchmarkWho built itScoresCurrent
LegalBenchStanford et al., 2023ModelsYes, but saturating
Vals Legal Research BenchVals AI + law firmsModelsAug 2026
Harvey LABHarvey, 2026ModelsAug 2026
BigLaw BenchHarvey, 2024ProductsVendor-run
VLAIRVals AI + 8 firmsProductsFeb 2025, not repeated
Contract Review BenchmarkLegalOn, 2026ProductsVendor-run
ContractScrubThomson Reuters + ImperialModelsAug 2026
LEXam, PLawBench, Multi-Legal-BenchAcademicModels2025–26

On LegalBench, the top model currently sits at 88.6% and the top eight cluster within about two and a half points. When every frontier model scores within a rounding error, the benchmark has stopped distinguishing them. Its own authors have moved on to harder things.

Four vendors built benchmarks. Four vendors won.

Harvey's BigLaw Bench (2024) reported Harvey completing about 74% of lawyer-quality work, ahead of the models it tested. Harvey wrote the tasks, wrote the rubrics, ran the evaluation and reported the result.

LegalOn's Contract Review Benchmark (June 2026) put LegalOn 87 ELO points above the next model and 400 above the best GPT. Ivo's April 2026 study scored Ivo at 4.52 against human lawyers at 4.56. GC AI's May 2026 study scored GC AI at 86.8% against ChatGPT's 79.8%.

Four benchmarks, four home wins. None of this means the products are bad. It means the numbers carry no procurement weight, and a buyer should treat them as marketing with arithmetic attached.

The honourable exception is Harvey's second attempt. LAB, released May 2026, open-sourced 1,200+ tasks across 24 practice areas with roughly 75,000 expert-written rubric criteria — and Harvey pointedly did not publish a score for its own product. That is how it should be done, and it still came from a market participant, which the author of LegalBench's own follow-up work would tell you is the structural problem.

The number that actually matters

On Vals' Legal Research Bench, the leading model scores 90.6% under partial credit and 55.3% when every required element must be present. On Harvey's LAB, models satisfy roughly 90% of individual rubric criteria but complete around 20% of whole tasks.

That gap is the whole story of legal AI reliability.

Legal work is conjunctive. A memo that gets nine things right and misses the tenth is not 90% correct; it is wrong, and possibly wrong in a way that costs a client. Every headline accuracy figure you have read is the generous number. Ask for the strict one.

The same benchmark also found the hardest category was reconciliation — synthesising authorities that conflict — at 20.7% all-pass. Which is, not coincidentally, a decent description of what lawyers are for.

You cannot compare the products you would actually buy

The most thorough independent product study is still VLAIR, from February 2025. It measured Harvey, CoCounsel, Vincent and Oliver on real law-firm data against a lawyer control group. Harvey took the top score on five of six tasks it entered, hitting 94.8% on document Q&A against a lawyer baseline of 70.1%. Lawyers beat every tool on redlining.

It is eighteen months old, it did not include Legora, and Lexis+ AI enrolled and then withdrew from most tasks.

There has been no second edition. Vals ran a legal research study in October 2025 in which Thomson Reuters, LexisNexis and vLex all declined to take part. Legora, now valued at $5.55B, has not appeared in any independent public benchmark we could find.

That is the structural flaw, and it is not methodological. Vendors choose whether to be measured, and the ones with the most to lose choose not to be.

"Hallucination-free" was not

The most consequential legal AI research of the last two years was not a benchmark. Stanford's RegLab tested the major legal research tools on 202 queries and published the results in the Journal of Empirical Legal Studies: Lexis+ AI hallucinated 17% of the time, Westlaw's AI-Assisted Research 33%, where hallucination included confidently citing a source that did not support the claim.

Both had been marketed on the absence of that problem.

Retrieval reduces hallucination substantially. It does not eliminate it. A Stanford follow-up in February 2026 testing statutory research found Lexis+ AI at 64% and Westlaw at 58%, below a generic retrieval baseline at 70% — commercial legal platforms underperforming an off-the-shelf approach on that task.

Worth remembering the same lesson applies to the bar exam. GPT-4's famous "90th percentile" result was re-examined against licensed attorneys rather than a pool padded with repeat takers: roughly 45th percentile overall, and 15th on the essays — the half that resembles legal writing.

What varies more than the model

Two findings should end the phrase "best legal AI."

Vals' research bench found scores by practice area ranging from 38.4% in health law to 11.4% in family law on the same model. And Multi-Legal-Bench, testing across six European jurisdictions in May 2026, found no model dominating in any language — rankings reorder by task and by country.

The right question is not which model is best. It is which model is best at your work, in your jurisdiction, and nobody's leaderboard answers that.

The thing benchmarks do not capture

One result is worth more than every accuracy score here. In an independent study of contract drafting, legal-specific tools raised risk warnings in 83% of high-risk scenarios against 55% for general-purpose AI — while the general-purpose models scored higher on raw output quality.

Knowing when to stop and flag something is not an accuracy metric. It may be the most lawyerly capability there is, and the leaderboards do not score it.

How to read the next score you are shown

Ask four questions. Who wrote the benchmark, and do they sell a product in it. Is that the partial-credit number or the all-pass number. Which practice area and jurisdiction was it measured in. And who declined to participate.

In January 2026 the authors of LegalBench published a paper in PNAS warning that benchmarks "can be captured, watered down, and abused." They built the field's most-cited legal benchmark, and that is their own assessment of what happens next.

We publish our own accuracy numbers and the caveats that go with them, which puts us in the same position as everyone else here: a market participant with a score to quote. Read ours the same way.

Figures are as published by the sources named, on the dates named. Nothing here is legal advice, and we have not independently replicated any of these benchmarks.