IQdocApplied AI for legal.
← Blog

· IQdoc

We scored 95 legal AI companies on what a buyer can actually verify

Six observable signals, applied mechanically to every company in our funding survey. The average score is 1.99 out of 10. Not one of the 95 has independent research published about its product.

Legal AI has raised $6.2 billion. We have spent the last few weeks building the funding ledger and the product index that catalogue where it went and what it bought. This post is about a gap those two pages kept running into.

A general counsel who wants to buy one of these products cannot check almost anything about it. Not the price. Often not which model is underneath. Almost never a measured accuracy figure from anyone other than the vendor. We went looking for the benchmarks and found four vendors who built their own and four vendors who won.

So we did the only ranking we think is defensible: not "which is best" — we have not tested these products and would not pretend otherwise — but how much of its own claim each company lets you check.

The rubric

Six signals, all observable from a public website, applied mechanically to all 95 companies on 25 August 2026.

SignalScoring
Publishes a pricePublished 2 · Partial 1 · None 0
Names specific security certificationsYes 1 · No 0
States which models it runs onNamed 2 · Vague pledge 1 · Silent 0
Publishes its own benchmarkOpen, with tasks and rubrics 2 · Score only 0
Participated in an independent evaluationYes 2 · Not found 0
Independent research measures the productYes 1 · No 0

Maximum 10. Two notes on the design. A closed benchmark — a headline score with no published method — scores zero rather than negative: it adds nothing a buyer can check, but publishing one is not a demerit. And absence of independent benchmarking is recorded as not found, never as refusal; most of these companies have simply never appeared in a published evaluation round.

The result

Mean 1.99. Median 2. The highest score in the field is 7.

ScoreCompanies
72
55
47
313
232
119
017

Nobody scored 6, 8, 9 or 10. An industry selling judgement to professionals whose entire discipline is evidence publishes, on average, a fifth of the evidence about itself that it could.

The price wall, and who is behind it

Seventy-three of 95 publish no price at all. Twenty-one publish a real figure; one publishes a partial one.

The interesting part is who. Companies that publish a price have raised $30 million on average. Companies that publish none have raised $76 million. The wall goes up as the cheque size does — which is what you would expect from enterprise sales motions, and is also precisely the segment where a buyer most needs a reference point.

The lowest published price in the set is midpage's $30 a month. The cleanest is Zuva's: $10 a document, $1.25 by API, with a working free tier — the only true per-unit publisher in 95 companies. At the other end, Athennian publishes a $25,000 entry tier, which at that price point is rare candour. AI is a paid add-on on top of it.

Then there is the category of pricing page that isn't one. Hona's nav item labelled "Pricing" resolves to a demo booking form. Juro's is a quote calculator containing no figures. Pocketlaw's — the company is now called Miramis — is a qualifying questionnaire that ends at "Get a demo." Conveyd markets "transparent pricing" and states no figure anywhere on the site. LegalOn names three pricing tiers and prices none of them.

Certifications, and the ones that aren't

Sixty-six of 95 name a specific certification, which is the healthiest number in this exercise. But the signal is only worth what the words mean, and a slice of the field is using the vocabulary without the audit.

As of 25 August 2026, on their own public sites: Casium's badge row advertises "SOC 2 Type 3", which does not exist. Turbo Law describes itself as "SOC 2 Type II observant", which is not an audit status (the same homepage ships its statistics counters unpopulated, reading "0+"). Eve says it follows "the SOC 2 Framework" — a framework claim, not an attestation. Finch's "SOC 2 compliant" is a self-assertion with no named auditor. NLPatent — now Clerq — claims ISO 27001 certification on its homepage and ISO 27001 alignment on its security page.

Compare Genie AI, which publishes its ISO certificate number and expiry date and states plainly that it does not hold SOC 2. That disclosure is worth more than any badge row on this list, and it is the only one of its kind we found.

Two more worth naming. Flank holds ISO 42001 — the standard that audits AI governance rather than infosec — and is the only company in the cohort that does. Avvoka has held ISO 27001 continuously since 2017.

And on the other side: LawGeex, a contract-review vendor, is serving its site on an expired TLS certificate. Darrow has no security or trust page; both conventional paths 404. Crosby is a law firm selling AI-delivered legal work with no security page at all.

Almost nobody will tell you what is underneath

Model provenanceCompanies
Names its providers14
Vague pledge ("enterprise-grade models", "your data is never trained on")36
Silent45

Forty-five companies selling AI legal work will not say whose AI it is. This is the signal that has changed least since we started tracking the market, and the one with the clearest reason to change: model choice determines cost, latency, jurisdiction of processing and failure mode, and a firm with a data-residency obligation cannot satisfy it against a vendor that won't name the stack.

The ones that do name it are specific about it. DeepIP names OpenAI on Azure with zero data retention. Wordsmith AI names the contractual basis for zero retention with each provider. Noxtua names its own 111-billion-parameter in-house model. TrialView names OpenAI and Anthropic outright and holds no certifications at all — the exact inverse of everyone else in this survey.

Two open benchmarks. Zero independent studies.

  • 2 of 95 publish a benchmark with the tasks and rubrics open: Harvey and Paxton AI.
  • 30 of 95 publish a score with no method attached.
  • 6 of 95 have participated in an evaluation run by someone else.
  • 0 of 95 have independent research published measuring their product.

That last line is the finding this whole post exists for. Not one of these 95 companies — across $6.2 billion and 269 products — has had its accuracy measured, in public, by a party with no commercial interest in the result.

The closest anything comes: Clearbrief was scored 40.5 out of 50 by the State Bar of Nevada's AI Workgroup. Alexi participated in the Vals legal research study in October 2025 and landed at the participant average, above the lawyer baseline. IPRally has a third-party study, but IPRally supplied the tool and trained the searchers who used it. And DoNotPay has an FTC order from February 2025 barring deceptive AI-lawyer claims, which is enforcement rather than research, and counts for nothing here.

Paxton AI deserves specific credit. It ran Stanford's hallucination benchmark on itself and released the method along with the result. It is still self-run, but it is reproducible, which is the whole point. Legora, for its part, declines to open its benchmark corpus and says why: published cases leak into the next model's training data. That is a real objection and we record it as a real one.

The top of the table

ScoreCompanyWhat earned it
7GC AI$500/mo published, named models, appears in two phases of an independent benchmark
7HarveyThe only company with both an open benchmark and independent participation — and it published no score for itself
5Blue J$1,498/yr published, named models
5Clearbrief$300/mo published, independently scored by a state bar
5Genie AICertificate numbers, expiry dates, and an honest negative
5Paxton AIReproducible self-run benchmark
5Wordsmith AINamed the contractual basis for retention with each provider

Harvey scoring top while publishing no price is an unusual shape, and worth sitting with: it is the best-funded company in the field, it open-sourced 1,200-plus tasks with roughly 75,000 expert-written rubric criteria, and it pointedly did not report a score for its own product. That is how it should be done. It also still came from a market participant, which remains the structural problem.

The bottom of the table

Seventeen companies score zero: Aavalynx, Bench IQ, Casium, Conveyd, Crosby, Darrow, Della AI, DoNotPay, Eve, Haast, Justpoint, LawGeex, LegalMation, Nexl, Semeris, Soxton, Theo Ai. Between them they have raised $459 million — 7.4% of all the capital we track.

We publish the names because a ranking with its bottom half redacted is not a ranking, and because everything scored here is a fact about a public website that anyone can check today. But a zero is a narrow statement and we want to be precise about what it does not say:

  • It does not say the product is bad. We have not tested any of these products.
  • It does not say the company is unsafe. It says we could not verify a claim from public sources.
  • Several zeros are honest ones. Aavalynx states plainly that it is "working toward ISO and SOC certifications" — candour that scores the same as silence, which is a limitation of the rubric, not of Aavalynx.
  • At least one may be a measurement artefact: Bench IQ's security link points at a Drata trust centre blocked by robots.txt, unreadable rather than necessarily absent.

And a few zeros mean something else entirely. Della AI's domain is parked and listed for sale at $98,000. LawGeex is a dead brand behind an expired certificate. Theo Ai has repositioned from outcome prediction to portfolio management and removed every quantified claim in the process. Those are not disclosure failures. Those are companies that are gone or going, and the score is picking up the silence on the way out.

Money does not buy disclosure

The ten best-funded companies average 3.10. The other 85 average 1.86. So capital correlates with transparency — and the correlation is real but the level is not a defence. The best-funded decile of an industry built on evidence scores three out of ten on whether its own claims can be checked.

By segment, the spread is wider than the averages suggest:

SegmentCompaniesMean score
Tax15.00
Firm platform43.50
Research & knowledge92.56
IP & patents72.29
Contracts & in-house322.25
Consumer & small law31.67
Litigation & PI221.55
Practice ops & billing51.40
Real estate31.33
AI-native law firm51.20
Immigration21.00
Compliance20.50

Litigation and personal injury — the segment selling into contingency practices where a wrong answer is a malpractice claim — sits at 1.55.

Every score, with its inputs

Our rule for this series is that no ranking ships without its inputs visible. Here is all 95, sorted by score then alphabetically. "Models" is what the company says about its providers; "Benchmarks" is open (its own, with method), closed (its own, score only) or independent (participated in someone else's). There is no column for third-party research because the answer is the same for every row.

ScoreCompanyPriceCertificationsModelsBenchmarks
7GC AIPublishedNamedNamedClosed + Independent
7HarveyNamedNamedOpen + Independent
5Blue JPublishedNamedNamed
5ClearbriefPublishedNamedIndependent
5Genie AIPublishedNamedNamed
5Paxton AIPublishedNamedOpen
5Wordsmith AINamedNamedIndependent
4Callidus Legal AIPublishedNamedVagueClosed
4IPRallyNamedVagueIndependent
4LexroomPublishedNamedVague
4midpagePublishedNamedVague
4Streamline AIPublishedNamedVague
4Twin1PublishedNamedVague
4ZuvaPublishedNamedVagueClosed
3AlexiNamedIndependent
3DeepIPNamedNamedClosed
3EnterNamedNamedClosed
3EudiaNamedNamed
3Hello DivorcePublishedVague
3LegalOn TechnologiesNamedNamedClosed
3LegoraNamedNamedClosed
3MoritzPublishedNamed
3NoxtuaNamedNamedClosed
3SpellbookNamedNamed
3SummizeNamedNamedClosed
3TermScoutPublishedNamed
3TrellisPublishedNamed
2AthennianPartialNamed
2AvvokaNamedVague
2BoundlessPublishedClosed
2BryterNamedVague
2CentariNamedVague
2DeepJudgeNamedVagueClosed
2EvenUpNamedVague
2FilereadNamedVague
2FlankNamedVague
2General LegalPublished
2IvoNamedVagueClosed
2Jus MundiPublished
2KeithPublishedClosed
2KlarityNamedVague
2LaurelNamedVague
2LawhivePublished
2LuminanceNamedVague
2MarveriNamedVague
2NLPatentNamedVagueClosed
2OrbitalNamedVague
2PandektesNamedVague
2PatlyticsNamedVagueClosed
2PocketlawNamedVague
2PointOneNamedVague
2Robin AINamedVague
2SandstoneNamedVague
2Skribe.aiPublishedClosed
2Solve IntelligenceNamedVagueClosed
2SupioNamedVague
2SyntractsNamedVague
2TradespaceNamedVague
2TrialViewNamed
1Case StatusNamed
1ChamelioNamed
1CheckboxNamed
1DealstackNamed
1DefinelyNamed
1FinchVagueClosed
1HonaNamed
1JosefNamed
1JuroNamed
1LegalFlyNamed
1Manifest OSNamedClosed
1MarqVisionNamedClosed
1Mary TechnologyVagueClosed
1NewcodeVague
1Norm AiNamedClosed
1SpotDraftNamed
1TomorroNamed
1Turbo LawVague
1Wexler AINamedClosed
0AavalynxClosed
0Bench IQ
0CasiumClosed
0ConveydClosed
0Crosby
0Darrow
0Della AI
0DoNotPay
0EveClosed
0Haast
0Justpoint
0LawGeexClosed
0LegalMation
0NexlClosed
0SemerisClosed
0Soxton
0Theo Ai

Method, and how to correct us

Every signal was read from the company's own public website between 20 and 25 August 2026 by four independent passes, then compiled and scored mechanically — no judgement calls at the scoring stage, so the same inputs always produce the same number. The dataset lives in the repo alongside the funding and product data and carries a per-company note wherever the finding needed one.

The rubric has known limits. It rewards publishing over doing: a company with excellent security and no trust page scores as though it had neither. It cannot see anything behind a login, a trust-centre gate or a sales call, and several vendors will tell you they hand over pricing and audit reports the moment you ask. That is exactly the point — a signal a buyer has to book a call to receive is not a signal the market can use.

If we have a fact wrong about your company, tell us and we will fix it and say that we did. If your score is low because the evidence sits behind a form, publish it and the score moves on the next audit. The whole exercise is designed so that it can.

This is the first of the research blocks going into The State of Legal AI, August 2026 — a full report on agents, models, benchmarks and the companies building them, organised by who they sell to. The evidence score is its centre of gravity.