· IQdoc
We scored 95 legal AI companies on what a buyer can actually verify
Six observable signals, applied mechanically to every company in our funding survey. The average score is 1.99 out of 10. Not one of the 95 has independent research published about its product.
Legal AI has raised $6.2 billion. We have spent the last few weeks building the funding ledger and the product index that catalogue where it went and what it bought. This post is about a gap those two pages kept running into.
A general counsel who wants to buy one of these products cannot check almost anything about it. Not the price. Often not which model is underneath. Almost never a measured accuracy figure from anyone other than the vendor. We went looking for the benchmarks and found four vendors who built their own and four vendors who won.
So we did the only ranking we think is defensible: not "which is best" — we have not tested these products and would not pretend otherwise — but how much of its own claim each company lets you check.
The rubric
Six signals, all observable from a public website, applied mechanically to all 95 companies on 25 August 2026.
| Signal | Scoring |
|---|---|
| Publishes a price | Published 2 · Partial 1 · None 0 |
| Names specific security certifications | Yes 1 · No 0 |
| States which models it runs on | Named 2 · Vague pledge 1 · Silent 0 |
| Publishes its own benchmark | Open, with tasks and rubrics 2 · Score only 0 |
| Participated in an independent evaluation | Yes 2 · Not found 0 |
| Independent research measures the product | Yes 1 · No 0 |
Maximum 10. Two notes on the design. A closed benchmark — a headline score with no published method — scores zero rather than negative: it adds nothing a buyer can check, but publishing one is not a demerit. And absence of independent benchmarking is recorded as not found, never as refusal; most of these companies have simply never appeared in a published evaluation round.
The result
Mean 1.99. Median 2. The highest score in the field is 7.
| Score | Companies |
|---|---|
| 7 | 2 |
| 5 | 5 |
| 4 | 7 |
| 3 | 13 |
| 2 | 32 |
| 1 | 19 |
| 0 | 17 |
Nobody scored 6, 8, 9 or 10. An industry selling judgement to professionals whose entire discipline is evidence publishes, on average, a fifth of the evidence about itself that it could.
The price wall, and who is behind it
Seventy-three of 95 publish no price at all. Twenty-one publish a real figure; one publishes a partial one.
The interesting part is who. Companies that publish a price have raised $30 million on average. Companies that publish none have raised $76 million. The wall goes up as the cheque size does — which is what you would expect from enterprise sales motions, and is also precisely the segment where a buyer most needs a reference point.
The lowest published price in the set is midpage's $30 a month. The cleanest is Zuva's: $10 a document, $1.25 by API, with a working free tier — the only true per-unit publisher in 95 companies. At the other end, Athennian publishes a $25,000 entry tier, which at that price point is rare candour. AI is a paid add-on on top of it.
Then there is the category of pricing page that isn't one. Hona's nav item labelled "Pricing" resolves to a demo booking form. Juro's is a quote calculator containing no figures. Pocketlaw's — the company is now called Miramis — is a qualifying questionnaire that ends at "Get a demo." Conveyd markets "transparent pricing" and states no figure anywhere on the site. LegalOn names three pricing tiers and prices none of them.
Certifications, and the ones that aren't
Sixty-six of 95 name a specific certification, which is the healthiest number in this exercise. But the signal is only worth what the words mean, and a slice of the field is using the vocabulary without the audit.
As of 25 August 2026, on their own public sites: Casium's badge row advertises "SOC 2 Type 3", which does not exist. Turbo Law describes itself as "SOC 2 Type II observant", which is not an audit status (the same homepage ships its statistics counters unpopulated, reading "0+"). Eve says it follows "the SOC 2 Framework" — a framework claim, not an attestation. Finch's "SOC 2 compliant" is a self-assertion with no named auditor. NLPatent — now Clerq — claims ISO 27001 certification on its homepage and ISO 27001 alignment on its security page.
Compare Genie AI, which publishes its ISO certificate number and expiry date and states plainly that it does not hold SOC 2. That disclosure is worth more than any badge row on this list, and it is the only one of its kind we found.
Two more worth naming. Flank holds ISO 42001 — the standard that audits AI governance rather than infosec — and is the only company in the cohort that does. Avvoka has held ISO 27001 continuously since 2017.
And on the other side: LawGeex, a contract-review vendor, is serving its site on an expired TLS certificate. Darrow has no security or trust page; both conventional paths 404. Crosby is a law firm selling AI-delivered legal work with no security page at all.
Almost nobody will tell you what is underneath
| Model provenance | Companies |
|---|---|
| Names its providers | 14 |
| Vague pledge ("enterprise-grade models", "your data is never trained on") | 36 |
| Silent | 45 |
Forty-five companies selling AI legal work will not say whose AI it is. This is the signal that has changed least since we started tracking the market, and the one with the clearest reason to change: model choice determines cost, latency, jurisdiction of processing and failure mode, and a firm with a data-residency obligation cannot satisfy it against a vendor that won't name the stack.
The ones that do name it are specific about it. DeepIP names OpenAI on Azure with zero data retention. Wordsmith AI names the contractual basis for zero retention with each provider. Noxtua names its own 111-billion-parameter in-house model. TrialView names OpenAI and Anthropic outright and holds no certifications at all — the exact inverse of everyone else in this survey.
Two open benchmarks. Zero independent studies.
- 2 of 95 publish a benchmark with the tasks and rubrics open: Harvey and Paxton AI.
- 30 of 95 publish a score with no method attached.
- 6 of 95 have participated in an evaluation run by someone else.
- 0 of 95 have independent research published measuring their product.
That last line is the finding this whole post exists for. Not one of these 95 companies — across $6.2 billion and 269 products — has had its accuracy measured, in public, by a party with no commercial interest in the result.
The closest anything comes: Clearbrief was scored 40.5 out of 50 by the State Bar of Nevada's AI Workgroup. Alexi participated in the Vals legal research study in October 2025 and landed at the participant average, above the lawyer baseline. IPRally has a third-party study, but IPRally supplied the tool and trained the searchers who used it. And DoNotPay has an FTC order from February 2025 barring deceptive AI-lawyer claims, which is enforcement rather than research, and counts for nothing here.
Paxton AI deserves specific credit. It ran Stanford's hallucination benchmark on itself and released the method along with the result. It is still self-run, but it is reproducible, which is the whole point. Legora, for its part, declines to open its benchmark corpus and says why: published cases leak into the next model's training data. That is a real objection and we record it as a real one.
The top of the table
| Score | Company | What earned it |
|---|---|---|
| 7 | GC AI | $500/mo published, named models, appears in two phases of an independent benchmark |
| 7 | Harvey | The only company with both an open benchmark and independent participation — and it published no score for itself |
| 5 | Blue J | $1,498/yr published, named models |
| 5 | Clearbrief | $300/mo published, independently scored by a state bar |
| 5 | Genie AI | Certificate numbers, expiry dates, and an honest negative |
| 5 | Paxton AI | Reproducible self-run benchmark |
| 5 | Wordsmith AI | Named the contractual basis for retention with each provider |
Harvey scoring top while publishing no price is an unusual shape, and worth sitting with: it is the best-funded company in the field, it open-sourced 1,200-plus tasks with roughly 75,000 expert-written rubric criteria, and it pointedly did not report a score for its own product. That is how it should be done. It also still came from a market participant, which remains the structural problem.
The bottom of the table
Seventeen companies score zero: Aavalynx, Bench IQ, Casium, Conveyd, Crosby, Darrow, Della AI, DoNotPay, Eve, Haast, Justpoint, LawGeex, LegalMation, Nexl, Semeris, Soxton, Theo Ai. Between them they have raised $459 million — 7.4% of all the capital we track.
We publish the names because a ranking with its bottom half redacted is not a ranking, and because everything scored here is a fact about a public website that anyone can check today. But a zero is a narrow statement and we want to be precise about what it does not say:
- It does not say the product is bad. We have not tested any of these products.
- It does not say the company is unsafe. It says we could not verify a claim from public sources.
- Several zeros are honest ones. Aavalynx states plainly that it is "working toward ISO and SOC certifications" — candour that scores the same as silence, which is a limitation of the rubric, not of Aavalynx.
- At least one may be a measurement artefact: Bench IQ's security link points at a Drata trust centre blocked by
robots.txt, unreadable rather than necessarily absent.
And a few zeros mean something else entirely. Della AI's domain is parked and listed for sale at $98,000. LawGeex is a dead brand behind an expired certificate. Theo Ai has repositioned from outcome prediction to portfolio management and removed every quantified claim in the process. Those are not disclosure failures. Those are companies that are gone or going, and the score is picking up the silence on the way out.
Money does not buy disclosure
The ten best-funded companies average 3.10. The other 85 average 1.86. So capital correlates with transparency — and the correlation is real but the level is not a defence. The best-funded decile of an industry built on evidence scores three out of ten on whether its own claims can be checked.
By segment, the spread is wider than the averages suggest:
| Segment | Companies | Mean score |
|---|---|---|
| Tax | 1 | 5.00 |
| Firm platform | 4 | 3.50 |
| Research & knowledge | 9 | 2.56 |
| IP & patents | 7 | 2.29 |
| Contracts & in-house | 32 | 2.25 |
| Consumer & small law | 3 | 1.67 |
| Litigation & PI | 22 | 1.55 |
| Practice ops & billing | 5 | 1.40 |
| Real estate | 3 | 1.33 |
| AI-native law firm | 5 | 1.20 |
| Immigration | 2 | 1.00 |
| Compliance | 2 | 0.50 |
Litigation and personal injury — the segment selling into contingency practices where a wrong answer is a malpractice claim — sits at 1.55.
Every score, with its inputs
Our rule for this series is that no ranking ships without its inputs visible. Here is all 95, sorted by score then alphabetically. "Models" is what the company says about its providers; "Benchmarks" is open (its own, with method), closed (its own, score only) or independent (participated in someone else's). There is no column for third-party research because the answer is the same for every row.
| Score | Company | Price | Certifications | Models | Benchmarks |
|---|---|---|---|---|---|
| 7 | GC AI | Published | Named | Named | Closed + Independent |
| 7 | Harvey | — | Named | Named | Open + Independent |
| 5 | Blue J | Published | Named | Named | — |
| 5 | Clearbrief | Published | Named | — | Independent |
| 5 | Genie AI | Published | Named | Named | — |
| 5 | Paxton AI | Published | Named | — | Open |
| 5 | Wordsmith AI | — | Named | Named | Independent |
| 4 | Callidus Legal AI | Published | Named | Vague | Closed |
| 4 | IPRally | — | Named | Vague | Independent |
| 4 | Lexroom | Published | Named | Vague | — |
| 4 | midpage | Published | Named | Vague | — |
| 4 | Streamline AI | Published | Named | Vague | — |
| 4 | Twin1 | Published | Named | Vague | — |
| 4 | Zuva | Published | Named | Vague | Closed |
| 3 | Alexi | — | Named | — | Independent |
| 3 | DeepIP | — | Named | Named | Closed |
| 3 | Enter | — | Named | Named | Closed |
| 3 | Eudia | — | Named | Named | — |
| 3 | Hello Divorce | Published | — | Vague | — |
| 3 | LegalOn Technologies | — | Named | Named | Closed |
| 3 | Legora | — | Named | Named | Closed |
| 3 | Moritz | Published | Named | — | — |
| 3 | Noxtua | — | Named | Named | Closed |
| 3 | Spellbook | — | Named | Named | — |
| 3 | Summize | — | Named | Named | Closed |
| 3 | TermScout | Published | Named | — | — |
| 3 | Trellis | Published | Named | — | — |
| 2 | Athennian | Partial | Named | — | — |
| 2 | Avvoka | — | Named | Vague | — |
| 2 | Boundless | Published | — | — | Closed |
| 2 | Bryter | — | Named | Vague | — |
| 2 | Centari | — | Named | Vague | — |
| 2 | DeepJudge | — | Named | Vague | Closed |
| 2 | EvenUp | — | Named | Vague | — |
| 2 | Fileread | — | Named | Vague | — |
| 2 | Flank | — | Named | Vague | — |
| 2 | General Legal | Published | — | — | — |
| 2 | Ivo | — | Named | Vague | Closed |
| 2 | Jus Mundi | Published | — | — | — |
| 2 | Keith | Published | — | — | Closed |
| 2 | Klarity | — | Named | Vague | — |
| 2 | Laurel | — | Named | Vague | — |
| 2 | Lawhive | Published | — | — | — |
| 2 | Luminance | — | Named | Vague | — |
| 2 | Marveri | — | Named | Vague | — |
| 2 | NLPatent | — | Named | Vague | Closed |
| 2 | Orbital | — | Named | Vague | — |
| 2 | Pandektes | — | Named | Vague | — |
| 2 | Patlytics | — | Named | Vague | Closed |
| 2 | Pocketlaw | — | Named | Vague | — |
| 2 | PointOne | — | Named | Vague | — |
| 2 | Robin AI | — | Named | Vague | — |
| 2 | Sandstone | — | Named | Vague | — |
| 2 | Skribe.ai | Published | — | — | Closed |
| 2 | Solve Intelligence | — | Named | Vague | Closed |
| 2 | Supio | — | Named | Vague | — |
| 2 | Syntracts | — | Named | Vague | — |
| 2 | Tradespace | — | Named | Vague | — |
| 2 | TrialView | — | — | Named | — |
| 1 | Case Status | — | Named | — | — |
| 1 | Chamelio | — | Named | — | — |
| 1 | Checkbox | — | Named | — | — |
| 1 | Dealstack | — | Named | — | — |
| 1 | Definely | — | Named | — | — |
| 1 | Finch | — | — | Vague | Closed |
| 1 | Hona | — | Named | — | — |
| 1 | Josef | — | Named | — | — |
| 1 | Juro | — | Named | — | — |
| 1 | LegalFly | — | Named | — | — |
| 1 | Manifest OS | — | Named | — | Closed |
| 1 | MarqVision | — | Named | — | Closed |
| 1 | Mary Technology | — | — | Vague | Closed |
| 1 | Newcode | — | — | Vague | — |
| 1 | Norm Ai | — | Named | — | Closed |
| 1 | SpotDraft | — | Named | — | — |
| 1 | Tomorro | — | Named | — | — |
| 1 | Turbo Law | — | — | Vague | — |
| 1 | Wexler AI | — | Named | — | Closed |
| 0 | Aavalynx | — | — | — | Closed |
| 0 | Bench IQ | — | — | — | — |
| 0 | Casium | — | — | — | Closed |
| 0 | Conveyd | — | — | — | Closed |
| 0 | Crosby | — | — | — | — |
| 0 | Darrow | — | — | — | — |
| 0 | Della AI | — | — | — | — |
| 0 | DoNotPay | — | — | — | — |
| 0 | Eve | — | — | — | Closed |
| 0 | Haast | — | — | — | — |
| 0 | Justpoint | — | — | — | — |
| 0 | LawGeex | — | — | — | Closed |
| 0 | LegalMation | — | — | — | — |
| 0 | Nexl | — | — | — | Closed |
| 0 | Semeris | — | — | — | Closed |
| 0 | Soxton | — | — | — | — |
| 0 | Theo Ai | — | — | — | — |
Method, and how to correct us
Every signal was read from the company's own public website between 20 and 25 August 2026 by four independent passes, then compiled and scored mechanically — no judgement calls at the scoring stage, so the same inputs always produce the same number. The dataset lives in the repo alongside the funding and product data and carries a per-company note wherever the finding needed one.
The rubric has known limits. It rewards publishing over doing: a company with excellent security and no trust page scores as though it had neither. It cannot see anything behind a login, a trust-centre gate or a sales call, and several vendors will tell you they hand over pricing and audit reports the moment you ask. That is exactly the point — a signal a buyer has to book a call to receive is not a signal the market can use.
If we have a fact wrong about your company, tell us and we will fix it and say that we did. If your score is low because the evidence sits behind a form, publish it and the score moves on the next audit. The whole exercise is designed so that it can.
This is the first of the research blocks going into The State of Legal AI, August 2026 — a full report on agents, models, benchmarks and the companies building them, organised by who they sell to. The evidence score is its centre of gravity.