· IQdoc
Everybody's transcription is 97% accurate
We benchmarked seven transcription methods against certified Supreme Court transcripts. They all landed within a point of each other. The gap that matters for legal work is somewhere else entirely.
Pick a transcription service and you will be told it is 99% accurate. Rev advertises 99%+ on its human tier and 99.6% for court-ready work. GoTranscript guarantees 99% with money back below it. Verbit says 99% for legal. TranscribeMe says 99%.
Read the small print and almost every one of those numbers describes a human or hybrid tier. Rev's AI-only product is advertised at 96%+, which is a different product at a different price.
So we measured. Two full U.S. Supreme Court oral arguments, 23,544 scored words, graded against the official Heritage Reporting certified transcripts, with every contender put through an identical scoring path. The full leaderboard is here.
The results are boringly close
| Method | Errors | Accuracy |
|---|---|---|
| Our extreme mode + review of flagged words | 328 | 98.6% |
| Our advanced mode (2-engine) | 665 | 97.2% |
| ElevenLabs Scribe, alone | 668 | 97.2% |
| Our extreme mode (3-engine) | 670 | 97.2% |
| Our good mode (AssemblyAI, alone) | 701 | 97.0% |
| Rev.com | 761 | 96.8% |
| Speechmatics, alone | 866 | 96.3% |
Every fully automated method, ours included, lands between 96.3% and 97.2%. Rev came in at 96.8% — right where Rev says its AI tier lands. Rev is not overselling anything. It is doing what it says on clean audio.
The spread between the best automated engine and the worst is 198 errors in 23,544 words. If you are choosing a transcription vendor on headline accuracy, you are choosing between things that are nearly the same.
97% is a much worse number than it sounds
Three percent of 23,544 words is around 700 errors. That is 145 minutes of audio — well-miked appellate argument with orderly turn-taking, judges who wait their turn, no crosstalk, no bad microphone at the back of a municipal courtroom. It is the easiest court audio that exists.
A full day of deposition is several times that length and considerably messier. On the same error rate, you are looking at thousands of wrong words in a document that is supposed to be what happened.
Nobody markets it that way because 97% sounds like an A.
The errors are not randomly distributed
This is the part that matters for legal work and that general transcription is not built to care about.
Errors cluster on proper nouns, case names, numbers and terms of art — precisely the words with legal consequence. On the case-caption words in one argument, we measured 16 errors out of 1,077 against Rev's 22. Both are small percentages. Both are the wrong party's name in a document somebody will rely on.
Then there is who said it. General speech-to-text optimises for what was said, because that is what a podcast or a meeting summary needs. Legal work needs attribution: the same sentence means opposite things depending on whether the judge or counsel said it. On speaker error we measured 0.054 against Rev's 0.065 on one argument — roughly 17% fewer words caught in a speaker mix-up.
Those are close numbers too. We are not claiming a chasm. We are saying the metric general services optimise is not the metric that decides whether a transcript is usable in court. IQdoc exists because of a speaker error — a passage read aloud in a hearing with one person's words attributed to another.
What actually moved the number
Look at the top row of the leaderboard again. Going from 665 errors to 328 — cutting them in half — did not come from a better model. Every engine we tried was stuck around 97%.
It came from review: having the system flag the words it is least sure about, and having a human resolve those. Examining roughly 3% of the words caught 70–75% of the errors we could prove. That is the whole trick, and it is unglamorous. Not a smarter transcriber. A transcriber that knows what it does not know, and an honest place to put a person.
General services do offer human tiers, and theirs is how they get to 99% too. The difference is what the human is asked to do: read everything at a linear cost, or check the 3% that is actually in doubt.
And then it has to be a document
Word accuracy is only the first problem. A court transcript is not text. It is a formatted instrument: line numbers, a caption matching the court's requirements, an index, a certification page, pagination the citing brief depends on. A general service returns you a text file and a formatting job.
What this benchmark does not show
Being specific about limits is the only reason to publish numbers at all.
The audio is the ceiling condition — clean appellate argument. Your county's motion calendar will be worse for everyone in that table. We tested Rev's standard automated product, not its human tier, so this is not a comparison with Rev's 99% offering. The review row assumes a flawless reviewer, which is a generous assumption. It is two arguments, not a corpus. And we have already had to correct these numbers once after finding scoring errors in our own method.
It is a relative benchmark against a certified reference. It is not an absolute claim about your hearing.
That is more hedging than a marketing page would carry. It is also the only kind of accuracy claim worth reading, which was rather the point.