AI-Assisted Redlining vs Manual Review Accuracy
A 2018 study found AI outpaced lawyers on contract review, but the real story is more complicated.

The number everyone quotes comes from a 2018 study run by LawGeex. Twenty experienced corporate lawyers, pulled from firms including Goldman Sachs, Cisco, Alston & Bird, and K&L Gates, reviewed five NDAs totaling 153 paragraphs of fairly dense legal language, under controlled conditions designed to make the comparison fair. The AI system scored 94% on a combined precision-and-recall metric called the F-measure; the lawyers averaged 85%.
The accuracy gap made headlines, but the speed gap deserves equal attention. The AI finished in 26 seconds, while the lawyers averaged 92 minutes, with the slowest clocking in at 156 minutes and the fastest at 51. Put another way: the fastest human in the room was still roughly 118 times slower than the machine, and that human was presumably having a very good day.
The spread among lawyers deserves its own attention, separate from the average. The best individual lawyer in the study matched the AI's 94% exactly, while the worst hit 67%, a 27-point gap between colleagues doing the identical task on the identical documents. That spread is arguably the more interesting finding than the AI's win. The case for AI rests as much on eliminating a range of outcomes wide enough to drive a truck through as on beating an average human.
But what does this study actually license someone to conclude? The task was narrow by design: standardized NDAs, a document type with predictable structure, sitting squarely in AI's comfort zone. It's now several years old, and both large language models and purpose-built legal AI tools have moved a long way since 2018. Nowhere in the study is there a test of complex commercial agreements, multi-party deals, jurisdiction-specific clauses, or the kind of strategic redlining judgment a senior associate brings to a term sheet negotiation. The accuracy advantage documented here is real, but it's real within a fenced-in yard. The honest question, and the one the rest of this piece tries to answer, is how far that fence actually extends.
How accuracy figures have evolved since 2018 (and what the newer benchmarks measure)
Dedicated redlining platforms now claim numbers that make the LawGeex figures look almost modest. LexCheck's 2024 benchmarks report AI clause-identification accuracy between 94% and 97% on standard commercial contracts, against roughly 80% for manual review. Dioptra reports 95% accuracy on first-party paper revisions, 92% on third-party papers, and 94% on issue detection. Worth flagging plainly: these are vendor-reported figures, useful as a directional signal, not as independent verification, and a company grading its own homework tends to grade generously.
Independent benchmarking tells a messier, more interesting story. A 2025 LegalBenchmarks study looking at contract first drafts found the top-performing AI system hit 73.3% reliability against 56.7% for human lawyers overall, with the single best human lawyer reaching 70%. AI still leads here, but the margin over the best individual human nearly disappears once drafting, rather than clause identification, is the task being measured. That's a meaningfully different exercise than spotting a missing indemnification clause; drafting requires generating language that fits, not just flagging language that's off.
General-purpose AI tells a starker cautionary tale. Gemini 1.5 Pro, tested on a 242-page commercial contract across 72 checklist questions, reached 64% accuracy, and only with advanced step-by-step prompting engineered specifically to coax better performance out of it. Without that scaffolding, roughly one in three answers needed a human to fix them. That's a meaningful gap, one that draws a fairly bright line between AI systems built and trained specifically for legal contract work and general-purpose language models pressed into service for a task they were never purpose-built to handle.
Where the productivity case gets genuinely clean is speed at scale. AI first-pass assistance cut average per-contract review time from 92 minutes to 22 minutes, a 76% reduction on standard commercial agreements. An eight-week pilot across 28 in-house legal teams using Axiom's DraftPilot tool reported 40% to 60% average time savings on routine review: contract summaries that used to take five days came back in one, redlines that used to take two hours took thirty minutes, and 89% of the attorneys involved said quality and consistency also improved. Part of the mechanism here is that AI reads faster clause by clause, but a bigger part is that the queue disappears. Issues get surfaced before a lawyer even opens the document, which changes the whole shape of the workday.
Where AI holds its accuracy edge and where the advantage collapses by contract type
Here's where the tiering starts to matter, because "AI accuracy" is not one number, and depends entirely on what's being reviewed.
Standardized, high-volume agreements are AI's home turf, no contest. NDAs, simple vendor agreements, form contracts: these follow patterns predictable enough that AI reportedly delivers 90% to 95% time reduction with minimal human review needed afterward. The reason is structural rather than mysterious. AI applies the same playbook rule to the same clause every single time, at any volume, without getting tired on contract 4,000 of a 5,000-contract batch. This is the territory where the LawGeex numbers hold up best and where organizations tend to see returns almost immediately.
Move up one tier, into licensing deals, service agreements, and commercial contracts carrying bespoke negotiated terms, and the picture becomes a shared project rather than a machine victory. Reported time reduction drops to a still-substantial 80% to 85%, and the division of labor becomes explicit: AI handles routine clause identification and flags deviations from a playbook, while the lawyer focuses on negotiation strategy and the business context that no dataset captures. This is arguably the healthiest version of the hybrid model, because neither side is pretending to do the other's job.
Then there's the top tier, where AI's accuracy claims mostly just stop applying. Complex, high-stakes, or jurisdiction-specific contracts require things AI genuinely does not have access to: negotiation strategy (which clause the client is actually willing to trade away), business context (a clause that looks alarming in the abstract might be completely standard in the client's specific industry), and relationship history with the counterparty. Deliberate ambiguity, the kind a skilled negotiator writes into a clause on purpose, requires a human to notice it was deliberate at all. Most organizations report needing two to three months of feedback-based refinement before AI performs reliably even on their own standard contract types, which is a useful reminder that published accuracy figures describe mature, tuned deployments, not what happens the day the software gets installed.
There's a structural dependency underneath all of this worth calling out directly. Roughly 60% of legal teams, per a 2024 LegalOn survey, do not operate with a playbook, the internal rulebook that tells AI what counts as "standard" language and what counts as a "deviation" worth flagging. Without that playbook, AI redlining ends up applying generic industry rules to an organization's specific contracts, a mismatch that tends to produce a flood of false positives and, eventually, an attorney who stops trusting the tool's output altogether. It is hard to overstate how much of the accuracy conversation quietly assumes a playbook exists.
The hallucination problem (what the research actually shows and why it is harder to dismiss than vendors suggest)
Hallucination is the word everyone in legal tech would prefer stayed out of the conversation, and yet it's the most rigorously studied part of it. Stanford's RegLab ran a study testing legal AI tools across a set of queries, each hand-scored by legal experts. The hallucination rate for Lexis+ AI came in at 17%, with other tools and general-purpose models performing materially worse. The errors varied in type and subtlety, and some were precisely the kind a rushed reviewer might miss.
Two caveats keep this honest. Stanford tested legal research tools, answering legal questions, not dedicated contract-redlining platforms doing clause comparison; the tasks are related but not identical. And the tool versions tested represent a point-in-time snapshot; both products have since been updated, so the exact numbers may no longer describe the current state of either product.
That said, the underlying machinery (a large language model reasoning over legal text and producing an answer with total confidence whether or not it's right) is shared across research tools and redlining tools alike. The failure mode Stanford documented, a confidently wrong answer indistinguishable on its face from a confidently right one, is the same risk anyone runs when an AI model interprets contract language it was not specifically trained to interpret. Reported instances of AI-driven legal hallucinations have grown alongside broader AI adoption in legal practice; some of that growth is just more AI use generally, but it's not nothing.
Why do general-purpose models hallucinate legal specifics so readily? Because they were not trained specifically on contract documents and do not have any built-in understanding of how legal risk gets allocated between two parties. Confidence, in these systems, is not correlated with accuracy; a model will state a fabricated citation with exactly the same tone it uses for a real one. That matters more at scale than it sounds like it should. A hallucination rate that reads as small on a single contract (17%, say) stops being a statistical footnote once it's running across a portfolio of several hundred agreements, and one in six wrong is not a rounding error when the volume is high enough.
Why human oversight is not a workaround but a structural part of how AI redlining actually works
Here's a fact that should reframe how the whole accuracy debate gets read: a consistent finding across surveys is that the majority of legal professionals review AI contract output before acting on it. This reflects the workflow working exactly as designed. AI surfaces the issues; a human decides what to do about them. Neither step is optional in the current model, and neither one is pretending to be.
This changes what those headline accuracy figures are actually measuring. When a vendor reports 94% to 97% clause-identification accuracy, that's accuracy on a first pass a lawyer still reviews afterward, distinct from accuracy as a stand-in for the review itself. The real gain shows up in the math from earlier: the lawyer spends 22 minutes confirming and adjusting instead of 92 minutes reading from a blank slate. That's the productivity story, and it's a genuinely good one, though a different story from AI replacing legal judgment.
One might argue that high accuracy is itself an argument for lighter oversight. The reason that argument falls short is worth sitting with: AI errors tend to be systematic rather than random. A human reviewer having a bad day might miss a clause in contract 12 and catch it fine in contract 13, but AI misses the same clause the same way, every single time, across every contract that shares that pattern, quietly and consistently, until a lawyer happens to notice the pattern itself. That is a different kind of risk than ordinary human inconsistency, and arguably a more dangerous one, because it hides behind a high accuracy score rather than showing up as an obvious outlier.
The cognitive division that actually makes this hybrid model function looks something like this. AI handles clause detection, deviation-flagging against a defined playbook, consistency across large volume, and first-draft redlines on standard boilerplate language, while humans handle business context, negotiation strategy, judgment informed by the actual relationship with the counterparty, ambiguity that was written into the contract on purpose, and final sign-off. The Microsoft anecdote from earlier cuts in both directions here: AI does erase the kind of inconsistency that came from five paralegals doing the same job five different ways, but somebody still has to sit down and write the playbook that defines what "consistent" is even supposed to mean in the first place. The machine enforces the standard someone else designed.
How to allocate review work between AI and human reviewers given what the accuracy data shows
Everything above points toward a tiering system rather than a single up-or-down vote on AI. High-volume, low-complexity contracts, NDAs and standard vendor agreements, are where AI does a first pass and gets a light human check afterward; this is where the 90% to 95% time reduction shows up and where the accuracy claims hold their shape best. Moderately complex agreements get AI doing clause identification and deviation flagging, with human attention concentrated on the clauses that actually matter for negotiation; that's the 80% to 85% range described earlier, and it's arguably the model most organizations should be building toward. High-stakes, novel, or jurisdiction-sensitive contracts flip the order: human-led review, with AI running as a consistency check rather than the primary screener.
Three things determine whether an organization actually gets the accuracy numbers the benchmarks describe, rather than a disappointing imitation of them. Playbook quality comes first: AI is only ever as consistent as the rules it's fed, and the 60% of teams operating without one are, by definition, not positioned to trust whatever accuracy score their tool reports. Training period comes second: two to three months of feedback-based tuning is the reported norm before AI performs reliably at scale, which means the published accuracy figures describe a calibrated, mature deployment rather than day-one performance out of the box. Tool selection comes third, and it's not a small detail: purpose-built redlining platforms, trained specifically on legal-domain data with retrieval architecture designed for contract review, perform quite differently from general-purpose language models asked to moonlight as legal reviewers. The Gemini 1.5 Pro result from earlier (64% accuracy even with careful prompting) is the clearest illustration of that gap available.
What does meaningful return look like once these pieces are in place? Ironclad's 2025 research on enterprises processing 500 or more contracts a month found an average 51% reduction in legal review spend after full AI deployment. That number compounds because it's a volume game: modest time savings on a single contract barely register, but multiplied across hundreds of contracts a month, they add up fast, and a reported 39% reduction in overall contract lifecycle time reflects that compounding effect on deal velocity rather than just hours saved reading.
The right comparison was never really AI against human in some cage match. It's an AI-assisted team against a manual-only team, and the evidence keeps landing on the side of the hybrid. Teams that build a playbook, define contract tiers, and set clear review checkpoints before turning AI loose tend to realize the accuracy gains the benchmarks promise. Teams that bolt AI onto an undocumented, inconsistent review process mostly just get faster inconsistency, speed without the accuracy, and an attorney corps that starts distrusting the tool within a few bad redlines. The conclusion that follows is a little unglamorous but probably the most useful one in this entire piece: the bottleneck in most legal departments has less to do with the AI's capability and more to do with process design, the boring, unsexy work of deciding what "standard" means before asking a machine to enforce it.
