How we measure accuracy

Initialed catches 31% more risky clauses than ChatGPT on long contracts

Head to head on the same commercial contracts, scored the same way. On documents over 30 pages — where the expensive clauses hide — the lead reaches 31%. Across all 30 contracts tested, 19% more.

Risky clauses caught

19% more

302 of 341 clauses vs 253 of 341

Initialed88.6%
ChatGPT74.2%

High-severity clauses caught

14% more

107 of 124 clauses vs 94 of 124

Initialed86.3%
ChatGPT75.8%

Contracts over 30 pages

31% more

across the 8 longest contracts tested

Initialed97.4%
ChatGPT74.6%

Contracts where it caught more

4 tied

out of 30 contracts

Initialed21
ChatGPT5

The longer the contract, the wider the gap

Initialed reviews your contract section by section, so it keeps up as documents get longer. A single chat prompt does not — it caught roughly the same share whether the contract was 5 pages or 50. On the longest contracts we tested, Initialed found 31% more.

Under 10 pages10 contracts
Initialed
82.1%
ChatGPT
75.0%
10 to 30 pages12 contracts
Initialed
85.3%
ChatGPT
73.4%
Over 30 pages8 contracts
Initialed
97.4%
ChatGPT
74.6%

Try it on a contract you actually care about

Your first review is free — no credit card, no plugin, any file.

How these numbers were measured

Measured August 2026 using Initialed’s production review pipeline against OpenAI’s then-current flagship model as the ChatGPT baseline, over 30 commercial contracts from the Contract Understanding Atticus Dataset (CUAD v1, The Atticus Project), used under CC BY 4.0 — a public research dataset in which lawyers annotated the clauses that matter. The answer key was written by those annotators before we ran anything, not chosen by us afterwards.

Percentages are clause-level: clauses found divided by clauses the annotators marked, pooled across all 30 contracts. “19% more” and “31% more” are relative — 302 clauses found against the baseline’s 253, and 97.4% against 74.6% on the eight contracts over 30 pages. A clause counts as found only when the quote the review points to overlaps the annotated clause; the same scoring code and threshold scored both sides. The baseline received the whole contract, the side being represented, and the same red-flag checklist Initialed uses — more guidance than a typical user provides. No contract was shortened to fit.

Scope and limits: this measures clause detection, not how often either tool raises an issue that turns out not to matter — we do not yet publish a figure for that, and Initialed reports substantially more findings per contract. CUAD records which clauses appear in a contract; which of those count as risky depends on which side you are on, and that judgment is ours. Both products change over time and neither returns identical output twice, so results will vary. Nothing here is legal advice or a guarantee about any individual contract.

ChatGPT and GPT are trademarks of OpenAI. Initialed is not affiliated with, endorsed by, or sponsored by OpenAI. The comparison reflects a single prompt sent to the model named above via its public API.