The model flagged risky clauses with uniform confidence. Lawyers found two wrong flags in the first session and stopped trusting all of them - including the correct ones, which were the majority.
STUDY 02 / 30 · ANONYMISED · NDA · AI TOOLS · NETHERLANDS · SERIES A
AI TOOLS · NETHERLANDS Where the model is wrong, and how the interface says so.

Made model disagreement visible. Where the ensemble split, the interface says so and shows both readings, rather than resolving to a single confident answer the lawyer will disprove.
Review throughput rose because lawyers stopped re-reading clauses the model was confident about, and started spending their attention where it had said it was unsure.
THE ARGUMENT
Why the obvious solution was wrong.
The study matters because the product problem was reframed before the interface was polished.
A professional will forgive a tool that is uncertain and will not forgive one that is confidently wrong. The product was doing the second thing structurally: an ensemble of models with genuine internal disagreement was being collapsed into a single flag with a single colour. When the collapsed answer was wrong - as it must sometimes be - the lawyer had no way to know it had been a close call, so they generalised the failure to everything.
We surfaced the split. Where model agreement is high, the clause carries a solid marker and a compact rationale. Where it is contested, the marker is drawn open and both readings are presented side by side, each with its supporting passage from the contract. The lawyer resolves it, and the resolution is logged against the clause type - which is now the product's most valuable training data. The interface stopped pretending to a certainty it did not have, and the users started believing the certainty it did.
THE INTERFACE CRAFT
The interaction, rendered as a working product surface.
The specimen below is code-native and uses the study's own design logic. The client interface remains protected.
Action preview ready
DESIGN DECISIONS
Positions we would defend.
Each decision names the principle and the product consequence, not a stylistic preference.
Never collapse a genuine split
A confident wrong answer costs more trust than an honest uncertain one, and the cost compounds across the whole product.
Rationale is ambient
Reasoning that requires a click is reasoning that does not exist at professional working speed.
The reviewer is the label
Design the correction path as the primary data-collection mechanism, not as an error state.
PRODUCT LEADER READOUT
What transfers, and what should remain specific to this product.
A case study is useful when its operating principle travels without turning the original interface into a template.
Read the operating condition
For Contract review AI · legal teams at 90 corporates, the transferable lesson is not a copied screen. It is the condition the interface had to make legible: An open marker where the ensemble disagreed, showing both readings with their supporting contract passages, rather than a resolved single answer. Rebuild that visibility for your own roles, risk, terminology, and operating cadence.
Protect the design rule
A confident wrong answer costs more trust than an honest uncertain one, and the cost compounds across the whole product. Keep that rule in the acceptance criteria, component states, and production QA record so later visual cleanup cannot erase why the interaction exists.
Measure behaviour after ship
The evidence record is 31 for clauses reviewed / hour, from 18. Recreate the baseline and outcome window before rollout, segment the result by role and context, and state clearly what the measure cannot prove.
RESEARCH RECORD
The work behind the interface.
These artefacts connect the final interaction back to the evidence and product model that produced it.
Clause-risk grammar
Built a shared vocabulary of 22 risk types with the client's legal team, replacing model-internal categories nobody outside ML understood.
Ensemble disagreement analysis
Mapped where the model ensemble split across 4,000 historical clauses; disagreement clustered in 6 clause types.
Wrong-flag post-mortem
Sat with three lawyers replaying the exact sessions where trust broke; both failures were contested clauses shown as certain.
Throughput baseline
Timed 40 review sessions pre-redesign to establish clauses-per-hour and re-read rate as the measurable metrics.
“It tells me when it isn't sure. That's the only reason I believe it when it is.”
Senior Counsel, listed manufacturer · under NDA
NEXT STUDY · 03 / 30 · AI TOOLS