STUDY 02 / 30 · ANONYMISED · NDA · AI TOOLS · NETHERLANDS · SERIES A

AI TOOLS · NETHERLANDS Where the model is wrong, and how the interface says so.

RoleLead designer
Timeline10 weeks
Team1 designer, 1 PM, 1 ML engineer, 2 legal SMEs
VerticalContract review AI · legal teams at 90 corporates
ai tools
THE FAILURE

The model flagged risky clauses with uniform confidence. Lawyers found two wrong flags in the first session and stopped trusting all of them - including the correct ones, which were the majority.

THE INTERVENTION

Made model disagreement visible. Where the ensemble split, the interface says so and shows both readings, rather than resolving to a single confident answer the lawyer will disprove.

WHAT CHANGED

Review throughput rose because lawyers stopped re-reading clauses the model was confident about, and started spending their attention where it had said it was unsure.

31CLAUSES REVIEWED / HOURfrom 18
6%FLAGS DISMISSED AS WRONGfrom 34%
12%HIGH-CONFIDENCE FLAGS RE-READfrom 91%

THE ARGUMENT

Why the obvious solution was wrong.

The study matters because the product problem was reframed before the interface was polished.

A professional will forgive a tool that is uncertain and will not forgive one that is confidently wrong. The product was doing the second thing structurally: an ensemble of models with genuine internal disagreement was being collapsed into a single flag with a single colour. When the collapsed answer was wrong - as it must sometimes be - the lawyer had no way to know it had been a close call, so they generalised the failure to everything.

We surfaced the split. Where model agreement is high, the clause carries a solid marker and a compact rationale. Where it is contested, the marker is drawn open and both readings are presented side by side, each with its supporting passage from the contract. The lawyer resolves it, and the resolution is logged against the clause type - which is now the product's most valuable training data. The interface stopped pretending to a certainty it did not have, and the users started believing the certainty it did.

THE INTERFACE CRAFT

The interaction, rendered as a working product surface.

The specimen below is code-native and uses the study's own design logic. The client interface remains protected.

AI TOOLS
CONFIDENCE 82%

Action preview ready

DETAIL 01Split markers for contested clauses

An open marker where the ensemble disagreed, showing both readings with their supporting contract passages, rather than a resolved single answer.

DETAIL 02Rationale on hover, not on click

The reason a clause was flagged is one hover away at all times. Clicking through to a rationale is a cost lawyers will not pay at volume.

DETAIL 03Disagreement log as training surface

Every lawyer resolution of a contested clause writes back to a labelled dataset, visible to the client's own legal ops lead.

DESIGN DECISIONS

Positions we would defend.

Each decision names the principle and the product consequence, not a stylistic preference.

01

Never collapse a genuine split

A confident wrong answer costs more trust than an honest uncertain one, and the cost compounds across the whole product.

02

Rationale is ambient

Reasoning that requires a click is reasoning that does not exist at professional working speed.

03

The reviewer is the label

Design the correction path as the primary data-collection mechanism, not as an error state.

PRODUCT LEADER READOUT

What transfers, and what should remain specific to this product.

A case study is useful when its operating principle travels without turning the original interface into a template.

01

Read the operating condition

For Contract review AI · legal teams at 90 corporates, the transferable lesson is not a copied screen. It is the condition the interface had to make legible: An open marker where the ensemble disagreed, showing both readings with their supporting contract passages, rather than a resolved single answer. Rebuild that visibility for your own roles, risk, terminology, and operating cadence.

02

Protect the design rule

A confident wrong answer costs more trust than an honest uncertain one, and the cost compounds across the whole product. Keep that rule in the acceptance criteria, component states, and production QA record so later visual cleanup cannot erase why the interaction exists.

03

Measure behaviour after ship

The evidence record is 31 for clauses reviewed / hour, from 18. Recreate the baseline and outcome window before rollout, segment the result by role and context, and state clearly what the measure cannot prove.

RESEARCH RECORD

The work behind the interface.

These artefacts connect the final interaction back to the evidence and product model that produced it.

ARTEFACT 01

Clause-risk grammar

Built a shared vocabulary of 22 risk types with the client's legal team, replacing model-internal categories nobody outside ML understood.

ARTEFACT 02

Ensemble disagreement analysis

Mapped where the model ensemble split across 4,000 historical clauses; disagreement clustered in 6 clause types.

ARTEFACT 03

Wrong-flag post-mortem

Sat with three lawyers replaying the exact sessions where trust broke; both failures were contested clauses shown as certain.

ARTEFACT 04

Throughput baseline

Timed 40 review sessions pre-redesign to establish clauses-per-hour and re-read rate as the measurable metrics.

“It tells me when it isn't sure. That's the only reason I believe it when it is.”

Senior Counsel, listed manufacturer · under NDA

NEXT STUDY · 03 / 30 · AI TOOLS

AI TOOLS · INDIA Retrieval you can audit, for a regulator who will ask.

Read next study