STUDY 01 / 30 · ANONYMISED · NDA · AI TOOLS · US · SERIES B

AI TOOLS · US An AI agent that shows its work before it acts, not after.

RoleLead designer + interaction architect
Timeline14 weeks
Team2 designers, 1 PM, 2 ML engineers, 1 frontend
VerticalAutonomous GTM agent · 340 enterprise seats
ai tools
THE FAILURE

The agent could send email, update CRM records and book meetings. Users disabled it within a week because they could not see what it was about to do - only what it had already done.

THE INTERVENTION

Inverted the interaction: every autonomous action renders as a reviewable draft with a countdown, a diff against current state, and a one-key hold. Trust is earned before execution, not apologised for after.

WHAT CHANGED

Agent stayed enabled past week four for the first time. Users granted broader permissions voluntarily as confidence accumulated.

78%AGENT RETAINED @ WK 4from 11%
94%ACTIONS REVIEWED BEFORE SENDby design
3.4xPERMISSION SCOPE GRANTEDvoluntary escalation

THE ARGUMENT

Why the obvious solution was wrong.

The study matters because the product problem was reframed before the interface was polished.

Every agent product we audited made the same wager - that autonomy is the feature and visibility is friction. The wager loses. A sales operator whose job depends on what lands in a customer's inbox will not delegate to a system whose next move is invisible, however good its last move was. The disabling was not a trust failure in the model. It was an absence of any surface on which trust could be built.

We built the preview as the primary object and execution as its consequence. Each queued action shows a structural diff - this field changes from this to that, this email goes to this person with this text - and a fifteen-second hold. Holding is one keypress and reversible; letting it run requires nothing. The asymmetry is deliberate. Over four weeks the median user's review rate fell naturally from every action to roughly one in six, and the permission scope they granted rose. They were not trained into trust. They were shown enough to build it.

THE INTERFACE CRAFT

The interaction, rendered as a working product surface.

The specimen below is code-native and uses the study's own design logic. The client interface remains protected.

AI TOOLS
CONFIDENCE 82%

Action preview ready

DETAIL 01Action preview with structural diff

Queued actions render the before and after state of every field they touch, with changed values highlighted. No natural-language summary of what will happen - the actual change, rendered.

DETAIL 02Fifteen-second hold, one key

Space holds an action indefinitely. Nothing is required to let it proceed. The cheapest gesture is the safe one, which inverts the usual arrangement.

DETAIL 03Permission ladder, not a toggle

Six graduated scopes rather than on/off. Users climb the ladder as confidence accrues, and the interface shows what the next rung would have done last week.

DESIGN DECISIONS

Positions we would defend.

Each decision names the principle and the product consequence, not a stylistic preference.

01

Preview is the product

Execution is a consequence of the preview, not the other way around. The queue is the primary screen, not a settings page.

02

Diff over summary

A generated summary of a change is another thing to verify. The change itself is not.

03

Undo is not enough

Reversibility after send does not restore a customer relationship. The interface intervenes before, not after.

PRODUCT LEADER READOUT

What transfers, and what should remain specific to this product.

A case study is useful when its operating principle travels without turning the original interface into a template.

01

Read the operating condition

For Autonomous GTM agent · 340 enterprise seats, the transferable lesson is not a copied screen. It is the condition the interface had to make legible: Queued actions render the before and after state of every field they touch, with changed values highlighted. No natural-language summary of what will happen - the actual change, rendered. Rebuild that visibility for your own roles, risk, terminology, and operating cadence.

02

Protect the design rule

Execution is a consequence of the preview, not the other way around. The queue is the primary screen, not a settings page. Keep that rule in the acceptance criteria, component states, and production QA record so later visual cleanup cannot erase why the interaction exists.

03

Measure behaviour after ship

The evidence record is 78% for agent retained @ wk 4, from 11%. Recreate the baseline and outcome window before rollout, segment the result by role and context, and state clearly what the measure cannot prove.

RESEARCH RECORD

The work behind the interface.

These artefacts connect the final interaction back to the evidence and product model that produced it.

ARTEFACT 01

Agent action taxonomy

Classified 41 agent action types by reversibility and blast radius; four required hard confirmation regardless of permission level.

ARTEFACT 02

Trust-decay study

Diary study across 12 operators tracking the exact moment they disabled the agent and what preceded it.

ARTEFACT 03

Hold-timing calibration

Tested 5s, 15s, 30s and manual holds against both error interception and perceived friction. Fifteen won on both.

ARTEFACT 04

Permission ladder prototype

Six-rung scope model tested against a binary toggle for voluntary escalation over four weeks.

“They stopped turning it off. That was the entire product problem and nobody had framed it as an interface problem.”

VP Product, GTM automation platform · under NDA

NEXT STUDY · 02 / 30 · AI TOOLS

AI TOOLS · NETHERLANDS Where the model is wrong, and how the interface says so.

Read next study