STUDY 05 / 30 · ANONYMISED · NDA · AI TOOLS · US · SEED

AI TOOLS · US Onboarding an AI product is teaching a mental model, not a feature tour.

RoleLead designer
Timeline8 weeks
Team1 designer, 1 PM, 1 frontend
VerticalAI research assistant · knowledge workers
ai tools
THE FAILURE

Activation was 9%. New users asked the product to do things it could not do, got a confident wrong answer, and left. The tour explained features; nobody had explained the boundary.

THE INTERVENTION

Designed the first session to teach where the product's competence ends. Three guided tasks - one it does brilliantly, one it does adequately, one it declines - with the decline as the deliberate centrepiece.

WHAT CHANGED

Activation more than tripled. The task that taught users what the product cannot do produced the largest single lift.

31%ACTIVATION @ SESSION 1from 9%
44%DAY-30 RETENTIONfrom 12%
-72%OUT-OF-SCOPE REQUESTS @ WK 2vs baseline

THE ARGUMENT

Why the obvious solution was wrong.

The study matters because the product problem was reframed before the interface was polished.

Every failed AI onboarding we have studied fails the same way: it demonstrates capability and conceals boundary, so the user constructs a mental model that is too generous, tests it within a day, and is disappointed by a system that was never claiming to do that. The disappointment is caused by the onboarding, not by the model. A tour that only shows wins is a tour that guarantees a loss on day two.

The first session runs three tasks in a fixed order. The first is chosen to succeed impressively and calibrate expectation upward. The second succeeds with visible effort - the user sees the model working at its limit. The third is designed to be declined, with the product explaining precisely why and what to use instead. Users who completed the third task retained at nearly twice the rate of users who skipped it. Teaching the edge of competence is the highest-leverage thing an AI onboarding can do, and almost nobody does it.

THE INTERFACE CRAFT

The interaction, rendered as a working product surface.

The specimen below is code-native and uses the study's own design logic. The client interface remains protected.

AI TOOLS
CONFIDENCE 82%

Action preview ready

DETAIL 01The declined task

Task three asks for something out of scope. The product declines clearly, explains the boundary, and suggests the correct tool. This is the centrepiece, not an error state.

DETAIL 02Effortful success made visible

Task two shows the model's working - retrieval, reasoning steps, revision - so users calibrate what 'hard' looks like for this product.

DETAIL 03Boundary card, permanently reachable

A single card listing what the product does and does not do, written plainly, reachable from every screen forever.

DESIGN DECISIONS

Positions we would defend.

Each decision names the principle and the product consequence, not a stylistic preference.

01

Teach the edge, not the centre

Users find the capable centre themselves. They discover the edge by falling off it, which is where churn happens.

02

Decline is a feature surface

A well-designed refusal builds more trust than an eighth demonstration of success.

03

No feature tour

Features are discoverable. Mental models are not. The session teaches the second and ignores the first.

PRODUCT LEADER READOUT

What transfers, and what should remain specific to this product.

A case study is useful when its operating principle travels without turning the original interface into a template.

01

Read the operating condition

For AI research assistant · knowledge workers, the transferable lesson is not a copied screen. It is the condition the interface had to make legible: Task three asks for something out of scope. The product declines clearly, explains the boundary, and suggests the correct tool. This is the centrepiece, not an error state. Rebuild that visibility for your own roles, risk, terminology, and operating cadence.

02

Protect the design rule

Users find the capable centre themselves. They discover the edge by falling off it, which is where churn happens. Keep that rule in the acceptance criteria, component states, and production QA record so later visual cleanup cannot erase why the interaction exists.

03

Measure behaviour after ship

The evidence record is 31% for activation @ session 1, from 9%. Recreate the baseline and outcome window before rollout, segment the result by role and context, and state clearly what the measure cannot prove.

RESEARCH RECORD

The work behind the interface.

These artefacts connect the final interaction back to the evidence and product model that produced it.

ARTEFACT 01

Failed-activation replay

Session recordings of 60 users who churned in week one; 71% had made an out-of-scope request in their first ten minutes.

ARTEFACT 02

Capability boundary map

Worked with the ML team to define the honest edge of competence across 30 task types, including the ambiguous middle.

ARTEFACT 03

Three-task sequencing test

Tested six orderings of the three tasks; success-effort-decline outperformed all alternatives on day-30 retention.

ARTEFACT 04

Decline copy study

Eleven versions of the refusal message tested for trust impact; specificity beat apology in every pairing.

“The task where it says no is the reason they stay. We would never have shipped that on our own.”

Founder, AI research tool · under NDA

NEXT STUDY · 06 / 30 · SALES AUTOMATION

SALES AUTOMATION · US A sequence builder that shows the human on the other end.

Read next study