STUDY 17 / 30 · ANONYMISED · NDA · ENTERPRISE INTERNAL · INDIA · 18,000 EMPLOYEES

ENTERPRISE INTERNAL · INDIA Performance review software that produces a conversation.

RoleLead designer
Timeline12 weeks
Team2 designers, 1 researcher, 1 PM, 2 frontend
VerticalHR platform · IT services firm
enterprise internal
THE FAILURE

Managers wrote reviews from memory in the last week of the cycle. Ratings clustered meaninglessly, calibration meetings ran six hours, and employees experienced the outcome as arbitrary.

THE INTERVENTION

Made evidence accumulation continuous and low-cost across the year, so the review composes from a record rather than from recall, and calibration argues from artefacts.

WHAT CHANGED

Rating distribution spread meaningfully for the first time. Calibration halved in length because disagreements had evidence attached.

91%REVIEWS CITING EVIDENCEfrom 12%
2.8 hrsCALIBRATION MEETING LENGTHfrom 6.2
3.9 / 5EMPLOYEE-PERCEIVED FAIRNESSfrom 2.2

THE ARGUMENT

Why the obvious solution was wrong.

The study matters because the product problem was reframed before the interface was polished.

Annual review tools are built as a form to be completed in a window, which guarantees the content is recall-based. A manager with fourteen reports writing in the final week is reconstructing a year from what they can remember, and what they remember is the last six weeks and any incident. The rating clustering everyone complains about is not managerial laziness. It is the predictable output of an interface that only exists for two weeks a year.

We made capture continuous and nearly free: a manager can attach a note to a person from Slack, from a ticket, from a document, in one action, and the note carries its source. Over a year that produces a record. The review then composes from the record with evidence attached to each claim, and calibration meetings argue from artefacts rather than impressions. Meeting length halved because the disagreements that used to consume hours - I see them differently - now resolve against a shared record in minutes.

THE INTERFACE CRAFT

The interaction, rendered as a working product surface.

The specimen below is code-native and uses the study's own design logic. The client interface remains protected.

ENTERPRISE INTERNAL
SHIFT / SATURDAY
08:00Coverage 92%
16:002 open
00:00Covered
DETAIL 01One-action capture, anywhere

Attach an observation to a person from Slack, a ticket or a document in a single action. The note keeps its source and timestamp.

DETAIL 02Claims carry evidence

Every statement in a review links to the observations that support it. An unsupported claim is visibly unsupported.

DETAIL 03Calibration argues from artefacts

The calibration console shows evidence density and distribution side by side, so disagreement resolves against a record.

DESIGN DECISIONS

Positions we would defend.

Each decision names the principle and the product consequence, not a stylistic preference.

01

Capture must be nearly free

Any capture mechanism costing more than one action will not be used, and the review will be recall-based regardless of the form's design.

02

Make the unsupported claim visible

Not by blocking it, but by rendering it plainly as unsupported. Social pressure does the rest in a calibration room.

03

Design for the year, not the window

A tool that exists for two weeks produces two weeks of thinking. Continuous capture is the whole intervention.

PRODUCT LEADER READOUT

What transfers, and what should remain specific to this product.

A case study is useful when its operating principle travels without turning the original interface into a template.

01

Read the operating condition

For HR platform · IT services firm, the transferable lesson is not a copied screen. It is the condition the interface had to make legible: Attach an observation to a person from Slack, a ticket or a document in a single action. The note keeps its source and timestamp. Rebuild that visibility for your own roles, risk, terminology, and operating cadence.

02

Protect the design rule

Any capture mechanism costing more than one action will not be used, and the review will be recall-based regardless of the form's design. Keep that rule in the acceptance criteria, component states, and production QA record so later visual cleanup cannot erase why the interaction exists.

03

Measure behaviour after ship

The evidence record is 91% for reviews citing evidence, from 12%. Recreate the baseline and outcome window before rollout, segment the result by role and context, and state clearly what the measure cannot prove.

RESEARCH RECORD

The work behind the interface.

These artefacts connect the final interaction back to the evidence and product model that produced it.

ARTEFACT 01

Review corpus analysis

Analysed 2,400 prior reviews for recency bias; 68% of cited incidents fell in the final quarter of the period.

ARTEFACT 02

Manager time study

Timed the review-writing burden at 14 hours per manager per cycle, concentrated entirely in one week.

ARTEFACT 03

Capture friction test

Tested five capture mechanisms; anything above one action fell below 10% sustained usage within a month.

ARTEFACT 04

Calibration observation

Sat through six calibration meetings recording how disagreements were resolved; 71% resolved by seniority, not evidence.

“The calibration meeting used to be about who argued hardest. Now it is about what is written down.”

Head of Talent, IT services · under NDA

NEXT STUDY · 18 / 30 · ENTERPRISE INTERNAL

ENTERPRISE INTERNAL · NETHERLANDS Access requests that expire by default.

Read next study