Managers wrote reviews from memory in the last week of the cycle. Ratings clustered meaninglessly, calibration meetings ran six hours, and employees experienced the outcome as arbitrary.
STUDY 17 / 30 · ANONYMISED · NDA · ENTERPRISE INTERNAL · INDIA · 18,000 EMPLOYEES
ENTERPRISE INTERNAL · INDIA Performance review software that produces a conversation.

Made evidence accumulation continuous and low-cost across the year, so the review composes from a record rather than from recall, and calibration argues from artefacts.
Rating distribution spread meaningfully for the first time. Calibration halved in length because disagreements had evidence attached.
THE ARGUMENT
Why the obvious solution was wrong.
The study matters because the product problem was reframed before the interface was polished.
Annual review tools are built as a form to be completed in a window, which guarantees the content is recall-based. A manager with fourteen reports writing in the final week is reconstructing a year from what they can remember, and what they remember is the last six weeks and any incident. The rating clustering everyone complains about is not managerial laziness. It is the predictable output of an interface that only exists for two weeks a year.
We made capture continuous and nearly free: a manager can attach a note to a person from Slack, from a ticket, from a document, in one action, and the note carries its source. Over a year that produces a record. The review then composes from the record with evidence attached to each claim, and calibration meetings argue from artefacts rather than impressions. Meeting length halved because the disagreements that used to consume hours - I see them differently - now resolve against a shared record in minutes.
THE INTERFACE CRAFT
The interaction, rendered as a working product surface.
The specimen below is code-native and uses the study's own design logic. The client interface remains protected.
DESIGN DECISIONS
Positions we would defend.
Each decision names the principle and the product consequence, not a stylistic preference.
Capture must be nearly free
Any capture mechanism costing more than one action will not be used, and the review will be recall-based regardless of the form's design.
Make the unsupported claim visible
Not by blocking it, but by rendering it plainly as unsupported. Social pressure does the rest in a calibration room.
Design for the year, not the window
A tool that exists for two weeks produces two weeks of thinking. Continuous capture is the whole intervention.
PRODUCT LEADER READOUT
What transfers, and what should remain specific to this product.
A case study is useful when its operating principle travels without turning the original interface into a template.
Read the operating condition
For HR platform · IT services firm, the transferable lesson is not a copied screen. It is the condition the interface had to make legible: Attach an observation to a person from Slack, a ticket or a document in a single action. The note keeps its source and timestamp. Rebuild that visibility for your own roles, risk, terminology, and operating cadence.
Protect the design rule
Any capture mechanism costing more than one action will not be used, and the review will be recall-based regardless of the form's design. Keep that rule in the acceptance criteria, component states, and production QA record so later visual cleanup cannot erase why the interaction exists.
Measure behaviour after ship
The evidence record is 91% for reviews citing evidence, from 12%. Recreate the baseline and outcome window before rollout, segment the result by role and context, and state clearly what the measure cannot prove.
RESEARCH RECORD
The work behind the interface.
These artefacts connect the final interaction back to the evidence and product model that produced it.
Review corpus analysis
Analysed 2,400 prior reviews for recency bias; 68% of cited incidents fell in the final quarter of the period.
Manager time study
Timed the review-writing burden at 14 hours per manager per cycle, concentrated entirely in one week.
Capture friction test
Tested five capture mechanisms; anything above one action fell below 10% sustained usage within a month.
Calibration observation
Sat through six calibration meetings recording how disagreements were resolved; 71% resolved by seniority, not evidence.
“The calibration meeting used to be about who argued hardest. Now it is about what is written down.”
Head of Talent, IT services · under NDA
NEXT STUDY · 18 / 30 · ENTERPRISE INTERNAL