Cross-functional evaluation panels scored suppliers in isolated spreadsheets. Scores were averaged, divergence was invisible, and award decisions could not be justified when challenged.
STUDY 29 / 30 · ANONYMISED · NDA · PROCUREMENT · MUNICH · INDUSTRIAL GROUP
PROCUREMENT · MUNICH Supplier evaluation where the scoring is visible to everyone scoring.

Made scoring collaborative and divergence explicit - where evaluators disagreed materially, the system surfaced it and required resolution before award.
Award challenges fell and the justification pack generated automatically, converting weeks of retrospective assembly into an artefact produced by the process itself.
THE ARGUMENT
Why the obvious solution was wrong.
The study matters because the product problem was reframed before the interface was polished.
Averaging hides the only information that matters. When engineering scores a supplier 9 on technical capability and quality scores the same supplier 3, the average of 6 is a number that describes no one's view and conceals a disagreement that will surface after award as a failed delivery. Every panel process we examined destroyed its most valuable signal at the moment of aggregation.
Scores are entered independently - anchoring is real - and then divergence is surfaced before any average is computed. Where two evaluators differ by more than a threshold on a weighted criterion, the system requires a recorded resolution: one moves, or both record why they hold. That resolution becomes part of the award justification, which now assembles automatically from the process rather than being reconstructed weeks later by someone who was not in the room. Fourteen months, no upheld challenge.
THE INTERFACE CRAFT
The interaction, rendered as a working product surface.
The specimen below is code-native and uses the study's own design logic. The client interface remains protected.
DESIGN DECISIONS
Positions we would defend.
Each decision names the principle and the product consequence, not a stylistic preference.
Never average away a disagreement
The divergence between two expert evaluators is more informative than their mean, and aggregation destroys it permanently.
Anchoring is real
Visible scores converge for social reasons rather than evidential ones. Independent entry first is a methodological requirement.
The record is the defence
An award justification assembled retrospectively is weaker than one produced by the process. Design the artefact into the workflow.
PRODUCT LEADER READOUT
What transfers, and what should remain specific to this product.
A case study is useful when its operating principle travels without turning the original interface into a template.
Read the operating condition
For Strategic sourcing platform · €2.4B annual spend, the transferable lesson is not a copied screen. It is the condition the interface had to make legible: Scores entered without visibility of others to avoid anchoring, then material disagreement surfaced before any aggregation. Rebuild that visibility for your own roles, risk, terminology, and operating cadence.
Protect the design rule
The divergence between two expert evaluators is more informative than their mean, and aggregation destroys it permanently. Keep that rule in the acceptance criteria, component states, and production QA record so later visual cleanup cannot erase why the interaction exists.
Measure behaviour after ship
The evidence record is 0 for award challenges upheld, in 14 months. Recreate the baseline and outcome window before rollout, segment the result by role and context, and state clearly what the measure cannot prove.
RESEARCH RECORD
The work behind the interface.
These artefacts connect the final interaction back to the evidence and product model that produced it.
Panel observation
Observed nine evaluation panels across three categories, recording where disagreement was raised and how it was resolved or buried.
Challenge post-mortem
Reviewed four historical award challenges; three turned on divergence that had been averaged away and never recorded.
Weighting transparency test
Tested whether showing criterion weights during scoring biased results; it did not, and it improved score rationale quality.
Justification pack review
Ran the generated pack past internal legal and an external procurement counsel for challenge-resistance.
“The disagreement between engineering and quality used to disappear into an average. Now it has to be resolved on the record.”
Head of Strategic Sourcing, industrial group · under NDA
NEXT STUDY · 30 / 30 · PROCUREMENT