Twice a year, a group of design managers sit in a room and decide what a year of someone's work was worth. They do it from memory, under time pressure, arguing from whichever examples happened to stick.
Everyone involved knows this is a bad instrument. The people who study performance calibration describe the failure mode without much diplomacy: sessions run on memory and narrative rather than structured data, and the conversation defaults to advocacy, which rewards proximity and visibility rather than actual contribution.
That sentence should be uncomfortable for anyone who runs a design organization, because it describes a mechanism, not an accident. And design has it worse than most functions.
Why design has it worse
Three things compound.
The best work is the least visible. A designer who prevents six months of wasted engineering by killing a bad direction in week two has done the most valuable thing anyone on the team did that quarter. There is no artifact. There is a conversation, in a critique, that three people remember and nobody wrote down. Meanwhile a designer who shipped four mediocre screens has four screens to point at.
Contribution is distributed and design work is joint. Almost nothing in a mature design organization has a single author. A flow is shaped by a researcher's finding, a systems designer's component, a critique from someone on another team, and a content designer who fixed the thing that was actually confusing. When the outcome lands, the record shows one name: whoever owned the file.
Calibration is cross-functional and design is the odd one out. Comparing designers to designers is hard enough. Comparing a design org against product and engineering in the same session, using ratings that are supposed to mean the same thing, is harder still, because the other functions arrive with delivery data and design arrives with a story.
The result is predictable. The designers who do well in calibration are disproportionately the ones whose managers are good advocates, who work on visible surfaces, and who sit close to leadership. That is not a moral failure of anyone in the room. It is what happens when the only available input is recall.
The evidence exists, it just is not collected
Here is what makes this frustrating rather than tragic: the record of contribution is generated continuously, and then discarded.
Someone shaped the direction in a critique. That critique happened, in a thread or a call or a Figma comment, with a timestamp and participants. Someone's research changed a roadmap. That study exists, in a repository, and so does the decision that cites it. Someone built a component four teams now use. The component exists, and so does every file that consumes it. Someone mentored a junior into a promotion. That shows up as review comments, pairing, and a visible change in another person's output.
Every one of those is evidence. All of it is real, dated and attributable. None of it is linked, so none of it is available six months later when it matters. By calibration season the only surviving copy is in somebody's head, and heads are lossy in a specific direction: they retain what was recent, what was visible, and what belonged to people they talked to often.
What a better record looks like
The goal is not a productivity score. It is worth being emphatic about that, because the moment a design organization starts ranking people by output volume it has recreated the artifact-counting problem with worse consequences: it will systematically promote the people doing the least valuable work.
A useful contribution record has three properties.
It is evidenced, not inferred. Every entry traces to a specific thing that happened: this critique, this study, this component, this decision. A manager should be able to click through from a claim to the artifact behind it. This is the property that makes the record survive an argument.
It accumulates continuously. The record is built as the work happens, not assembled in the two weeks before review season. This is the entire ballgame. A record compiled retroactively is just memory with extra steps, and it inherits every bias memory has.
It captures influence, not just ownership. The distinction between "who owned this file" and "who shaped this outcome" is where most of design's real contribution lives. A record that only knows ownership will keep rewarding the people whose names are on things.
What this changes in the room
A calibration session with an evidenced record does not become mechanical, and it should not. Judgment is still the point. What changes is what judgment is applied to.
Instead of "I think Priya had a strong year, she did the checkout work," the conversation starts from a record showing that Priya shaped direction on three initiatives she did not own, that her research changed the roadmap for a team she is not on, and that the component she built is consumed by four squads. The manager still has to interpret that. But they are interpreting evidence rather than substituting for it.
The second-order effect matters more. When contribution is visible as it accrues, it stops being a twice-yearly performance and becomes a continuous fact. Managers coach against it in one-on-ones. Promotion cases get built over months rather than assembled in a panic. People stop having to self-promote to be seen, which disproportionately helps the ones who are bad at self-promotion and good at the job.
The thing not to build
A closing caution, because this is the failure mode that will discredit the whole idea.
Do not build individual surveillance. The unit of analysis is how the organization works, and contribution is context for recognition, growth and planning. It is not a leaderboard, it is not a stack rank, and it is not a productivity metric with a designer's name on it. The moment your team believes it is being scored, the record stops describing the work and starts describing what people think is being counted, which is the same way time tracking dies.
The instrument is there to make good work visible. If it does anything else, it has failed at the only thing worth doing.