There is a specific moment every design leader has been in. You say design drove the improvement. Someone from another function says, politely, that a lot of things changed that quarter. The room moves on, and your claim is now worth slightly less than if you had never made it.
The instinct afterwards is to find better numbers. Usually the problem is not the numbers, it is the shape of the claim.
Why causal claims lose
"Design drove the 12 percent lift" is a claim about causation, and in almost every real case it is unprovable. Marketing ran a campaign. Pricing changed. Engineering fixed the loading time. Seasonality exists. Whoever is listening knows all of this, so the claim gets discounted, and rightly.
Worse, it is a competitive claim. Every function in the building is pointing at the same lift. When four teams claim the same outcome, the room stops believing any of them and falls back to whoever has the most political capital. That is a game design usually loses.
The trap is that the more confident the claim, the more it invites the challenge. "Design was responsible for" is doing more work than the evidence supports, and everyone can feel it.
Trace the path instead
The claim that survives is smaller and structural. It has four parts, in order:
This priority the company already said mattered.
These decisions were made in service of it, on these dates.
These people shaped those decisions, and here is what each contributed.
This changed, measured on a metric the business already tracks.
Notice what is missing. There is no assertion that design caused the change. What you have shown is a documented chain from a stated priority through specific decisions to a measured outcome, with the design contribution located inside it.
This is a weaker claim and a far more useful one. It cannot be attacked on attribution grounds because it does not claim sole attribution. What it demonstrates is that design was operating on the things the company said were important, in a traceable way, and that those things moved. Over three or four quarters, a pattern of that is more persuasive than any single causal claim, because it shows a function that is aimed correctly and can account for itself.
Pick metrics the business already runs on
A recurring self-inflicted wound is inventing a design-specific metric and asking leadership to care about it.
Design quality scores, system adoption percentages, usability indices. Each is defensible internally and each converts every conversation into an argument about the metric rather than the work. You end up defending your instrument instead of your contribution, and you look like you brought your own scoreboard.
Use what the company already runs on. If it runs on activation, speak in activation. If it runs on gross margin, find the path to gross margin. Yes, translation is harder, and yes, some design work maps awkwardly. That difficulty is the job. A metric nobody outside the function recognizes cannot function as evidence, however rigorous it is.
There is one exception worth protecting. Internal operating measures like cycle time and rework rate are for running the organization, and they belong in your own reviews. Just do not present them as impact. They describe how well the machine runs, not what the machine produced.
The chain breaks in a specific place
Here is why this is harder than it sounds.
Each link exists somewhere. The priority is in a planning doc. The decisions are in tickets and files and threads. The people are on those artifacts. The outcome is in the analytics tool.
What does not exist is the connections between them. Figma knows who edited a file but not which initiative it served. Jira knows the epic but not the research that shaped it. The analytics tool knows the metric moved but nothing about what shipped into it. Each system holds one link of the chain and no system holds the chain.
So building it means a human sits down and reconstructs the whole thing from memory, once a quarter, in the week before the review. That reconstruction is lossy in a predictable direction: it captures what was memorable and recent, and loses the contribution of anyone who was not in the room when you were remembering.
This is the actual reason design impact goes unproven. Not that the work is intangible. The evidence is real, dated and attributable, and it is unlinked.
Start with one initiative
Do not try to instrument everything. Pick one initiative from last quarter and trace it end to end.
Write down the priority it served, in the company's own language. List the decisions that shipped against it, with dates. For each, note who shaped it, including people who did not own the file. Find the metric it was meant to move and what it actually did. Then write the four-part chain in a paragraph.
Two things usually happen. The first is that the chain has a hole, most often between decision and outcome, because nobody wrote down what the decision was supposed to move. That hole is the finding. It means the decision shipped without a stated expectation, which is worth fixing regardless of measurement.
The second is that reconstructing one initiative takes half a day. Multiply by everything your organization shipped and you have the reason this does not happen.
Include the ones that did not work
The strongest thing you can put in front of leadership is a decision that shipped, did not move what it was supposed to, and taught you something.
This costs less than it feels like it will. It also does something no amount of positive reporting achieves: it demonstrates you are measuring rather than selling. A leader presenting five wins reads as advocacy, because they chose the five. A leader presenting four wins and a miss reads as someone with an instrument.
Design has spent a long time arguing for its value. McKinsey's design index work found top-quartile design performers achieving 32 percentage points higher revenue growth than industry peers over five years, and that finding gets cited constantly in exactly these conversations. It is a good number and it is somebody else's evidence. The organizations that benefit from it are the ones that can show the same shape of thing about themselves.
That takes a traceable chain, kept as the work happens rather than reconstructed once a quarter from what anyone can still remember.