The Measurement Trap
An essay on how organizations confuse measuring performance with managing it — and why the metrics that feel most actionable are often the ones that obscure the most.
Organizations that measure a lot often know less than they think. The dashboards are full. The KPIs are tracked. The weekly business reviews move through a consistent set of numbers with a consistent set of commentary. And yet, when something goes wrong — when a key account churns, when a launch underperforms, when a team loses momentum — the data did not predict it. The metrics looked fine until they did not, and by then the damage was done.
This is not a data quality problem. The data was accurate. It is a measurement design problem. The organization was measuring the wrong things with confidence — building its operating rhythm around signals that were legible, trackable, and largely disconnected from the outcomes it was trying to produce.
Why organizations measure what they measure
Measurement is not neutral. The things organizations measure reflect what they believed mattered when they built the measurement system — which is usually a combination of what was technically easy to track, what the previous leadership cared about, and what showed up in a benchmark or industry report. These forces compound over time into a measurement infrastructure that outlives the conditions that created it.
This is how organizations end up with dashboards full of activity metrics — calls made, tickets closed, emails sent, features shipped — while the outcomes those activities are supposed to produce remain opaque or are measured with a significant lag. Activity is measurable in real time. Outcomes take longer, involve more variables, and are harder to attribute cleanly. So organizations measure activity and report on it as though it were a proxy for outcomes. Sometimes it is. Often it is not.
The substitution is rarely deliberate. It accumulates through the path of least resistance. The CRM captures calls. The ticketing system captures resolutions. The release log captures deployments. Building dashboards from these systems is straightforward. Building a measurement system that captures whether the right accounts are being called, whether tickets are being resolved in ways that prevent recurrence, whether releases are actually improving the metrics they were designed to move — that requires design effort most organizations defer.
The legibility trap
The most dangerous metrics are not the ones that are obviously wrong. They are the ones that are obviously right. A metric that is clean, timely, and easy to act on gets treated as authoritative. Over time, the metric stops being a proxy for the outcome and becomes the outcome itself. Managing to the metric replaces managing toward the underlying result.
This dynamic has a name — Goodhart's Law — but naming it does not prevent it. It emerges from a rational incentive structure. Leaders want clarity. Managers want to demonstrate progress. Teams want to know what winning looks like. A legible metric provides all of this. The pressure to perform against the metric produces behavior that is optimized for the metric. If the metric is well-designed, this optimization produces better outcomes. If the metric is not well-designed — if it is a proxy that diverges from the underlying reality under optimization pressure — the performance is real and the outcome is not.
Net Promoter Score is one well-documented example: teams trained to ask for high scores at the right moment produce high scores that do not predict retention. Sales activity metrics are another: reps optimized for call volume produce call volume that does not produce pipeline. The pattern is consistent across functions. Any metric that is made visible and consequential will be managed. The question is whether managing it produces the thing that was supposed to matter.
Leading indicators that do not lead
The standard response to lagging metrics is to add leading ones. If revenue is a lagging indicator, measure pipeline. If pipeline is lagging, measure activities. If activities are lagging, measure inputs to the activities. The logic is that measuring earlier in the causal chain gives earlier warning and more time to intervene.
This logic is correct but requires a condition that is rarely verified: the causal chain must actually connect. The leading indicator must actually lead to the outcome in the specific context the organization operates in. If the causal link is assumed rather than validated, leading indicators produce the same problem as lagging ones — with an additional cost. Leaders believe they are ahead of the problem because they are watching the leading metric. They are not ahead of the problem. They are watching a metric that feels predictive but is not.
Validating a causal chain requires historical data, analytical capacity, and a willingness to discover that the assumed connection does not hold. These requirements are frequently skipped because the alternative — accepting that we do not know what drives outcomes — is uncomfortable. A leading indicator that is declared without validation is not a forecasting tool. It is a confidence mechanism.
The design requirements for a useful measurement system
A measurement system is useful to the extent that it changes decisions. This is the test that most measurement systems fail — not because the data is wrong, but because the path from the data to a different decision is never designed.
Useful measurement starts with the decision, not the data. What is the governing decision this metric is meant to inform? Who makes it? At what frequency? What threshold would change the decision, and what change would follow? Without answers to these questions, a metric is an observation — interesting, possibly, but not decision-relevant. Building a dashboard before answering them produces a dashboard that gets reviewed and does not change anything.
The second design requirement is that the metric must be sensitive to the thing being managed. A metric that does not move when the underlying reality moves is not a measurement — it is a baseline held constant. Teams working in the metric's domain should be able to see their decisions reflected in it. When they cannot, the metric belongs to someone else's decisions, or to no one's.
The third requirement is that the metric must be honest about what it does not capture. Every metric is a simplification. The useful question is not what the metric shows but what it does not show that the decision-maker needs to know. Organizations that treat their metrics as complete pictures rather than partial ones develop blind spots that are proportional to their confidence in the data.
What happens when measurement fails
The failure is not usually visible in the data. It is visible in the texture of how the organization operates. Leaders spend more time explaining why the metrics look good while outcomes look bad. Teams develop parallel informal tracking systems — spreadsheets, personal notes, word-of-mouth reporting — that carry the real performance signal the formal system does not. Decisions are made based on anecdote and instinct because the metrics do not illuminate the right things.
These are symptoms of a measurement system that has decoupled from the organization's actual operating reality. The decoupling is gradual. It accelerates when the organization grows — because scale increases the distance between the leaders who designed the measurement system and the work the system is supposed to measure. It accelerates further when strategy changes and the metrics do not, because no one has established a process for reviewing whether the measurement system still fits the strategic direction.
The organizations that avoid this failure maintain a live connection between the questions leadership is trying to answer and the metrics that purport to answer them. This requires periodic review — not of the numbers, but of the questions. Are these still the right things to measure? Do these metrics still connect to the decisions they were designed to inform? Have the conditions that made these proxies valid changed? Most organizations do not have a process for asking these questions. The measurement system persists because it is there, not because it is right.
The path to measurement that works
Improving a measurement system is not primarily a technical project. The technology for capturing and displaying data is not the constraint. The constraint is the clarity of thought required to identify what decisions need to be made, what information would change those decisions, and whether that information is actually being captured.
The work starts with an audit of the decisions, not the dashboards. For each significant operating decision — how resources are allocated, which accounts get attention, when to escalate, when to change direction — the question is: what data currently informs this decision, and does that data actually illuminate the right factors? The gap between what informs the decision and what should inform it is the design brief for a better measurement system.
The redesign does not require new infrastructure in most cases. It requires being precise about what question each metric answers, who makes the decision that question is relevant to, and what threshold in the data would produce a different decision. Metrics that do not have clear answers to these questions are candidates for removal, not refinement. A shorter measurement system with genuine decision relevance is more valuable than a comprehensive one that is reviewed and disregarded.
Measurement does not create clarity by itself. It creates the appearance of clarity — which is more dangerous than acknowledged uncertainty, because it forecloses the questions that uncertainty would have forced. The organizations that measure well are not the ones with the most data. They are the ones that remain skeptical of the data they have.
Is this happening inside a live workflow?
Rivington diagnoses and redesigns one important workflow in three weeks.