When the Measure Becomes the Target: Why Quantifying Everything Makes It Harder to See

当度量变成目标:为什么把一切量化之后,反而看不清了

From central-bank targets to quarterly earnings — the same mechanism keeps recurring: once a number is used to steer action, it stops being a good description of reality

A Concrete Problem: Why More Measurement Can Mean Less Clarity

Modern institutions translate almost everything into numbers. Schools track progression rates, hospitals track bed turnover, companies track quarterly earnings per share, investors track Sharpe ratios, and individuals are advised to log their sleep and their steps. The promise behind this translation is straightforward: numbers make comparison possible, they give responsibility somewhere to sit, and they replace vague intuition with evidence that can be checked.

Yet almost every number that stays in use long enough develops the same symptom: it stops being true. Scores stop standing for learning, click-through rates stop standing for value, reported profit stops standing for the quality of a business. The cheapest explanation for this symptom is that somebody is cheating, and so the remedy is always the same — more auditing, harsher penalties, finer rules.

The judgement this essay defends is different: the distortion of measures is not a moral failure at the point of execution but a mechanical consequence of measurement itself. Whenever a single number is used both to describe reality and to allocate reward, distortion arises on its own, regardless of anyone's character.

To see this clearly, three questions have to be answered in order: why numbers distort; whether distortion means measurement should be abandoned; and under what conditions a measure can still be trusted.

The Double Life of a Number: Description and Command

An indicator normally does two jobs at once. It describes reality, and it distributes reward.

The trouble is that these two jobs impose opposite requirements. As a description, an indicator should track reality as faithfully as possible. As a command, it must be simple, comparable, and attributable. Reality is multidimensional; anything capable of commanding action has to be one-dimensional. Compression necessarily destroys information.

The people being measured are not compressed along with it. They continue to live in the high-dimensional reality, and can therefore always move along a dimension that is not being counted. That space of movement is where distortion comes from.

Two separate fields discovered this independently. Studying British monetary targets in 1975, Charles Goodhart observed that any statistical regularity that is observed will tend to collapse once the authorities place pressure on it for purposes of control. The formulation was later condensed into something more memorable: when a measure becomes a target, it ceases to be a good measure. Four years later, assessing the evaluation of social programmes, Donald Campbell reached a near-isomorphic conclusion: the more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures, and the more apt it will be to distort the very process it was meant to monitor.

Neither is really about cheating. Both describe a structural fact: a proxy variable and the goal it stands for are never identical, and incentives direct all available energy into the gap.

What follows is a characteristic split state. Everyone knows the number is lying. Nobody can stop acting on it, because whoever stops first immediately falls behind on the number.

The Counterargument: Without Measurement, Things Get Worse

Before going further, the strongest opposing case deserves a hearing, because the claim that measures distort slides very easily into a lazy conclusion: if measures distort, trust intuition instead.

That conclusion does not survive the evidence.

In Noise, Daniel Kahneman and his co-authors show that what is genuinely underestimated in human judgement is not bias but noise — the meaningless random variation between judgements that ought to be identical. The same file given to different judges, the same résumé given to different interviewers, can produce widely different conclusions, and the causes are often nothing more than time of day, order, and mood. Their conclusion is unkind to professional intuition: across most repeatable decisions, a simple mechanical rule outperforms an experienced expert, and does so more consistently.

Philip Tetlock supplies evidence pointing the same way in Superforecasting. He asked participants to state forecasts as checkable probabilities, then scored every forecast by a fixed rule. The finding: measurable feedback is a precondition for calibration. The people who eventually performed best were not unusually gifted; they revised their judgements continuously and in small increments, and they could see their own track record.

Together these constitute a strong counterargument: abandoning measurement hands decisions back to intuition that cannot be reproduced and cannot be improved.

So the real choice is not measurement versus intuition. It is how measurement is used. Treating "measures distort" as a reason to oppose all measurement is the same laziness as treating "intuition is unreliable" as a reason to oppose all judgement.

Where the Gap Turns Lethal

If measurement cannot be abandoned, the question becomes: under what conditions is distortion most expensive?

Three conditions amplify the risk sharply.

First, the further the proxy sits from the goal, the more dangerous it is. An exam score is one step from learning and several steps from competence; a click-through rate is several steps from user value. The longer the chain, the more room there is to optimise the measure without improving the goal.

Second, the faster the feedback and the more immediate the reward, the more dangerous it is. Annual reviews are easier to game than ten-year assessments. This shows up most clearly in capital allocation. In The Outsiders, William Thorndike examines eight chief executives with extraordinary long-run returns and finds that what they shared was precisely a refusal to run the business on Wall Street's quarterly rhythm: they optimised the long-run compounding of cash flow per share rather than the uninterrupted growth of earnings per share. Yet the market's default way of grading a CEO is the latter. A system that grades on quarterly earnings will therefore systematically select executives who are good at managing expectations rather than executives who are good at allocating capital.

Third, the more the measured and the measurer have divergent interests and asymmetric information, the more dangerous it is. In Poor Charlie's Almanack, Charlie Munger places incentive-caused bias first among the tendencies that lead people astray, for a direct reason: incentives do not need to persuade anyone, they change what a person is able to see. He also returns repeatedly to the point that accounting figures are a set of conventions rather than a direct reading of reality — and when a convention is mistaken for reality itself, manipulation stops being cheating and becomes compliance.

The mechanism is not confined to business. In The Dictator's Handbook, Bruce Bueno de Mesquita and Alastair Smith argue that a politician's first objective is to maintain the coalition that keeps them in power, so public indicators get produced and presented selectively, in the service of political survival rather than of accurate reporting. Where information is most asymmetric, measurement most easily becomes a narrative instrument.

Taken together, these give a rough but useful boundary: measurement is most reliable when describing what has already happened, is repeatable, and is low-dimensional; it is most dangerous when used to command what has not happened, happens once, and is high-dimensional. Financial accounts belong to the first category; strategic judgement belongs to the second. Applying the evaluation method of one to the other causes trouble in either direction.

The Judgement: An Instrument Panel Is Not a Steering Wheel

In Thinking in Systems, Donella Meadows ranks both "the goals of a system" and "information flows" very high among leverage points. The implication is direct: changing one indicator can change the behaviour of an entire system — which also means that whoever controls the indicator controls, in practice, the distribution of behaviour.

Three operating principles follow.

  1. Separate the instrument panel from the steering wheel. Indicators are good at detecting anomalies and bad at simultaneously allocating reward. When one number serves as both, the reading fails. A workable arrangement keeps the measures used for evaluation distinct from the measures used for observation, and updates the latter more often than the former.
  1. Use a portfolio rather than a single number. As long as one uncounted dimension exists, moving along it is the rational choice. Pairing a primary measure with at least one measure from a different dimension raises the cost of gaming substantially.
  1. Preserve one channel of judgement that is not scored. This is not the abolition of judgement but an admission that some questions must be settled by argument rather than by score — the long-run contribution of a person, the soundness of a strategic direction, the difference between a lucky quarter and a durable capability.

Finally, the limits. This judgement has boundaries of its own. In safety, medicine, and engineering, hard measurement has saved a great many lives and irreversibly raised the floor of whole industries — aviation accident rates and surgical infection rates were driven down by measurement, not by intuition. So the test of a measure has never been whether it is quantitative, but whether it still permits itself to be falsified: when someone points out that the indicator has drifted away from reality, does the system revise the indicator, or does it revise the person?

The most dangerous state of any system is not the absence of indicators. It is a state in which everyone knows the numbers are lying, and nobody can stop acting on them.

Related books