When does a number become a target?

A number can describe performance. Once it becomes the key to a bonus, a budget, a promotion, or a reputation, it also shapes behavior. Over time, a team may stop optimizing for value delivered to a customer, patient, product, or institution and start optimizing for the number that stands in for that value.

A key performance indicator, or KPI, is a measure selected to track progress toward an objective. A good KPI does not replace judgment. It tells decision-makers where they stand, which risk deserves attention, and what needs investigation. The problem is not measurement itself. The problem is confusing the measure with the value it represents.

My central inference is simple: once measurement is connected to rewards and control, it stops being a passive record. It becomes part of the behavior it is meant to observe. A good KPI therefore has to do more than calculate correctly. It has to serve the right decision.

Three warnings, one mechanism

This is not a recent management fad. In 1975, Charles Goodhart wrote about the problem in the context of monetary policy. A statistical regularity can become less reliable when authorities use it as a control target. The original issue was the relationship between money and the British economy: once the relationship became a policy instrument, it did not remain behaviorally independent. Goodhart's sentence is not a magic law that explains every management failure. It is a warning that a measure used for control will not remain outside the system it controls.

Donald Campbell made a related point in a 1976 paper on social policy. The more a quantitative social indicator is used to allocate jobs, money, or power, the more pressure there is to shape the process that produces it. If improving the indicator becomes easier than improving the social outcome it is meant to track, the number starts to displace reality.

Steven Kerr's classic article makes the design problem more direct: an organization can reward one behavior while hoping for another result. Reward sales volume while hoping for trust. Reward call speed while hoping for durable resolution. Reward quarterly profit while hoping for long-term resilience. These are different versions of the same mismatch.

The common point is not that the original measure was useless. It may have been useful as a signal. A signal can be read alongside other signals. Once it becomes a target, a budget gate, or a compensation trigger, people learn how the number is produced and adjust their behavior around it.

Midas and the price of the proxy

Midas is not only a convenient metaphor, but history and myth need to remain separate.

Herodotus refers to Midas, son of Gordias, as a king of Phrygia and says that Midas dedicated the royal seat from which he judged at Delphi. UNESCO describes Gordion as the political and cultural center of ancient Phrygia and identifies the Midas Mound as a major archaeological feature, with a burial chamber dating to around 740 BCE. The Metropolitan Museum distinguishes the historical king known as Midas to the Greeks and Mita in Assyrian sources from later Greek legends and myths.

The golden-touch story is a literary account in Ovid's Metamorphoses, Book 11. Midas asks for the power to turn everything he touches into gold, discovers that food is included, and asks to be released from the gift. The story lasts because visible gain destroys the system that made life possible.

The existence of a historical Midas does not make the golden touch a historical event. Archaeology can establish a political center and an elite burial; it cannot establish a literal transformation of objects into gold. Midas is an interpretive lens here, not evidence.

The gap between measurement and value

George Baker's work on incentive contracts frames the core problem: if an employee's payoff is tied to a performance measure rather than to the principal's underlying objective, the contract may fail to create the right incentives. A metric and an objective are not the same thing. The closer the metric is to the objective, the more useful it may be, but the gap does not disappear.

Bengt Holmstrom and Paul Milgrom's model of multi-task work points to a harder consequence. When some parts of a job are easy to measure and others are not, strong rewards for the measurable parts can pull effort away from the rest. This does not require bad intent. A person with limited time will often move effort toward the place where the reward is most visible.

That is why a single sales metric can hide customer relationships, a delivery metric can hide quality, a call-time metric can hide durable resolution, and a quarterly profit figure can hide investment in resilience. Work that is invisible may be neglected while the dashboard looks healthier. The cost can appear later, in another team or another customer group.

Adding more KPIs is not an automatic fix. Gaming one measure may be easy; managing the weights and thresholds of twenty measures may also be easy. Longer lists can create competing priorities and reporting burden. A balanced scorecard is not a larger pile of boxes. It is a way to make the important tensions in a decision's value chain visible.

Case: the number grew while customer value fell

Wells Fargo's retail banking sales practices are a bounded public example of how measurement can become a behavioral and information problem.

In 2016, the Consumer Financial Protection Bureau found that sales targets and compensation incentives spurred employees to open deposit and credit-card accounts without customers' knowledge or consent. The agency said the bank's own analysis identified more than two million deposit and credit-card accounts that may not have been authorized by customers.

In 2020, the Securities and Exchange Commission found that Wells Fargo had promoted a cross-sell metric to investors even though it was inflated by accounts and services that were unused, unneeded, or unauthorized. The SEC announced a $500 million civil penalty.

This cannot be reduced to "a bad KPI." Wells Fargo's independent directors investigated broader root causes, including corporate structure, culture, and individual actions. The lesson is not that one number explains everything. It is that when a target, a bonus, job security, and the investor story all depend on the same number, the number can detach from customer value and make that detachment look like success.

Strongest competing explanation: when measurement helps

Ignoring the positive cases would make the thesis weaker. Measurement and goals can be useful.

A BMJ review of targets in the English National Health Service says reported performance against key targets improved, while the genuineness of those improvements and the costs to other services remained uncertain. A related study by Bevan and Hood examines gaming around targets in the English public health care system. The distinction matters: a target can improve and still fail to capture the value of the whole service.

Kaplan and Norton's balanced-scorecard approach began with a related concern: a single financial measure can miss the capabilities a company needs for the future. It proposed looking at customer experience, internal processes, learning, and innovation alongside financial results. That is not a universal formula. In a within-firm study, Griffith, Neely, and Smith found that a balanced-scorecard-based performance-pay scheme had some effect, but that the effect varied with branch characteristics and managerial experience. A broader meta-analysis found positive average relationships for strategic performance measurement systems, while also finding substantial heterogeneity and an important role for whether systems were linked to rewards.

So the question is not "Should we use KPIs?" Better questions are:

  • What decision will this measure improve?
  • Which part of the value chain does it see?
  • Which value or risk does it miss?
  • How will its meaning change once people adapt to it?

Better decision architecture, not more KPIs

Before tying a measure to a reward system, leaders need at least seven checks.

  1. Define the decision. Write down which decision the number is supposed to improve. If the purpose is only "increase performance," the measure is not ready.
  2. Separate the controllable from the uncontrollable. Who can influence the outcome, and which part depends on outside conditions? A result that one team cannot control should not automatically determine its pay.
  3. Complete the value chain. Look at activity, output, outcome, quality, risk, and longer-term effect together. Sales volume is an output. Product use, problem resolution, and retained trust are outcomes.
  4. Write the measurement dictionary. Define the numerator, denominator, period, data owner, exceptions, and verification method before setting the target. Otherwise the team may manage the definition rather than the result.
  5. Choose the consequence. A signal can support learning, inform a manager's judgment, or directly affect pay. Those uses require different levels of confidence.
  6. Add guardrails. Customer harm, quality loss, safety failures, or regulatory violations should not count as success merely because a volume target was met.
  7. Set a review and retirement date. People adapt to systems. A measure that worked in year one may weaken in year two as behavior changes. Its authority should depend on periodic review and a clear option to redesign or retire it.

This framework does not ask for more measures. It makes a measure's role, limits, and shelf life explicit.

Three scenarios

Scenario 1: The measure remains a signal. Customer satisfaction, durable resolution, cost, and quality are read together. Definitions are clear, employees know how results will be reviewed, and managers investigate anomalies. Measurement focuses attention without becoming the entire verdict.

Scenario 2: The proxy becomes Midas. Management rewards transaction volume while quality and customer outcomes arrive too late to matter. The team raises the number while the value behind it deteriorates. This is not a prediction. It is a conditional scenario that can arise when reward and visibility are designed around one proxy.

Scenario 3: Measurement disappears. An organization removes targets to avoid their harms but does not replace them with shared definitions, recorded review, or independent challenge. Decisions may not improve. They may simply become less visible. The answer to imperfect measurement is not unexamined discretion.

Questions leaders should ask

Before attaching a KPI to a budget, promotion, or bonus, leaders should ask:

  • What decision is this number meant to improve?
  • What valuable work does it not see?
  • If the number rises while customer or long-term outcomes worsen, which signal wins?
  • Who can influence the number, and who bears the cost of gaming it?
  • Are the definition, denominator, period, and data source clear to the people being judged?
  • Is this a diagnostic tool, an input to managerial review, or an automatic reward trigger?
  • What evidence shows that the measure tracks value rather than a convenient proxy?
  • How will the organization separate outside conditions from the owner's contribution?
  • Who will challenge quiet quality loss or gaming?
  • When will the organization retire the measure?

Judgment: a KPI is evidence, not a verdict

Good management does not abandon measurement. It gives measurement a more modest and more accurate role.

A KPI can be a radar, not a court ruling. A number focuses attention, starts an investigation, and narrows the next question. Value does not live inside the number. It appears in the result an institution protects over time.

Midas's mistake was not wanting gold. It was mistaking a visible transformation for the thing that made life possible. Organizations make a similar mistake when they replace the value represented by a proxy with the proxy itself.

Protecting the value of measurement requires three habits: know what the system cannot see, separate rewards from signals when the signal is weak, and accept that a measure changes as behavior adapts to it. A good scorecard does not guarantee success. It makes false success easier to see and gives leaders a chance to correct the system before the cost becomes irreversible.

Sources