A useful measure improves a decision. It helps a team notice change, compare
an outcome with an expectation, investigate meaningful variance, or decide
whether to continue, adapt, or stop.
A number without a decision can still create work and status. Once a target
affects reputation or reward, people will adapt to it. Measurement therefore
requires judgment about purpose, definition, behavior, and the costs the
number does not show.
Begin with the decision
Before choosing a metric, write the question and action it should support:
Each month, the service owner will use this measure to decide whether the
current intake process needs a capacity, quality, or sequencing change.
If no one can name the owner, cadence, and plausible action, the measure may be
informative but it is not yet part of a management system.
Do not collect data merely because it is available. Easy activity counts can
crowd out harder evidence about whether anything improved.
Pair outcomes with earlier signals
A lagging measure shows the outcome after it occurs. A leading indicator offers
earlier evidence but is useful only when its relationship to the outcome is
tested.
For example, customer time-to-value may be the outcome. Age of active work,
handoff delay, or incomplete intake may be earlier signals. Track the leading
measure as a hypothesis, not a law.
Add a guardrail for a cost the main measure might hide. Faster completion is
not an improvement if rework, exclusion, risk, or employee harm increases.
Make the definition reproducible
A metric needs:
- a clear purpose and decision owner;
- formula, unit, population, and time window;
- source and refresh cadence;
- target or comparison, when justified;
- material exclusions and known limitations;
- a guardrail;
- likely gaming or displacement behavior; and
- a review date.
Two people using the same definition should produce the same result. Precision
in the calculation does not eliminate uncertainty in what the measure means.
Inspect the behavior the metric creates
Targets shape action. A ticket-closure target can reward premature closure. An
uptime target can hide degraded user experience. A hiring-speed target can
weaken evidence and access.
Ask:
- How could a reasonable person improve this number without improving the
outcome?
- Whose cost is outside the denominator?
- What work becomes less visible when attention moves here?
- What balancing signal would reveal displacement?
- Should this be a target, a diagnostic, or context only?
Sometimes the best response is to change the measure. Sometimes it is to
remove the target while keeping the diagnostic.
Review variance for learning
Not every movement deserves a story. Define the scale or pattern that warrants
investigation. At review, separate:
- a real change in the system;
- ordinary variation;
- a data-quality problem;
- a definition or population change; and
- a one-time event.
Then ask what evidence would distinguish the explanations and what decision is
reversible now. Avoid presenting a confident narrative simply because the
chart moved.
Connect reflection to one owned change
Measurement becomes learning when a review changes behavior, a system
condition, or the next question.
Record one decision, owner, expected effect, guardrail, and review date. At the
next cadence, inspect whether the change produced the expected signal. A list
of observations without ownership creates retrospective theater.
Retire measures that no longer inform decisions. A metric portfolio has an
attention cost even when data collection is automated.
Practice with a metric card
Create one card for a live operational measure. Include purpose, decision
owner, formula, source, window, target or comparison, guardrail, limitations,
gaming risk, and review cadence.
Walk through two recent periods and ask whether the card would have changed a
decision. Have someone outside the work reproduce the interpretation from the
definition. Revise ambiguous terms and identify one way the measure could be
improved without improving the outcome.
Use resources with context
The DORA research program is useful for seeing how outcome measures and
organizational capabilities can guide sustainable improvement, especially in
technology delivery. Measure What Matters offers an accessible model for
connecting outcomes, transparency, check-ins, and learning.
Neither source supplies universal targets. Resource completion does not
establish measurement capability; a decision-ready, behavior-aware metric
system does.
Evidence of Practice
You may be ready to record Practice when you can:
- name the decision, owner, and cadence a measure supports;
- distinguish an outcome from an activity count;
- pair a lagging outcome with a tested leading hypothesis and guardrail;
- define formula, population, window, source, and exclusions reproducibly;
- identify how a target could be gamed or displace harm;
- distinguish system change from ordinary variation and data-quality issues;
- turn a review into one owned change with a later effectiveness check; and
- retire a measure that no longer improves a decision.
These prompts support an explicit proficiency judgment; they do not create it
automatically.
Continue through the tree
Goals, planning & operating cadence gives
measurement a decision rhythm.
Norms, incentives & culture exposes the behavior a
target may reward.
Performance & accountability applies evidence
to individual expectations without reducing judgment to a number.