05. Does what cannot be measured exist?
The question: What is missing from the dashboard?
Imagine the scene after a service outage at two in the morning. The monitoring screen is green again, and average response time has fallen back within its target. The error rate is below 0.1 percent. The on-call engineer leans back in the chair and lets out a long breath. But during that time, someone may have pressed the payment button over and over. Someone else may have called customer support and given up. Someone may simply have left for another service without saying anything. The line on the dashboard is calm again, but a user's day may already be ruined.
We are taught that measurable things are easier to manage. Numbers can be compared, graphs can be pinned to a meeting-room wall, and target values can be placed in tickets and quarterly review sheets. At some point, things that cannot be measured get pushed aside as if they do not matter. Satisfaction, trust, dignity, hesitation, and the feeling of being excluded are left outside the discussion because they seem too vague. Philosophy stops here. Does measurement discover an object, or does it cut out only some of what already exists?
The philosophical point: Is a number a window onto the world, or a frame?
Measurement does not copy the world as it is. We first have to decide what to count, where to begin and end, and which unit to use. If we measure a school's achievement through test scores, the scores become clear while curiosity and the ability to cooperate fade from view. If we measure a hospital's performance through average length of stay, the number becomes neat while the patient's anxiety at home after discharge disappears.
Goodhart's law states the danger briefly: once a measure becomes a target, it stops being a good measure. This is not merely an operations tip. It is a philosophical warning about how human beings handle the world. The moment we translate a value into a number, that number becomes both a sign intended to describe reality and a rule that directs behavior. People may change reality to meet a target, but sometimes they leave reality untouched and work around the measurement instead.
That does not mean we can give up on measurement. Without measurement, power can more easily present its own judgment as fact. Return rates and churn rates can provide a starting point for reviewing the claim that "users like it." The problem is not that numbers are wrong. It is that they pretend to be the whole when they are only a part. Measurement should not replace truth. It should be a limited observation through which we approach it.
An engineering case: The shadow of fast shipping
Suppose an e-commerce team chooses "the percentage of orders shipped within 24 hours" as its primary metric to reduce shipping delays. A few weeks later, the metric looks remarkably better. When an order arrives, the logistics team immediately packs the item and prints a shipping label. Even when the inventory is not actually ready or the address is wrong, the system marks the order as shipped first. The dashboard fills with successes while customers wait for empty boxes or cannot correct their shipping addresses. The team has met its target, but the service has become worse.
A more careful team might track not only the shipping rate, but also the time until actual delivery, cancellation rates, customer inquiries, and refund rates. Even that is not enough. Older users may have trouble but never submit a support request. Users from other countries may fail to understand an error message and quietly leave. A low volume of logs does not mean there is little trouble. Silence may signal satisfaction, but it may also mean that people could not find a way to speak.
The engineer's job does not end with adding more metrics. First ask: "Whose delay are we trying to reduce?" Are we lowering the system-wide average while leaving the worst experiences of a particular group untouched? Put the distribution of failure next to the success metric. Regularly collect voices from outside the numbers. Audit how the target was achieved, not just whether it was achieved. A value judgment is not a sentence written once in a requirements document. It has to be repeated in the schema, the alert rules, and the deployment approval process.
Counterpoint and tension: Don't we need to measure in order to act?
The objection that "values we cannot measure cannot move a team" is a strong one. It is difficult to secure a budget with nothing more than a call to protect trust, and a declaration to reduce discrimination does not tell us where to set a model's threshold. Engineering is a world of action, so some operationalization and quantification are necessary.
That is true. But operationalization should be treated as a hypothesis about how to approach a value, not as the value itself. If we define trust as a single measure of complaint-response time, we should also record what that definition leaves out. Metrics need an expiration date and a stated scope, along with counter-metrics and qualitative review. Humility does not mean eliminating numbers. It means limiting the authority we give them.
A question to leave with
Who was outside the graph we optimized today? Was there anyone who mattered but had not been given a language of measurement, rather than someone who did not matter because they had not been measured? What metric should we add in the next deployment? More fundamentally, which experiences should we decide not to turn into numbers?