Essay · measurement

Read the trade, not the target.

A metric is a stand-in: the floor sees the work, the report sees the number. Promote the number from signal to target and it starts buying its own improvement with the work underneath it.

No leader can hear every interaction. No leader can watch every handoff, or catch the moment a person chose the slower path because the customer needed the slower path. The operation builds instruments and reads those instead. A proxy is a number standing in for something the floor can see and the report cannot.

Then the instrument gets promoted and becomes the work. The dashboard does not have to lie for that. It only has to be partial, and then someone with authority has to treat partial as complete.

I learned that from the headset.

The leaderboard and the cube wall

I had moved from a supervisor seat at a BPO back to an agent seat at a large consumer operation. On paper it was a step backward. It was my second time on the phones, and this time I had seen the other side of the dashboard.

The people and the place in this story are deliberately blurred. The pattern is the point.

There was an agent on my team who hung up on customers. I knew the pattern. Her handle time would have raised the question even if I had not heard calls end through the cube wall. When an interaction started to get difficult, she dropped it.

Her CSAT was the highest on the team. The customers who got her surveys were the ones whose calls stayed easy. A satisfaction score picks up anger at a fee, a policy, or a broken process and hands it to whoever stayed on the line. It can also miss the customer whose call ended in the middle.

She was on the leaderboard.

I took the customers whose anger belonged to the process: fees they did not expect, policies they could not work around, refunds the rules would not allow. I stayed on the line, and the survey gave that anger nowhere to land except my score. My team leader coached me on a score in the mid-seventies while the hang-up agent was presented as the model.

The operation did not lack evidence. Her call-type distribution showed she handled virtually no complaints. Her short contacts sat in the same reporting environment that produced the leaderboard. I brought the complaint mix to my team leader and the concern was dismissed as jealousy. The headline numbers were green, so the headline numbers became the truth, and the evidence against them became a problem with the person raising it.

She was recognized. I was coached. The operation rewarded the gameable number and missed the work. My team leader was not the villain in this. She was reading the dashboard the operation had taught her to trust, and it was telling her something clean and wrong.

CSAT was not useless. One person's score could not carry the weight the operation was putting on it. That number should have started a conversation about her call mix. It ended one about my coaching.

Reading the pattern across measures would have shown the work. That took more effort than reading a ranked list, and the ranked list is what got read. What stayed with me is that a system with more data than ever saw me less clearly.

The mechanism has a name

Two laws older than the modern contact center describe what that leaderboard was doing. Goodhart's law: when a measure becomes a target, it ceases to be a good measure. (Charles Goodhart, 1975. The wording that travels is Marilyn Strathern's.) Campbell's law: the harder a number is leaned on for decisions, the more it corrupts the work it was watching.

Nobody has to cheat for this to run. Command a handle-time target and agents shorten calls: some learn better call control, some rush the customer, some drop the hard call. Command a service-level target without reading the trade and the queue gets answered faster while occupancy climbs, coaching gets canceled, and the operation buys the customer promise with the recovery time the floor was running on. Command perfect script compliance and agents read the script at the customer instead of solving the issue. QA climbs. Satisfaction collapses.

The metric improves and the work gets worse. The people producing that result are answering the question the operation actually asked them, so the failure sits in the management system and not on the floor. The incentive is the design question, and it runs upward too. The team leader who pushes handle time, the analyst who softens a forecast to make the budget palatable, and the executive who asks for one clean number are usually trying to make the operation manageable with what they were handed. What they were handed measures one thing at a time and pays out on it.

Dropping customers on purpose is conduct, not a skill gap, and none of this excuses it. Naming what the system rewarded does not clear the person who did it. It puts the burden on the operation to gather evidence strong enough to tell the difference before it attaches a consequence.

Restoring the contract

None of this argues for measuring less. An operation that refuses to measure is guessing. What has to be repaired is the contract between a measure and the reality behind it: a stated limit on what the measure is allowed to say, and a second reading that can argue back. The people who can see what the instrument cannot need a route to say so, and it has to survive being unwelcome.

Three questions before any metric becomes a target.

  1. What reality is this metric trying to reveal?
  2. What trade does this metric create when pressure arrives?
  3. What would a reasonable person do to make this number look good without improving the work underneath it?

The third one usually answers itself. When it does, the number is a signal and not the truth, and it should be worked as a signal. Coach the behavior it points at. Price what the target is buying and name who pays for it.

The same reading changes what happens when a number goes red. The first question stops being who owns it. It becomes what the number is a proxy for, what else could have produced it, and what evidence would tell those apart.

The words we use

Before the scorecard there are the words. A definition has to carry three things: the metric, the trade it makes, and the reason anyone measures it. Drop one and a familiar word turns into cover for an operating choice nobody said out loud.

Service level is the percentage of contacts answered within the promised time. The tradeoff is speed against what pays for quality later: staffing cost, recovery time, coaching time, the patience to work an actual problem to the end. The why is the customer promise, because waiting is real. When service level becomes the only truth, the operation starts buying speed with the conditions underneath it.

Occupancy is the percentage of on-queue time spent handling contacts rather than waiting for the next one. The tradeoff is utilization against human recovery. High occupancy looks efficient until nobody has a minute between a hard contact and the next one. The why is capacity discipline. Some of that apparent idle time is what makes the next call possible.

Adherence is whether people are where the schedule said, when it said. The tradeoff is reliability against context. An operation cannot run if the schedule is optional, and it turns dishonest when every deviation is treated as a character flaw instead of information. The why is trust between planning and execution. WFM can only build a usable plan if the floor treats the plan as real, and the floor will only treat it as real if the plan respects what the floor knows.

Failure demand, John Seddon's term, is the work the operation created for itself because the prior contact, process, or policy did not solve the problem. The tradeoff is short-term efficiency against long-term volume: today's call gets faster by leaving part of the issue unresolved, and tomorrow the same customer comes back and the operation pays again. The why is system truth. Failure demand shows where the operation is manufacturing its own workload and then blaming the floor for being busy.

These definitions are not the operating model. They are the shared language underneath it. When the whole room holds the metric, the tradeoff, and the why together, someone can say "we hit the number" and the next question gets asked out loud instead of absorbed by the floor: what did we bend to get there?

The rule

Every number can be green while capability is leaving. The veteran departs with knowledge nobody captured. The new hire arrives into a schedule with no room to learn from the veteran who left. The reporting system records each piece as a separate event. The floor lives the chain. The dashboard receives the pieces.

So treat a numerical target as unfinished until someone has named what it trades. A number nobody can game has not been invented. A number everybody in the room reads the same way is much harder to hide behind.

All of this hands every miss a defense. Every red number turns into an argument about definitions, and nothing gets decided. The risk is real. Naming a trade and then not holding the number is a slower way to miss the same target. Name what the number costs. Then hold it anyway. An operation that ran hard single-number targets for three years and kept its people would settle it against me.

Want to talk about this? Email me.