Essay · measurement
AI changed the queue. The scorecard did not.
When automation takes the easy work first, every number on the legacy scorecard moves. Read literally, the movement looks like decline.
Choose a single number, put weight on it, and the floor will chase the number instead of the work. Average handle time alone teaches people to end contacts. Adherence alone teaches people to be in the right state rather than in the right conversation. Occupancy alone teaches an operation to run people hot and file the result under efficiency. QA compliance alone teaches a checklist. That trap is a decade old and well documented, and most operations have at least built the habit of pairing a metric with a counter-metric so the gaming cancels.
The trap that arrived with AI deflection is subtler, and it is worse, because it does not require anyone to game anything.
"The Goodhart trap of 2015 was a single-metric pursuit… The Goodhart trap of the post-AI floor is subtler. When AI deflects the easy work, the human queue gets harder by definition. AHT rises because the remaining calls are more complex. FCR gets harder… CSAT becomes less cleanly attributable… The traditional scorecard reads that as a degraded operation. It may not be. It may be the same operation handling harder work with the same humans. If the targets are not recalibrated for what the human queue actually is now, the operation penalizes the floor for a math problem the floor did not create."
(excerpts from my working manuscript, here and below)
Read that slowly, because it is arithmetic and not a position. Deflection does not thin the queue evenly. It removes the contacts that were easiest to automate, which are the same contacts that were easiest to handle: the balance check, the reset, the where-is-it, the change of address. What is left behind is not a smaller version of the old queue. It is a different population, and every legacy metric is now measuring that different population while carrying a target that was cut against the old one.
Handle time rises. That is not a slower agent. It is a harder contact. First-contact resolution gets harder, because what remains is less likely to resolve in one touch. Resolution increasingly depends on a policy exception, a back-office correction, or a decision that sits with someone who is not on the call. Satisfaction becomes less cleanly attributable, because by the time a person is reached the customer has usually already spent effort somewhere else, and the score carries that effort whether or not the agent could have done anything about it.
None of those movements is a performance signal. They are composition signals. Read as performance, they describe an operation getting worse. Read as composition, they describe an operation absorbing precisely the work automation could not take.
The damage starts when the targets do not move. A baseline drawn from a queue that no longer exists becomes the standard the floor is held to, and everything downstream inherits the error. Coaching gets aimed at the wrong behavior, because the report says handle time and the actual problem is that nobody has decided who owns the exception. The agent trusted with the hardest contacts of the day appears on the sheet as the weakest performer, and the routing that put them there is invisible in the number. Performance conversations turn into arguments about the number rather than the work, which is the oldest failure mode in the building wearing new clothes. Then people leave, and the operation writes it down as market conditions. Why that write-down is wrong (and what it costs) gets an essay of its own.
The defense is not a new dashboard. It is a standing habit of interrogating any metric before it is allowed to carry weight: three questions, asked in order.
"1. What reality is this metric trying to reveal? 2. What trade does this metric create when pressure arrives? 3. What would a reasonable person do to make this number look good without improving the work underneath it? If the answer to the third question is obvious (and it usually is), the metric is not the truth. It is a signal."
The third question is the one that does the work, and it is the one most scorecard reviews skip. Ask it of an unchanged handle-time target sitting on a deflected queue. A reasonable person, under pressure, with that target and that queue, transfers earlier. Splits one problem into two contacts so each looks short. Keeps the wrap thin and the notes thinner. Steers toward whatever simple work is still reaching the human layer, and lets the hard contact wait for someone else. Every one of those moves is rational. Not one of them improves the work, and several of them make the next contact worse for the customer and for whoever picks it up. If the answer is that obvious, the target is no longer measuring the operation. It is instructing it.
Recalibration is not lowering the bar, and it should not be sold as generosity toward the floor. A target is a description of the work before it is a demand on the people doing it. When the work changes and the description does not, the difference does not vanish. It gets charged to the floor as underperformance. Re-cutting baselines against the queue as it actually is now is simply keeping the description accurate. What follows from an accurate description is a different set of questions: which contact types are the humans actually carrying, what does resolution now depend on that the agent does not control, and what would have to be true for first-contact resolution to be a fair thing to ask.
The instruments themselves are not the problem. Handle time, occupancy, adherence and QA all still point at something real. They point. Diagnosis decides. The discipline is to refuse to let the metric become a substitute for the work, which is harder now, not easier, because the metric has quietly changed what it is measuring while keeping its old name.
There is one question that catches most of this before it becomes a quarter of misdirected coaching. When someone says, "We hit the number," the question the old dashboard avoids is: what did we bend to get there?
Want to talk about this? Email me.