The customer service metrics worth tracking in 2026

Every one of these can be improved without helping a single customer. Which is the point — here is how each one lies, and what to pair it with.

Every support metric can be gamed, usually without anyone intending to. That is not an argument against measuring — it is an argument for knowing, for each number, exactly how it goes wrong.

Six that are worth having, what each is for, how each lies, and the second number that keeps it honest.

1. First response time

What it is for: whether the queue is staffed at the hours messages arrive.

How it lies: an automatic “thanks, we’ll get back to you” resets the clock without helping anyone. Teams under pressure discover this within a week, and after that the metric measures how fast you can acknowledge rather than how fast you can answer.

Pair it with: time to a useful answer — the first reply that actually addressed what was asked. Less flattering, harder to fake, and the one that correlates with whether the person stayed.

Read it as a distribution, not a mean. The average is dragged down by the daytime majority. The damage is in the slowest tenth, and that decile is almost always out-of-hours.

2. Resolution time

What it is for: whether conversations end, or just stop.

How it lies: closing a conversation is an action a person takes, so the number partly measures tidiness. A team told to improve it will close faster, which is not the same as resolving faster.

Pair it with: reopen rate. A conversation that comes back was closed, not resolved, and the pair together tells you which you are doing.

3. Abandoned conversations

What it is for: demand you did not serve. Someone asked and gave up.

How it lies: it does not, much — which is why it is the most undervalued number on this list. Its problem is that most setups do not count it at all. A split inbox cannot see it, because a conversation nobody opened looks the same as one nobody needed.

Pair it with: nothing. Just start counting it. If you track one new thing this quarter, this is it.

4. Share of conversations handled without a person

What it is for: whether automation is actually carrying load.

How it lies: spectacularly, if “handled” is defined loosely. A bot that replies and the customer leaves counts as handled. So does a conversation the customer abandoned mid-flow.

Pair it with: the abandoned rate within automated conversations specifically, and the reopen rate on them. High automation with high abandonment is not automation, it is deflection with a nicer name.

5. Escalation rate

What it is for: whether the automated layer is scoped correctly.

How it lies: people read it as bad. It is not a cost, it is a setting. Driving it down means the agent is answering things it should have handed over, which is the expensive failure.

Pair it with: post-escalation correction rate — how often a person had to correct something the agent already said. That is the number that should be near zero. A rising escalation rate with a flat correction rate is a system behaving well under an unfamiliar question.

Watch the derivative, not the level. A sudden climb mid-week means a new question is arriving in bulk, and it is usually fixable in ten minutes.

6. Detected intent distribution

What it is for: knowing what people actually ask, rather than what three colleagues each believe they ask.

How it lies: the categories are yours, so the distribution reflects your taxonomy as much as your customers. An “other” bucket above about a tenth means the taxonomy is wrong, not that customers are unusual.

Pair it with: the list of questions nothing could answer. That is the roadmap, and it is the only metric here that tells you what to do next rather than how you did.

The two that are not on this list

CSAT and NPS. Not because they are worthless, but because at the volumes most teams have, the response rate is low enough and the self-selection strong enough that the number moves on noise. Read them as directional over quarters, never as a weekly target, and never tie them to individual people — that is how you get a team asking customers for good scores.

Messages per conversation. It sounds like efficiency and it is mostly length. Long conversations are often the ones that closed a sale.

How to actually use these

Pick three. Nobody acts on six.

If you are starting from nothing, the useful three are time to a useful answer, abandoned conversations, and the list of questions nothing could answer. The first tells you whether the queue is staffed, the second tells you what it costs when it is not, and the third tells you what to fix.

And set them up so the numbers arrive without anyone assembling them. A metric that requires a monthly export is a metric you will look at twice.


Fellix reports all six from the conversations themselves, so nothing has to be maintained by hand. See what Analytics shows.