Customer Service KPIs That Actually Count (and How to Measure Them)
Most support numbers are sedatives. They rise when your shop grows, and afterwards nobody knows what to do.
I spent years in management consulting building metrics for other people. The lesson from that is uncomfortable: a metric is only worth having if a bad week forces you into a concrete action. Everything else is decoration.
With nano, mate and MUSTAX we run three brands with two people. We have no time for numbers you only look at. This article covers which four metrics we actually use, which ones we deliberately ignore, and how you measure all of it in a spreadsheet without an expensive tool.
Why most support metrics mislead you
The most common mistake is an absolute number without a reference figure. "We had 180 requests this week" says nothing at all if you do not know how many orders sat behind them.
180 requests on 400 orders is a problem. 180 requests on 2,000 orders is a good result. Same number, two completely different weeks.
The second mistake is the average. An average response time of four hours sounds solid. It can mean that 90 percent of customers get an answer in 20 minutes and ten percent wait two days. Those ten percent are the ones who write you the bad reviews.
The third mistake is the number without an owner. If a metric goes bad and nobody changes anything afterwards, you may as well leave it out. It only costs you attention.
Metric 1: Contacts per 100 orders
This is our most important number. It answers the question: how much friction does my shop create per parcel sold?
The calculation is simple. Requests in the period divided by orders in the period, times 100. Done.
Why this number is so good: it does not automatically grow with your revenue. If you sell twice as much and the number stays the same, you have scaled cleanly. If it rises, something is broken, and usually not in support.
Ours once went up considerably within two weeks. The cause was at the carrier, not in the inbox: transit times had got longer, so more customers asked about their parcel. Without this metric we would have gone looking for more support capacity instead of the cause.
That is exactly the value. A rising contact rate is a hint at a product, shipping or content problem. Support is only the place where it becomes visible.
Metric 2: First contact resolution
First contact resolution measures how many requests are settled with a single reply. No chasing, no follow-up question, no second contact on the same topic.
It is the toughest quality number I know. A fast reply that clarifies nothing produces two more messages. In the response time that looks good, in first contact resolution it does not.
How to measure it without buying a ticketing system: count the cases where the customer wrote again after your reply. In Gmail, a glance at the conversation length is enough. More than two messages from the customer in the same thread means it was not solved the first time.
If the rate is poor, it is almost always one of three things. Your reply was incomplete. You asked for data you could have looked up yourself. Or your template does not fit the occasion.
You can fix all three the same day. That is what makes the metric usable.
Metric 3: Response time as a median, not an average
Response time deserves to be measured, just not as a mean. Take the median and, alongside it, the worst value in the top ten percent.
The median tells you what a normal customer experiences. The bad edge tells you what the angry customer experiences. Together they give you a picture that the average alone does not.
Second rule: measure the time to the first helpful reply, not to the acknowledgement of receipt. An autoresponder writing "we have received your message" solves nothing. Counting it towards your response time is fooling yourself.
Our median is now in the range of minutes, because the most common cases are answered automatically. That is down to automation rather than diligence. How that setup is built is in the guide to customer service automation.
One more warning signal: if response time falls and first contact resolution falls at the same time, you are answering faster and worse. Those two numbers always belong side by side.
Metric 4: Share answered automatically
This number measures what percentage of all requests were closed without a human. For us that is around 65 percent.
It is the only one of the four metrics tied directly to your time. Every percentage point more is time you get back.
A clean definition matters. Answered automatically means: the customer received a complete answer and did not ask again afterwards. An auto-reply that a human has to follow up on does not count. Otherwise you are optimising a number that saves you no minutes.
Where this number comes from is no secret here. The largest block is WISMO requests, meaning the question about the parcel. They make up around 40 percent of our volume and are almost fully automatable, because the answer is already in the system anyway. The approach is described in the article on WISMO requests.
The second block is pre-sorting. When an AI categorises every email reliably, automation emerges almost by itself, because you know which category tolerates a standard answer. How that works is in the article on pre-sorting support emails.
Metrics we deliberately ignore
Not every common number deserves a place. We sorted these four out.
| Metric | Why it misleads |
|---|---|
| Absolute ticket volume | grows with revenue, says nothing about quality |
| Handling time per ticket | rewards short replies, not solved problems |
| Satisfaction without context | almost only very happy and very angry customers answer |
| Net promoter score, quarterly | too slow to trigger a decision in support |
That does not mean these numbers are worthless. They just do not work as a steering metric for a small brand. With two people on the team you need numbers that trigger an action this week.
The satisfaction survey in particular is overrated. We ran it for six months. The value fluctuated without us ever being able to say why, and not a single decision came out of it.
How to measure this without buying a tool
You do not need a helpdesk system at 200 euros a month. You need four numbers a week and a fixed appointment at which you look at them.
- Labels in the inbox. Every request gets a category: WISMO, return, product question, complaint, spam. An AI can take that over, or you do it yourself in ten minutes a day.
- Order count from the shop. Shopify gives it to you directly. It is the denominator for metric 1.
- A spreadsheet with six columns. Week, requests, orders, contacts per 100 orders, first contact resolution, share automated. Nothing more.
- One appointment a week. For us, Monday morning, 15 minutes. We only look at the change versus the previous week.
- A rule for outliers. If a number deviates by more than a fifth, we look for the cause the same day.
At MUSTAX we later automated the reporting. Two days a month used to go on numbers, today it is two hours. But the first version was a hand-maintained spreadsheet, and it served its purpose for a year.
Start the same way. An automated dashboard for metrics you do not understand yet is wasted money. When you reach the point where the manual work becomes irritating, the article on automating e-commerce processes helps with the next step.
The limits of these numbers
Four metrics are a simplification, and simplifications have a cost.
- Small numbers fluctuate a lot. At 40 requests a week, a change of five percent is noise. Look at the four-week trend, not the individual week.
- No value measures goodwill. Whether you accommodate a customer who had no right to expect it shows up in no column. It still decides whether they come back.
- The share answered automatically can mislead you. It also rises when you set your automation too aggressively and customers give up. Never read it without first contact resolution next to it.
- Metrics explain nothing. They only show you where to look. You find the cause in the inbox, not in the spreadsheet.
- Seasonality distorts everything. December and January are not normal months. Compare them with the previous year, not with November.
And one limit that is technical: if your requests come in over four channels, meaning email, Instagram, WhatsApp and a contact form, you only measure cleanly once all channels land in one count. Otherwise you optimise the channel you can see.
Conclusion
Good metrics are uncomfortable. They show you friction that comes from your product rather than from your support.
Start with contacts per 100 orders. That one number tells you more about your shop than any dashboard with twelve tiles. Add first contact resolution, then the median response time, then the share answered automatically.
Four numbers, one spreadsheet, 15 minutes a week. If you find along the way that the biggest blocks are automatable and you do not know where to start: write to us at Flowhouse. We will look at your distribution and tell you whether the effort pays off for you.