When a Human Has to Answer: The Limits of Support AI
A customer wrote to us saying her toothbrush had made her gums bleed. The AI would have produced a polite standard reply to that.
It did not, because that kind of enquiry never enters the automated path with us in the first place. That is what this is about.
At nano, mate and MUSTAX, AI answers around 65 percent of support enquiries automatically. That number gets celebrated a lot. I find the other 35 percent more interesting. They are a decision rather than a technical gap we still have to close.
The 65 percent is not an interim figure
When I talk to founders, almost the same question always comes up: why not 90 percent? Why not everything?
Because the remaining enquiries are not harder. They are more expensive when you get them wrong. A badly answered tracking question costs you a second email. A badly answered health question costs you the brand in the worst case.
The right question is never: what percentage runs automatically? It is: what happens in the worst case, and who carries it? We already worked that question through for automations generally in Human in the loop. In support it has its own edge, because at the other end sits a person who is currently unhappy.
What the automation handles for us is quickly told. Where is my parcel, which is around 40 percent of all enquiries. Plus delivery times, product questions with a clear answer, invoice copies, address changes before dispatch, and the statutory withdrawal notice German shops have to provide. All cases with exactly one right answer that exists in the system.
Four categories that may never answer automatically
Over the years we boiled this down to four groups. If one of them applies, the enquiry goes to a human immediately. No suggestion, no draft, no intermediate step.
First: health and safety. At nano we sell toothbrushes. Questions come in about bleeding gums, about implants, about children. An AI that even sounds like advice here is a liability risk. We answer with a human, and in case of doubt that human points to the dental practice.
Second: money without a clear rule. Refund outside the window, goodwill, partial credit, a double charge. Where a rule exists, the automation may apply it. Where somebody has to weigh things up, it may not.
Third: complaints with emotion. Not every claim is a complaint. But when somebody writes that they feel cheated, the perfectly worded reply is the problem. It reads like mockery.
Fourth: anything that sounds legal. Letters from lawyers, data protection requests, withdrawal with reference to a deadline, a threatened review used as leverage. Emails like that need vetted wording, not fast phrasing.
On top of that comes a fifth group that is not a category: anything the system cannot classify with confidence. Uncertainty is its own exit with us, never a reason to guess.
How the detection actually works
The categories are worthless if detection hangs on wording. So we work with three layers stacked on top of each other.
The bottom layer is hard stop words. Lawyer, lawsuit, doctor, allergy, hospital, child, fraud. If one of them appears, the automation is off. Full stop. That is crude, it produces false alarms, and it is still the most important layer, because it does not depend on the model.
The middle layer is classification by the AI itself. It assigns the enquiry to a category and states how confident it is. If confidence falls below our threshold, the case goes to a human.
The top layer is context from the shop. Is this customer here for the third time in four weeks? Has there already been a refund? Is the order above a threshold value? Then we escalate, even if the enquiry looks harmless. A technically simple case from a customer with three incidents is not a simple case.
These three layers are deliberately redundant. Each one catches something the others let through. The price is unnecessary escalations, and we pay it gladly. If you want to rebuild this in detail, the setup is in Pre-sorting support emails with AI.
The handover is the real sticking point
Most systems do not fail at detection. They fail in the second afterwards.
The classic: the AI writes two replies, then notices it is getting nowhere, and hands over. The customer now has two useless emails in their inbox and has to explain the question a third time. That is worse than no automation at all, because it burns trust on top.
Four rules we stick to:
- Hand over early. The decision falls on the first message, never after the third attempt.
- The customer is told. One sentence is enough: your message is now with a colleague, you will hear from us today. No excuse, no technical jargon.
- The context travels with it. Order number, history, what the AI has already checked, what it suspects. The human should decide rather than research.
- The way back is closed. A case that has landed with a human does not return to the automation on its own. Otherwise the AI writes into the same thread while the colleague is still typing.
And one more rule that is uncomfortable: the customer has to be able to tell they are writing to a machine. We do not put it in every line, and we never claim the opposite. Get caught pretending once and you lose exactly the customers who would otherwise have stayed.
What happens when you get it wrong
I have seen three mistakes, two of them at our own place.
The broken trust. A customer gets an obviously generated reply to a serious matter. They post the screenshot. The damage is bigger than the one reply. From then on every email of yours is under suspicion.
The quiet backlog. Escalation works, but nobody defined who handles the escalated case. The cases sit for two days. A question turns into a complaint. An automation project that builds the routing and forgets the ownership makes service measurably worse.
The false authority. The model invents a rule that does not exist and phrases it convincingly. Ours once promised a return window we never had. We honoured it, because everything else would have been worse. Why this sits in how these models work is in How to prevent AI hallucinations.
The counter-argument: too cautious is also wrong
I do not want to sound as though restraint is always right. It has a price, and that price rarely gets named.
Escalate everything out of fear and you have a slow system rather than a safe one. Response time is the number customers feel most directly. A correct answer after two days is worse for the customer than an automatic one after two minutes.
The second effect is internal. Anyone reviewing forty escalated cases a day stops reading properly around case ten. Then you have kept the risk and lost the time saving. Our rule of thumb: if more than one in three cases lands with a human, the rules are too coarse.
So we measure the escalation rate as seriously as the error rate. Which figures matter is in Customer service KPIs. The limit is a number you watch, never a setting you fix once.
What this means day to day
At mate a short list sits in Slack every morning. Of roughly thirty enquiries overnight, twenty are answered automatically, seven prepared as drafts, three escalated. Those three cost us maybe twenty minutes in total.
Those twenty minutes are the reason the whole thing works. The twenty automatic replies are not. Without the three serious cases in human hands, the automation would be dead after the first public misstep.
Conclusion
The line between machine and human is a question of how far you can fall, rather than a question of technology.
My advice if you are just starting: write down first what the AI must never answer. Only then what it may. The list of prohibitions is shorter, clearer and protects you better than any fine-tuning on the model. What the rest of the setup looks like is in the guide to customer service automation.
If you are unsure where your line should sit, talk to us at Flowhouse. We go through your real enquiries and also tell you what you can safely let run automatically.