Articles
Everyone has a dashboard. Almost nobody has a detector.
Most teams aren't flying blind. They're watching the wrong half of the conversation

Over the last few weeks I shared a series of patterns from 80+ conversations with people who run conversational AI in production. Those posts were about what breaks: bots graded on finishing a flow, fluent answers that are quietly wrong, containment numbers that go up while customers give up.
This is about the part I didn't write, which turned out to be the more uncomfortable half.
Going in, I assumed the story would be that teams lack visibility. That they'd tell me they were flying blind and wanted better instrumentation. Almost nobody said that. Most of the people I spoke to had dashboards, sometimes very good ones. They still found out late.
What this is. These were LinkedIn conversations, not research interviews. People self-selected by replying to a founder who was obviously building something in this space, and my questions varied depending on who I was talking to. So no percentages here and no claim that any of it generalises. It's what I noticed, and where something came up often enough that I stopped being surprised, I've said so.
The dashboards exist. They just don't go anywhere.
The line that reframed this for me came from someone who has spent years doing conversational AI inside large enterprises:
"One thing I definitely noticed working in large enterprises is that dashboards often catch dust. Dashboards exist but teams don't spend enough resources to extract insights and translate them into user stories for the backlog. Setting up proper governance and ownership is very important here."
— Floris Van Cauwenberge, Sr. Conversational AI Designer, CBTW
Dashboards catch dust. I've thought about that phrase more than anything else anyone said to me.
The bottleneck he's describing isn't technical. Someone built the dashboard, it works, the numbers on it are correct. What's missing is the path from a number moving to something actually changing — a named person whose job it is to look, a route into the backlog, and enough time budgeted for the looking to happen. None of that is a tooling problem, and no tool fixes it.
A consultant who works with organisations moving from pilot to scale described the same gap from further up:
"The biggest reason for failure is governance and quality measurement. Most are doing a POC because it's hot and the shiny penny, but the majority are not realising the gain or even scaling it, because they are not continuing the constant iteration of measuring results."
Someone else, from the delivery side, added the version I heard most often in private: the hard part usually isn't the technology, it's the stakeholders, the scope creep, and non-technical people making technical decisions.
A metric's job is to route attention, not to render a verdict
I'd been asking people about the situation where the dashboard says fine and customers say otherwise. One CX practitioner declined to accept my framing:
"I don't see dashboards as something that should tell you what is wrong. Their primary purpose is to tell you where to look."
That correction is worth more than the question I asked. If you expect a dashboard to decide, you'll be disappointed and eventually stop trusting it. If you expect it to point, it can only ever be as useful as what happens after the pointing.
Which sets a hard test for any metric: can it tell you which fifty conversations to go and read? If it can't, it isn't earning its place, no matter how healthy it looks. A green number that doesn't generate an action is decoration.
What the metrics structurally cannot see
The sharpest inventory of blind spots came from someone who listed what he actually watches for. I'll quote it in full because the shape matters more than any single item:
"Rephrasing loops. Customer asks the same thing two or three times just worded differently. High completion, low sat — this is likely logged as a resolution, but the experience sucked. Silent repeat contacts. It's technically resolved because nothing escalates, but the customer comes back within 48 hours for the same question. Quiet abandonment. No rage-quit, no 'give me a human.' They just leave. Your escalation metrics look clean but they're gone."
Look at what all four have in common. In every case the standard dashboard reads healthy. Quiet abandonment is close to a perfect failure: the customer is gone, no complaint was filed, and the absence of the complaint registers as success.
This is why "the dashboard says fine" is so often literally true rather than a reporting error. The dashboard is accurately reporting the things it measures. The failures are simply not among them.
Detection is downstream
Ask people how they actually discovered a real production problem, and almost nobody describes a system telling them. They describe inference. Ticket volume that didn't fall the way it should have. Transfer rates shifting after a release. A colleague who happened to know the right answer and noticed the bot's number didn't match. Someone senior escalating.
Several people described exactly this progression: the failure was caught by a human, then understood by reading conversations, and only later, if ever, automated. One person put the limit of that channel plainly — it only works for the small subset of users who can verify the output for themselves.
Everyone else absorbs the failure silently and never enters your data at all. Which means the failure rate you're looking at is a floor, not an estimate. It counts the failures that happened to produce a visible secondary effect.
And the lag is not small. One practitioner made the point that some issues are only found after they have affected enough people to become statistically obvious — which is another way of saying the detection threshold is measured in customers.
What the teams who've got this working do differently
Not everyone is stuck. The ones who weren't had things in common, and none of them were exotic.
They alert on deviation from trend, not on fixed thresholds. One practitioner tracks the ratio between correctly and incorrectly understood requests, and treats a departure from its own history as the trigger:
"An issue is typically identified when this KPI deviates significantly from its historical trend. Such deviations indicate a potential degradation in understanding performance and trigger a deeper analysis to identify the underlying cause."
— Giuseppe Russo, Digital Product Manager – AI & Conversational Channels
A bot that has quietly degraded fifteen percent rarely trips an absolute limit, but it always breaks its own trend. Another practitioner described threshold alerts firing hourly across the whole funnel — expected conversations started, expected transactions completed, expected conversion between them — with the relevant conversations stored automatically for whoever picks it up.
They instrument the journey, not just the conversation. And more than one said the hard part was cultural:
"Telemetry is the most underestimated challenge. It's one thing to launch the assistant — it's another to get the development team to internalize a culture of measuring every step of the customer journey. That mindset shift has been harder to achieve than the technical build itself."
— Rodrigo, Tech Manager
They measure whether the process finished, not whether the conversation did. One team running a public-service assistant described tracking whether people actually completed the thing they came for — a registration, a filing, a payment — rather than whether the exchange ended tidily. Then they deliberately review the conversations where users dropped off, asked for a human, or repeated themselves.
They automate triage, not judgement. Several described using automated evaluation to flag candidates and keeping a human on whatever gets flagged. Notably, nobody described full automation as sufficient. Even the most instrumented teams had someone senior still reading conversations personally, and treated that as a feature rather than an embarrassment.
The thread
Look at that list again and there's one thing underneath all of it. Every practice people described is, at bottom, a strategy for deciding which conversations a human should read. Metrics to narrow the field. Alerts to catch the moment. Sampling to make the volume survivable. Automated flagging to shrink the queue.
That's a sensible adaptation to a real constraint, and the teams doing it well are doing genuinely good work. But it's worth naming for what it is: everyone is building increasingly clever ways to ration a scarce resource, and the scarce resource is attention.
Two people, separately and without seeing each other's answers, pointed at the same missing signal. Both suggested that the evidence you want is already sitting in the user's own turns — rephrasing, backtracking, repeated clarification, unnecessary round-trips — and that friction shows up there long before it shows up in CSAT or containment. One of them proposed measuring it directly, something like a Customer Effort Score inferred from the conversation itself.
I think they're right, and I think it points at the real gap. Nearly every metric we've built is an observation about the bot's own turns — what it said, and how well it said it. Very few are observations about the user's turns, about what the customer actually had to do to get what they came for. The evidence is sitting right there in the transcript, on the other side of the conversation.





