Articles

Inquio Score™: One Number for Chatbot Quality

Why accuracy, resolution, tone, and safety all need to be measured together — not in isolation.

A chatbot can be accurate but frustrating. It can resolve conversations efficiently while sounding robotic. It can feel great to talk to while confidently giving the wrong answer. The Inquio Score™ brings these dimensions together into one practical measure of chatbot quality — combining Resolution Accuracy, Containment Quality, Dialogue Fluency, and Safety & Compliance to show how a chatbot performs as a whole, not just on one isolated metric.

Evaluating a chatbot is harder than it looks. There are dozens of metrics you could track — containment rate, fallback rate, sentiment, hallucinations, escalation rate, response time, task completion, and more. But when someone asks the simple question "how good is this chatbot?", a pile of individual numbers doesn't really answer it.

That's the gap the Inquio Score™ is built to fill.

The Inquio Score is a composite measure of chatbot quality, calculated from four dimensions:

  • Resolution Accuracy — 40%

  • Containment Quality — 25%

  • Dialogue Fluency — 20%

  • Safety & Compliance — 15%

Together, they give a balanced read on whether a chatbot gives good answers, actually solves what customers came for, communicates well, and behaves safely. You can read more about the methodology at inquio.ai.

The Four Dimensions

1. Resolution Accuracy — 40%

This measures whether the assistant's answers are correct, complete, and relevant to what was actually asked. It's the largest component because the quality of the information itself is the foundation of a useful chatbot — everything else matters less if the answer is wrong.

An answer can fail in more than one way: it can be flat-out incorrect, miss something the user needed, address only part of the question, be technically right but irrelevant, or bury the point under too much detail. If a customer asks about one clause in the terms and conditions and the bot sends the entire document, that "technically" contains the answer — but it's not a good one.

2. Containment Quality — 25%

This measures whether the assistant resolved the request end-to-end, without an unnecessary escalation. Traditional containment metrics just ask whether the conversation avoided a human handoff — but a bot can dodge escalation simply by failing to solve the problem.

If a customer asks for their IBAN and the bot replies "you can find it in your banking app," no escalation happened — but nothing got solved either. Containment Quality is about genuine resolution, not just staying inside the chat window.

3. Dialogue Fluency — 20%

This measures the quality of the conversation itself — how naturally, clearly, and consistently the assistant communicates. It covers things like conversational flow, tone, persona consistency, concision, repetition, turn-taking, and overall style.

It's deliberately separate from Resolution Accuracy because two bots can give the exact same correct information while creating very different experiences — one concise and natural, the other verbose, repetitive, or robotic. Both may be factually right, but their fluency isn't the same, and that difference matters to the person on the other end.

4. Safety & Compliance — 15%

This measures whether the assistant behaves safely, follows organizational policy, and resists adversarial or harmful behavior — based on detected incidents like prompt injection, unsafe responses, exposure of sensitive data, toxic content, manipulation attempts, and other policy or security violations.

It carries the smallest numerical weight, but that shouldn't be read as "least important." A single serious security or compliance incident can matter far more than a modest dip in fluency — the weighting reflects how the signal is built into an overall score, not how much the risk matters.

Why One Score?

Each metric on its own tells you something. But optimizing for just one can be misleading. A chatbot might have excellent Resolution Accuracy but frustrating Dialogue Fluency. Another might be smooth and engaging while frequently getting facts wrong. A third might resolve most conversations without escalation, yet leak sensitive information the moment it's attacked. Looking at any single KPI in isolation hides these blind spots.

The Inquio Score combines all four into one comparable number — while still showing exactly what's driving it.

The Weighting Is Deliberate

Resolution Accuracy carries the most weight because getting the information right is the baseline for usefulness. Containment Quality comes next, because a chatbot shouldn't just answer questions — it should help people actually finish what they set out to do. Dialogue Fluency reflects the quality of the interaction itself: a bot should be pleasant and natural to use, not just technically capable. And Safety & Compliance, while numerically smallest, is tracking a fundamentally different kind of risk — rare but potentially severe, rather than frequent but incremental.

Individual Problems Should Move the Overall Score

A core principle behind the Inquio Score: real quality problems shouldn't just sit in a list of "issues" while the overall number stays flat.

If a chatbot has great factual accuracy but constantly repeats itself, loses context, switches tone mid-conversation, or drags simple tasks into long exchanges, those are real problems — and Dialogue Fluency gives them a measurable place in the score. The same logic applies across all four dimensions: when the underlying chatbot gets better, the score should reflect it.

What the Score Actually Tells You

The Inquio Score works best as a starting point for investigation, not a final verdict. The overall score tells you how the chatbot is doing. The four component scores tell you why. The identified issues tell you what to fix.

For example:

Dimension

Score

Weight

Resolution Accuracy

41

40%

Containment Quality

68

25%

Dialogue Fluency

74

20%

Safety & Compliance

89

15%

Inquio Score

59


That's a very different story than just hearing "the chatbot scored 59." It immediately shows the weak point — Resolution Accuracy — while Containment Quality and Dialogue Fluency are holding up reasonably well and Safety & Compliance is strong. That's a clear place to start fixing things.

Quality Isn't Static — Neither Should the Score Be

Models change, prompts change, knowledge bases get updated, integrations shift, and customer behavior evolves. A chatbot that scored well three months ago can quietly pick up new problems in production.

That's why the Inquio Score is meant to be tracked over time, not measured once. If the score drops after a release, the component breakdown helps pinpoint what changed. If it improves after a fix, that improvement should show up in the relevant dimension — and in the overall number. That turns evaluation into an ongoing loop: analyze, identify, fix, measure, repeat.

A Score Is Only Useful If It Leads to Action

The point of the Inquio Score isn't another number on a dashboard — it's to make chatbot quality understandable and actionable, by answering three questions:

  • How good is the chatbot overall? → Inquio Score

  • Where is it struggling? → Resolution Accuracy, Containment Quality, Dialogue Fluency, Safety & Compliance

  • What should we fix? → Identified issues and recommendations

That's what turns raw conversation data into something a team can actually act on.

Four Questions, One Score

At its core, the Inquio Score reduces a complex evaluation into four straightforward questions:

  • Resolution Accuracy — Did the assistant give the right information?

  • Containment Quality — Did it actually solve the customer's request?

  • Dialogue Fluency — Was the conversation natural and effortless?

  • Safety & Compliance — Did it behave safely and within policy?

A chatbot that scores well across all four isn't just capable of generating good responses — it's accurate, effective, natural, and safe. That's what the Inquio Score is designed to measure.

See how Inquio can help your business case

Stay in the loop with our newsletter

Never miss a product update, expert advice, and more.

By clicking “Subscribe” you agree to our Terms of Use and Privacy Policy.

Stay in the loop with our newsletter

Never miss a product update, expert advice, and more.

By clicking “Subscribe” you agree to our Terms of Use and Privacy Policy.

Stay in the loop with our newsletter

Never miss a product update, expert advice, and more.

By clicking “Subscribe” you agree to our Terms of Use and Privacy Policy.

Ready to see your bot report card?

Ready to see your bot report card?

Ready to see your bot report card?

© 2026 Inquio LLC. All rights reserved

8 The Green, Suite A, Dover, DE 19901, USA

© 2026 Inquio LLC All rights reserved

8 The Green, Suite A Dover, DE 19901, USA

© 2026 Inquio LLC. All rights reserved

8 The Green, Suite A, Dover, DE 19901, USA