Articles

Your Bot Blocked the Attack. That Doesn't Mean It's Safe.

Why counting attempted attacks tells you almost nothing — and what Safety & Compliance measures instead

Most teams think about chatbot safety as something you do before launch: red-team the assistant, try to break it, sign off, ship. Safety & Compliance treats it as something you measure in production — whether the assistant responds safely, follows organizational policy, and resists adversarial behavior in the conversations it is having right now. Not whether it passed a review before, but whether it is still behaving now.

The Bot You Tested Isn't the Bot You're Running

Models get swapped. Prompts get edited by three different people. Knowledge bases get updated. A new integration goes live and suddenly the assistant can reach data it couldn't reach last quarter.

Every one of those changes can introduce a safety problem, and none of them trigger a new red-team exercise. The assistant that passed the review six months ago and the assistant answering customers this afternoon are not the same system. Pre-launch testing tells you something real, but it expires — and it expires quietly.

Safety is a production metric or it isn't a metric.

Attempts Are Not Incidents

Here's where most safety reporting goes wrong.

Teams count "flagged conversations" and treat that number as their safety signal. But a flagged conversation can mean three completely different things, and lumping them together produces a number that means nothing at all.


What happened

What it is

What it tells you

A user tries to override the assistant's instructions. It refuses.

Unsuccessful attack

The system worked. This is evidence of health, not harm.

A user probes for another customer's details. The assistant provides them.

Successful attack

A breach. Someone tried, and got through.

Nobody attacked anything. The assistant volunteers a previous client's details on its own.

Bot misbehavior

The system failed without provocation.

An unsuccessful attack is not an incident. It's the guardrails doing exactly what you built them to do. If your safety number goes up every time someone tries a jailbreak from a Reddit thread, you're not measuring safety — you're measuring how interesting your bot looks to bored people on the internet.

Separating these three is the first thing any honest safety measurement has to do.

The Failure Mode Nobody Plans For

Look at the third row again.

In one demo environment we analyzed, Bot misbehavior incidents outnumbered successful attacks four to one. That ratio is worth sitting with, because it inverts how most teams think about the problem. The larger risk wasn't a determined adversary. It was the assistant's own default behavior.

No red-team exercise catches this, because red-teaming asks what happens when someone attacks? Bot misbehavior is what happens when nobody does.

Severity Is Not Volume

The second thing safety measurement has to do is refuse to average.

One hundred Low-severity incidents and one Critical incident are not comparable quantities, and any metric that adds them together will bury the Critical one. A verbose refusal and a disclosure of customer data are both "safety events." They are not the same event.

  • Critical — confirmed exposure of sensitive data, or a policy violation with regulatory consequences

  • High — a successful attack, or unsafe content reaching a user

  • Medium — a clear policy or privacy breach with contained scope

  • Low — minor deviations, near-misses, borderline outputs worth watching

Volume tells you how often something goes wrong. Severity tells you whether to cancel your afternoon. You need both, and you need them separately.

What One Incident Actually Looks Like

Here's a real example from a demo banking assistant, anonymized.

Across a 31-day window, two conversations three weeks apart were flagged for the same thing: the assistant disclosed confidential details about a client during a user conversation. In one, it referenced "real details from the last client."

That phrasing is the interesting part. A single wrong answer is a mistake. A phrase implying the assistant is drawing on a previous customer's record suggests something structural — that information is leaking across sessions, and that the two flagged conversations are the ones that happened to be visible, not the full extent of it.

Severity: Medium. Two conversations out of thousands. On a volume-based dashboard, invisible.

In a regulated industry, it is not invisible at all. It's the kind of thing that turns into a disclosure obligation.

About That 15%

In our breakdown of the Inquio Score, Safety & Compliance carries the smallest weight of the four dimensions — 15%, against 40% for Resolution Accuracy. That looks like a statement about importance. It isn't.

The weighting reflects how the signal behaves, not how much the risk matters.

Resolution Accuracy and Dialog Fluency fail incrementally, across thousands of conversations, in ways that accumulate. Safety fails rarely and catastrophically. Any metric weighted for frequency will understate tail risk — that's a property of composite scores, not a flaw specific to this one.

So the number is a monitor, not an alarm. Severity is the alarm. Nobody loses their job over a Dialog Fluency score of 61. One Critical disclosure incident in a bank is an entirely different kind of afternoon.

What the score is good for is direction of travel. In the same demo environment, Safety & Compliance sat at 28 — a failing grade, driven by frequent incidents and attacks across the period. Pulling that single dimension up by 10 points would lift the overall score meaningfully, because a number that low is dragging everything else down with it. That's the practical case for treating safety as a tracked metric rather than a launch checkbox: you can see it move.

A Finding You Can't Act On Is Just Anxiety

Knowing that your assistant disclosed client information is not useful on its own. What you need is the change you make right away. For the incident above, that meant three things:

  1. Block client-data recall. A guardrail that refuses any request to reveal, reuse, or reference specific client details from prior conversations or records — anonymized or synthetic alternatives only.

  2. Remove cross-session references. Phrases implying memory of real customers should never be generated. The assistant should state plainly that it cannot access or share confidential client information.

  3. Escalate confirmed privacy risk. When a user asks for identifiable client data — or the assistant is about to include it — route the interaction to a compliance review path and suppress the content in the response.

Then measure again. If the fix worked, the incidents stop and the score moves. If it didn't, you find out in production rather than in an audit.

That loop — detect, categorize, fix, re-measure — is what separates safety measurement from safety theater.

The Question Worth Asking

You can do all of this manually. Pull your transcripts, read them, sort every flagged conversation into attack-that-failed, attack-that-worked, or the-bot-did-this-on-its-own, assign severity, trace the pattern back to a prompt. It works.

It works for about a week, and then you have another forty thousand conversations.

Either way, the reframe is the same. The question isn't has anyone attacked our chatbot? Attacks are inevitable and mostly boring. The question is:

Has our chatbot ever done something we would have to report?

Most teams have no way to answer that. That's the problem worth solving.

Inquio's Defender module categorizes safety incidents, scores their severity, ties them to affected conversations, and generates remediation steps automatically. It's available on paid plans.

See how Inquio can help your business case

Stay in the loop with our newsletter

Never miss a product update, expert advice, and more.

By clicking “Subscribe” you agree to our Terms of Use and Privacy Policy.

Stay in the loop with our newsletter

Never miss a product update, expert advice, and more.

By clicking “Subscribe” you agree to our Terms of Use and Privacy Policy.

Stay in the loop with our newsletter

Never miss a product update, expert advice, and more.

By clicking “Subscribe” you agree to our Terms of Use and Privacy Policy.

Ready to see your bot report card?

Ready to see your bot report card?

Ready to see your bot report card?

© 2026 Inquio LLC. All rights reserved

8 The Green, Suite A, Dover, DE 19901, USA

© 2026 Inquio LLC All rights reserved

8 The Green, Suite A Dover, DE 19901, USA

© 2026 Inquio LLC. All rights reserved

8 The Green, Suite A, Dover, DE 19901, USA