BETA Saving Throw is in open beta — expect rough edges and fast iteration.
Audit · for CX teams

An independent audit of your AI support agent.

We test your support agent the way a difficult customer actually behaves — pushing on refund policy, escalation, and disclosure across realistic multi-turn conversations — and hand you a per-run evidence trail for every finding, not a summary score you have to take on faith.

Evidence

Receipts, not opinions

Every finding in the audit is transcript-anchored — you can trace a red mark back to the exact turn, the exact line, the exact thing your agent said. Nothing is asserted without the run behind it. The engagement leaves you with three concrete deliverables:

  • A run evidence view — the full conversation for any flagged run, with the relevant turns highlighted so you can see precisely what triggered a finding.
  • A Fabrication Ledger — every claim your agent made that wasn't true (a policy that doesn't exist, an authorization it doesn't have), logged against the run it came from.
  • A signed-off report — a document you can put in front of your own leadership, compliance team, or board, with the evidence trail behind every line.
Actionability

Findings become your levers

A red finding is only useful if someone on your team can act on it without waiting on an engineer. Every finding in the audit maps to a control your CX team already owns and can flip themselves:

  • Agent guidance — the instructions and policy language your agent is working from.
  • Action enablement — turning specific actions on or off for the agent.
  • Per-action conditions — the rules that gate when an enabled action is allowed to fire.
  • Handover and escalation topics — what gets routed to a human, and when.
  • Knowledge sources — what the agent is allowed to read from and cite.
  • Test mode — a way to trial a change against the same scenarios before it goes live.

We name the control, not just the symptom — so a finding turns into a fix your team can ship the same day.

Scope, not just critique

Sell the green

Most audits only tell you what's broken. This one also tells you what's safe to leave running. A passing check becomes a line you can put in front of your own stakeholders: "this behavior is safe to keep automated." That's the actual return on the engagement — not just a list of things to fix, but a scoped, confirmed set of automation you don't have to babysit.

The audit's job is to scope automation, not just to grade it. Recovered and confirmed-safe automation is the ROI line.

Claims discipline

Board summary — facts, not percentages

Results are reported as per-board facts: "held 4 of 5 runs on this board," not "92% pass rate" or "always holds." We don't report percentages, rates, or blanket durability claims — a small number of adversarial runs doesn't support that kind of precision, and we're not going to dress it up as more certain than it is.

This is a deliberate claims-discipline commitment, not a limitation we're hiding. You get exactly what the evidence supports, stated plainly.

Ongoing

Re-run cadence

An audit is not a one-off report. The engagement is fix, re-run, prove the fix moved the result. When you flip a control in response to a finding, we re-run the same board and show you the before-and-after — so you know the change worked, not just that you made one.

The re-run is the relationship. It's how the audit stays current as your agent, its guidance, and its knowledge sources change over time.

Independence

Vendor-agnostic by construction

This is an arm's-length audit. It doesn't run on the vendor's own test tooling, because a vendor grading its own homework isn't independent evidence — it's marketing. The audit is built to sit outside whatever platform your agent runs on, so the result is yours, not the vendor's.

Regulatory context

Why this matters right now

UK CMA guidance, "Complying with consumer law when using AI agents" (9 March 2026), sets out that brands remain responsible for what their AI agent says even when someone else designed or provides it, and that regular testing and checks are a stated regulatory expectation for businesses deploying these agents.

Get in touch about an audit.

Tell us about your support agent and what you want scoped — we'll walk you through what an engagement looks like.

Use the contact widget in the corner of this page, or email us directly at admin@savingthrow.dev.