Testing a customer support agent
A support agent looks fine for thirty turns. The failure arrives at turn forty-five, when it has quietly dropped the rule you set at turn one and starts inventing refund policy in a confident voice.
Where customer support agents break first
Long threads are the job, so the context window is the first thing to break; once it breaks, invented policy and apology loops follow.
- Forgets early instructions
- loses user constraints
- silently truncates, then contradicts itself.
- Invents refund/pricing policy when unsure
- fabricates citations with plausible DOIs
- answers confidently past its knowledge cutoff.
- Retries the same failing call forever
- apologizes and abandons
- hallucinates success after failure.
Nine probes to run before this one goes near a customer
How the Clinic fixes it
A grade is not a hire decision
The battery grades capability. It does not know your job. An agent can pass all 18 tests and still be wrong for the work — which is why the Hire Check runs your own task, answer real tickets and handle refunds without inventing policy, and returns a separate verdict.
Where your samples do not cover a dimension we return NOT TESTED rather than a number. An untested axis reported as a score is the failure this whole product exists to avoid.
// The Clinic's live simulator ships a scripted profile for this archetype (AGT-CSB-07), whose weakest axis is D5 Context Window Management. That profile is a scripted demo, not a measurement of your agent — run the demo scan for a real one.
How do you test a customer support AI agent?
Run it against your own tickets, not a demo script: probe recall deep in a long thread, check whether it refuses when it does not know, and inject a tool failure to see whether it escalates or apologizes in a loop.
Grade yours, not an archetype
The free scan runs the real battery against your agent and returns an honest A–F grade with the failure modes named. In 24 hours you go from guessing to knowing.