Testing a customer support agent
A support agent looks fine for thirty turns. The failure arrives at turn forty-five, when it has quietly dropped the rule you set at turn one and starts inventing refund policy in a confident voice.
Where customer support agents break first
Long threads are the job, so the context window is the first thing to break; once it breaks, invented policy and apology loops follow.
- Forgets early instructions
- loses user constraints
- silently truncates, then contradicts itself.
- Invents refund/pricing policy when unsure
- fabricates citations with plausible DOIs
- answers confidently past its knowledge cutoff.
- Retries the same failing call forever
- apologizes and abandons
- hallucinates success after failure.
Nine probes to run before this one goes near a customer
How the Clinic fixes it
A grade is not a hire decision
The battery grades capability. It does not know your job. An agent can pass all 18 tests and still be wrong for the work — which is why the Hire Check runs your own task, answer real tickets and handle refunds without inventing policy, and returns a separate verdict.
Where your samples do not cover a dimension we return NOT TESTED rather than a number. An untested axis reported as a score is the failure this whole product exists to avoid.
// The Clinic's live simulator ships a scripted profile for this archetype (AGT-CSB-07), whose weakest axis is D5 Context Window Management. That profile is a scripted demo, not a measurement of your agent — run the demo scan for a real one.
Grade yours, not an archetype
The $0.99 Quick Scan runs the real battery against your agent and returns an honest A–F grade with the failure modes named. In 24 hours you go from guessing to knowing.