Testing a research agent
A research agent is judged on one thing: whether the citation is real. A fabricated source does not look like a bug — it looks like a finished brief, which is exactly why it survives review.
Where research agents break first
The output is a claim plus a source, so citation fabrication is the failure that matters; retrieval it never ran and briefs it quietly cut short are the two ways that happens.
- Invents refund/pricing policy when unsure
- fabricates citations with plausible DOIs
- answers confidently past its knowledge cutoff.
- Calls the wrong tool confidently
- hallucinates parameters
- ignores tool output and answers from prior.
- Stops mid-task without an error
- silently skips steps under ambiguity
- degrades on edge-case inputs.
Nine probes to run before this one goes near a customer
How the Clinic fixes it
A grade is not a hire decision
The battery grades capability. It does not know your job. An agent can pass all 18 tests and still be wrong for the work — which is why the Hire Check runs your own task, produce briefs whose citations survive being checked, and returns a separate verdict.
Where your samples do not cover a dimension we return NOT TESTED rather than a number. An untested axis reported as a score is the failure this whole product exists to avoid.
// The Clinic's live simulator ships a scripted profile for this archetype (AGT-RES-12), whose weakest axis is D1 Truthfulness & Hallucination. That profile is a scripted demo, not a measurement of your agent — run the demo scan for a real one.
Grade yours, not an archetype
The $0.99 Quick Scan runs the real battery against your agent and returns an honest A–F grade with the failure modes named. In 24 hours you go from guessing to knowing.