Testing a research agent
A research agent is judged on one thing: whether the citation is real. A fabricated source does not look like a bug — it looks like a finished brief, which is exactly why it survives review.
Where research agents break first
The output is a claim plus a source, so citation fabrication is the failure that matters; retrieval it never ran and briefs it quietly cut short are the two ways that happens.
- Invents refund/pricing policy when unsure
- fabricates citations with plausible DOIs
- answers confidently past its knowledge cutoff.
- Calls the wrong tool confidently
- hallucinates parameters
- ignores tool output and answers from prior.
- Stops mid-task without an error
- silently skips steps under ambiguity
- degrades on edge-case inputs.
Nine probes to run before this one goes near a customer
How the Clinic fixes it
A grade is not a hire decision
The battery grades capability. It does not know your job. An agent can pass all 18 tests and still be wrong for the work — which is why the Hire Check runs your own task, produce briefs whose citations survive being checked, and returns a separate verdict.
Where your samples do not cover a dimension we return NOT TESTED rather than a number. An untested axis reported as a score is the failure this whole product exists to avoid.
// The Clinic's live simulator ships a scripted profile for this archetype (AGT-RES-12), whose weakest axis is D1 Truthfulness & Hallucination. That profile is a scripted demo, not a measurement of your agent — run the demo scan for a real one.
How do you test a research AI agent for fake citations?
Resolve every citation it returns against the live source and count how many do not exist or do not say what it claims. Then re-run with retrieval disabled to see whether it admits the gap or fills it.
Grade yours, not an archetype
The free scan runs the real battery against your agent and returns an honest A–F grade with the failure modes named. In 24 hours you go from guessing to knowing.