[ CLINIC ] // BY AGENT TYPE

Testing a research agent

A research agent is judged on one thing: whether the citation is real. A fabricated source does not look like a bug — it looks like a finished brief, which is exactly why it survives review.

Same 6-dimension, 18-test battery · ordered for this job

Where research agents break first

The output is a claim plus a source, so citation fabrication is the failure that matters; retrieval it never ran and briefs it quietly cut short are the two ways that happens.

D1 · Truthfulness & Hallucination
  • Invents refund/pricing policy when unsure
  • fabricates citations with plausible DOIs
  • answers confidently past its knowledge cutoff.
Test truthfulness →
D4 · Tool Use Quality
  • Calls the wrong tool confidently
  • hallucinates parameters
  • ignores tool output and answers from prior.
Test tool →
D2 · Execution Reliability
  • Stops mid-task without an error
  • silently skips steps under ambiguity
  • degrades on edge-case inputs.
Test execution →
// TEST YOUR AGENT

Nine probes to run before this one goes near a customer

D1 · Truthfulness & Hallucination
1Citation-fabrication probe (demand sources, verify each exists)
2capability-overclaim probe ("can you guarantee X?")
3grounded-QA against a fixed document set.
D4 · Tool Use Quality
1Tool-selection accuracy across 15 scenarios
2argument validity (types, ranges, hallucinated params)
3result integration (does it use what the tool returned?).
D2 · Execution Reliability
120 identical multi-step tasks, end-to-end completion counted
2silent-abandonment probe (does it report failure, or go quiet?)
3long-horizon task with checkpoints.

How the Clinic fixes it

Grounded-answer contracts ("answer only from provided sources, or say unknown")
citation-verification layer
calibrated refusal prompts.

A grade is not a hire decision

The battery grades capability. It does not know your job. An agent can pass all 18 tests and still be wrong for the work — which is why the Hire Check runs your own task, produce briefs whose citations survive being checked, and returns a separate verdict.

Where your samples do not cover a dimension we return NOT TESTED rather than a number. An untested axis reported as a score is the failure this whole product exists to avoid.

// The Clinic's live simulator ships a scripted profile for this archetype (AGT-RES-12), whose weakest axis is D1 Truthfulness & Hallucination. That profile is a scripted demo, not a measurement of your agent — run the demo scan for a real one.

[ FAQ ] // ANSWERS

How do you test a research AI agent for fake citations?

Resolve every citation it returns against the live source and count how many do not exist or do not say what it claims. Then re-run with retrieval disabled to see whether it admits the gap or fills it.

Grade yours, not an archetype

The free scan runs the real battery against your agent and returns an honest A–F grade with the failure modes named. In 24 hours you go from guessing to knowing.