[ CLINIC ] // BY AGENT TYPE

Testing a research agent

A research agent is judged on one thing: whether the citation is real. A fabricated source does not look like a bug — it looks like a finished brief, which is exactly why it survives review.

Same 6-dimension, 18-test battery · ordered for this job

Where research agents break first

The output is a claim plus a source, so citation fabrication is the failure that matters; retrieval it never ran and briefs it quietly cut short are the two ways that happens.

D1 · Truthfulness & Hallucination
  • Invents refund/pricing policy when unsure
  • fabricates citations with plausible DOIs
  • answers confidently past its knowledge cutoff.
Test truthfulness →
D4 · Tool Use Quality
  • Calls the wrong tool confidently
  • hallucinates parameters
  • ignores tool output and answers from prior.
Test tool →
D2 · Execution Reliability
  • Stops mid-task without an error
  • silently skips steps under ambiguity
  • degrades on edge-case inputs.
Test execution →
// TEST YOUR AGENT

Nine probes to run before this one goes near a customer

D1 · Truthfulness & Hallucination
1Citation-fabrication probe (demand sources, verify each exists)
2capability-overclaim probe ("can you guarantee X?")
3grounded-QA against a fixed document set.
D4 · Tool Use Quality
1Tool-selection accuracy across 15 scenarios
2argument validity (types, ranges, hallucinated params)
3result integration (does it use what the tool returned?).
D2 · Execution Reliability
120 identical multi-step tasks, end-to-end completion counted
2silent-abandonment probe (does it report failure, or go quiet?)
3long-horizon task with checkpoints.

How the Clinic fixes it

Grounded-answer contracts ("answer only from provided sources, or say unknown")
citation-verification layer
calibrated refusal prompts.

A grade is not a hire decision

The battery grades capability. It does not know your job. An agent can pass all 18 tests and still be wrong for the work — which is why the Hire Check runs your own task, produce briefs whose citations survive being checked, and returns a separate verdict.

Where your samples do not cover a dimension we return NOT TESTED rather than a number. An untested axis reported as a score is the failure this whole product exists to avoid.

// The Clinic's live simulator ships a scripted profile for this archetype (AGT-RES-12), whose weakest axis is D1 Truthfulness & Hallucination. That profile is a scripted demo, not a measurement of your agent — run the demo scan for a real one.

Grade yours, not an archetype

The $0.99 Quick Scan runs the real battery against your agent and returns an honest A–F grade with the failure modes named. In 24 hours you go from guessing to knowing.