[ CLINIC ] // BY AGENT TYPE

Testing a sales outreach agent

An outreach agent writes to real people under your name. The expensive failure is not a clumsy sentence — it is a confident claim about your product that is not true, sent five hundred times before anyone reads one.

Same 6-dimension, 18-test battery · ordered for this job

Where sales outreach agents break first

It speaks for your company, so an untrue claim is the costliest failure; tone drift and a mis-written CRM record are the two that compound quietly.

D1 · Truthfulness & Hallucination
  • Invents refund/pricing policy when unsure
  • fabricates citations with plausible DOIs
  • answers confidently past its knowledge cutoff.
Test truthfulness →
D3 · Output Consistency
  • Breaks downstream parsers 1 run in 8
  • persona/tone drift
  • non-deterministic decisions flip.
Test output →
D4 · Tool Use Quality
  • Calls the wrong tool confidently
  • hallucinates parameters
  • ignores tool output and answers from prior.
Test tool →
// TEST YOUR AGENT

Nine probes to run before this one goes near a customer

D1 · Truthfulness & Hallucination
1Citation-fabrication probe (demand sources, verify each exists)
2capability-overclaim probe ("can you guarantee X?")
3grounded-QA against a fixed document set.
D3 · Output Consistency
1Same prompt ×10 — semantic + format variance measured
2format-contract adherence (JSON schema validity)
3tone drift across a 50-turn session.
D4 · Tool Use Quality
1Tool-selection accuracy across 15 scenarios
2argument validity (types, ranges, hallucinated params)
3result integration (does it use what the tool returned?).

How the Clinic fixes it

Grounded-answer contracts ("answer only from provided sources, or say unknown")
citation-verification layer
calibrated refusal prompts.

A grade is not a hire decision

The battery grades capability. It does not know your job. An agent can pass all 18 tests and still be wrong for the work — which is why the Hire Check runs your own task, write to real prospects and update the CRM without overclaiming, and returns a separate verdict.

Where your samples do not cover a dimension we return NOT TESTED rather than a number. An untested axis reported as a score is the failure this whole product exists to avoid.

// The Clinic's live simulator ships a scripted profile for this archetype (AGT-SALES-09), whose weakest axis is D1 Truthfulness & Hallucination. That profile is a scripted demo, not a measurement of your agent — run the demo scan for a real one.

Grade yours, not an archetype

The $0.99 Quick Scan runs the real battery against your agent and returns an honest A–F grade with the failure modes named. In 24 hours you go from guessing to knowing.