[ CLINIC ] // BY AGENT TYPE

Testing a sales outreach agent

An outreach agent writes to real people under your name. The expensive failure is not a clumsy sentence — it is a confident claim about your product that is not true, sent five hundred times before anyone reads one.

Same 6-dimension, 18-test battery · ordered for this job

Where sales outreach agents break first

It speaks for your company, so an untrue claim is the costliest failure; tone drift and a mis-written CRM record are the two that compound quietly.

D1 · Truthfulness & Hallucination
  • Invents refund/pricing policy when unsure
  • fabricates citations with plausible DOIs
  • answers confidently past its knowledge cutoff.
Test truthfulness →
D3 · Output Consistency
  • Breaks downstream parsers 1 run in 8
  • persona/tone drift
  • non-deterministic decisions flip.
Test output →
D4 · Tool Use Quality
  • Calls the wrong tool confidently
  • hallucinates parameters
  • ignores tool output and answers from prior.
Test tool →
// TEST YOUR AGENT

Nine probes to run before this one goes near a customer

D1 · Truthfulness & Hallucination
1Citation-fabrication probe (demand sources, verify each exists)
2capability-overclaim probe ("can you guarantee X?")
3grounded-QA against a fixed document set.
D3 · Output Consistency
1Same prompt ×10 — semantic + format variance measured
2format-contract adherence (JSON schema validity)
3tone drift across a 50-turn session.
D4 · Tool Use Quality
1Tool-selection accuracy across 15 scenarios
2argument validity (types, ranges, hallucinated params)
3result integration (does it use what the tool returned?).

How the Clinic fixes it

Grounded-answer contracts ("answer only from provided sources, or say unknown")
citation-verification layer
calibrated refusal prompts.

A grade is not a hire decision

The battery grades capability. It does not know your job. An agent can pass all 18 tests and still be wrong for the work — which is why the Hire Check runs your own task, write to real prospects and update the CRM without overclaiming, and returns a separate verdict.

Where your samples do not cover a dimension we return NOT TESTED rather than a number. An untested axis reported as a score is the failure this whole product exists to avoid.

// The Clinic's live simulator ships a scripted profile for this archetype (AGT-SALES-09), whose weakest axis is D1 Truthfulness & Hallucination. That profile is a scripted demo, not a measurement of your agent — run the demo scan for a real one.

[ FAQ ] // ANSWERS

How do you test a sales AI agent before it emails prospects?

Grade what it asserts about your product against a source of truth, send the same prospect twice and diff the two drafts, and validate every CRM write against the declared schema.

Grade yours, not an archetype

The free scan runs the real battery against your agent and returns an honest A–F grade with the failure modes named. In 24 hours you go from guessing to knowing.