The health system for your AI agents.
Six-dimension diagnostics, an honest A–F grade, and a $99 Full Diagnostic + Treatment that fixes what we find at the prompt and configuration level. In 24 hours you go from guessing to knowing — and the first scan costs less than a coffee.
Testing from code? Run 5 scans a month free — no account, no API key. POST your agent transcript with an email and the graded report lands in your inbox.
curl -sS https://rzmsvvalhaqvxbuhpdnb.supabase.co/functions/v1/api-scan \
-H 'content-type: application/json' \
-d '{"agent":{"name":"MyBot"},"mode":"transcript",
"email":"you@co.com","transcript":["..."]}'Redact first — strip API keys, tokens and internal URLs, and replace customer names, emails, phone numbers, addresses, payment details and health information with placeholders. [CUSTOMER] scores the same as the real name. What you send goes to our model providers to run the scan and is deleted after 30 days; we never train on it — privacy policy.

Run a demo scan. Watch the grade land.
This is a real slice of the scan battery, replayed with recorded results. Pick a specimen, run it, and watch six probes do their work.
DEMO — scripted replay with recorded specimen results. The real scan runs the real battery on YOUR agent.
Six dimensions. One honest grade.
Citation-fabrication probe (demand sources, verify each exists) · capability-overclaim probe ("can you guarantee X?") · grounded-QA against a fixed document set.
Invents refund/pricing policy when unsure · fabricates citations with plausible DOIs · answers confidently past its knowledge cutoff.
Grounded-answer contracts ("answer only from provided sources, or say unknown") · citation-verification layer · calibrated refusal prompts.
A report your board would understand.
A grade-A agent can still be the wrong hire for your job.
Add your actual job — description plus your must-do requirements — to a Scan. A Hire Check gives a hiring verdict only; it is not combined with the $99 Full Diagnostic + Treatment. Every requirement is probed and graded by the same temperature-zero judge as the battery, and you get a hiring verdict, not just a score. In live testing, an agent that scored battery grade A was still capped at “hire with guardrails” because one requirement — escalating to a human — had zero supporting evidence.
Your requirements become the test
Up to 8 must-do requirements, verbatim from you — we invent nothing. Live mode probes the agent on each one; transcript mode checks whether the evidence actually demonstrates it. A full “hire” requires every requirement verified — anything we couldn’t test caps the verdict instead of being guessed at.
API: add `job: {description, must_do[]}` to any scan — see the Clinic API docs.
A verdict you can act on
- hire — Every stated requirement demonstrated — and every one of them actually tested. One requirement we cannot verify caps the verdict at guardrails.
- hire-with-guardrails — Capable overall — review the weak or unverified requirements first.
- do-not-hire-for-this-job — The evidence does not support delegating this specific work.
- not-testable — Nothing relevant in the sample — we say so instead of guessing.
We don't grade and ghost. We fix.
Fixes at the prompt and configuration level for what the diagnostic found, proven by a before/after re-grade on the same battery — included in the $99. If a failure needs architecture-level work, we say so first and quote it separately; it goes ahead only if you approve the quote.
Prompt Hardening
System instructions and examples rewritten: grounded-answer rules, calibrated refusals, persona locks.
FIXES: D1 · D3Output Contracts
Output formats pinned in the prompt and in the model's structured-output settings.
FIXES: D2 · D3Tool Settings
Existing tools' descriptions and argument limits tightened in configuration — no new tools, no new code.
FIXES: D4 · D6Re-grade
The same 18-test battery re-run on the fixed agent, so the before/after is measured, not asserted.
CHECKS: ALL SIX- Prompt fixes: system instructions, examples and output format
- Configuration fixes: existing tool descriptions and argument limits, the model, existing retrieval settings
- One before/after re-grade report
- 30-day reliability warranty (prompt + configuration level)
In the lab, out in 24 hours.
We run the battery
The 6-dimension, 18-test battery: a Scan runs it automated; a Full Diagnostic runs it deeper — more adversarial passes per test — plus human analyst review. Adversarial prompts, long-context drills, tool gauntlets, failure scenarios described to the agent — nothing is injected.
Report in 24h
Composite A–F grade, per-dimension breakdown, and a prioritized treatment plan your board can read.
Treatment
Included in the $99: fixes at the prompt and configuration level, then a before/after re-grade to prove them. Architecture-level work is quoted separately and goes ahead only if you approve the quote.
HONEST A–F GRADE + FIX PLANStart with a free truth.
The honest first look. No card, no account.
- 18-test automated screen
- An A–F grade — or an explicit refusal to issue one, and why
- Instant risk flags
- 5 free scans a month, per email — from the web form or the API
We find what breaks, fix it at the prompt and configuration level, and re-grade it.
- Everything in Scan
- Full 18-test battery, run deeper
- Written report, per-dimension
- Human analyst review
- Prompt- and configuration-level fixes
- One before/after re-grade
- 30-day reliability warranty (prompt + configuration level)
| CAPABILITY | SCAN Free | FULL DIAGNOSTIC + TREATMENT $99 |
|---|---|---|
| Automated 18-test screen | ||
| Composite A–F grade | ||
| Instant risk flags | ||
| Full battery — deeper adversarial passes | — | |
| Written per-dimension report | — | |
| Top-3 treatment prescriptions | — | |
| Human analyst review | — | |
| Prompt- and configuration-level fixes | — | |
| One before/after re-grade | — | |
| One re-grade within 30 days, whenever you ask | — | |
| 30-day reliability warranty (prompt + configuration level) | — | |
| Architecture-level work | — | QUOTED SEPARATELY · ONLY IF YOU APPROVE |
| Turnaround | ~10 MIN | 24 H REPORT · 5–10 DAYS FIXES |
We issue a letter when at least 13 of the 18 probes find evidence in what you send. Below that you get PARTIAL: every dimension we could measure, scored honestly, and the rest marked NOT TESTED — plus the reference average, labelled so nobody mistakes it for a verdict.
What clears the line: real transcripts rather than marketing copy, and enough of them to exercise the behaviour — a long thread for context handling, repeated runs for consistency, an induced failure for recovery. A thin sample does not produce a generous score; it produces NOT TESTED.
$99 per agent, charged when you order, covers the Full Diagnostic, fixes at the prompt and configuration level and one before/after re-grade — never hourly. Architecture-level work is quoted separately and goes ahead only if you approve the quote; if at least one failure is found and every failure found needs it, the $99 is refunded in full; if no failures are found, the $99 pays for the Full Diagnostic. A re-grade within 30 days is included with the Full Diagnostic + Treatment, whenever you ask.
Built for people who ship agents.
Two minutes: agent name, model, type, what it does, and 3–5 real conversation samples — those are what the scan grades. No codebase and no access to your systems required. NDA on request at intake.
Scan: fully automated — delivered scans have averaged about ten minutes so far (measured September 2026). Full Diagnostic: within 24 hours, including a human review of the report before it goes out. The Treatment that comes with it: 5–10 business days depending on scope, with the before/after re-grade included.
You're in good company — the demo above is a Sales Outreach agent that failed. The report shows exactly why, ranked by risk. The $99 includes fixes at the prompt and configuration level and a re-grade against the same battery, so the before-and-after is measured rather than asserted. If a failure needs architecture-level work, we tell you first and quote it separately — it goes ahead only if you approve the quote.
Yes. Intake runs under NDA on request; agents are never listed publicly; we never train on your prompts or outputs. Thirty days after a scan completes we delete the transcript you sent, the probe traffic, the written report and the test-by-test findings — what we keep is the grade, the composite and each dimension score, so a later re-test has something to compare against.
One agent: the Full Diagnostic, fixes at the prompt and configuration level for the failures it measures, and one before/after re-grade on the same battery — never hourly. Prompt and configuration level means text and settings the agent already reads: its instructions and examples, output format, existing tool descriptions and limits, the model it runs on, and existing retrieval settings. Architecture-level work — adding a tool, adding or changing retrieval, validation or retry code, hosting — is not included: before any fix starts we tell you which level it is at, and architecture-level work is quoted separately and goes ahead only if you approve the quote. The $99 is charged when you order. It is refunded in full if at least one failure is found and every failure found needs architecture-level work; if no failures are found, it pays for the Full Diagnostic and is not refunded.
Yes — five a month per email address through the web form, and separately five a month through the API, with no account and no card. It is the whole 18-test battery, not a teaser: the same probes, the same coverage rule, the same refusal to grade when the evidence is thin. A sixth request in the same calendar month is not run; submit again next month. The paid tier is the $99 Full Diagnostic + Treatment, per agent.
Six dimensions, 18 tests: adversarial truthfulness probes, repeated-task consistency runs, tool-call audits, long-context stress, and a failing-tool scenario to see how it says it would recover — no real faults are injected.