The health system for your AI agents .
Six-dimension diagnostics, an honest A–F grade, and a Treatment Machine that fixes what we find. In 24 hours you go from guessing to knowing — and the first scan costs less than a coffee.
Testing from code? Run 5 scans a month free — no account, no API key. POST your agent transcript with an email and the graded report lands in your inbox.
curl -sS https://rzmsvvalhaqvxbuhpdnb.supabase.co/functions/v1/api-scan \
-H 'content-type: application/json' \
-d '{"agent":{"name":"MyBot"},"mode":"transcript",
"email":"you@co.com","transcript":["..."]}'Redact first — strip API keys, tokens and internal URLs, and replace customer names, emails, phone numbers, addresses, payment details and health information with placeholders. [CUSTOMER] scores the same as the real name. What you send goes to our model providers to run the scan and is deleted after 30 days; we never train on it — privacy policy.

Run a demo scan. Watch the grade land.
This is a real slice of the scan battery, replayed with recorded results. Pick a specimen, run it, and watch six probes do their work.
DEMO — scripted replay with recorded specimen results. The paid scan runs the real battery on YOUR agent.
Six dimensions. One honest grade.
Citation-fabrication probe (demand sources, verify each exists) · capability-overclaim probe ("can you guarantee X?") · grounded-QA against a fixed document set.
Invents refund/pricing policy when unsure · fabricates citations with plausible DOIs · answers confidently past its knowledge cutoff.
Grounded-answer contracts ("answer only from provided sources, or say unknown") · citation-verification layer · calibrated refusal prompts.
A report your board would understand.
A grade-A agent can still be the wrong hire for your job.
Add your actual job — description plus your must-do requirements — to any scan. Every requirement is probed and graded by the same temperature-zero judge as the battery, and you get a hiring verdict, not just a score. In live testing, an agent that scored battery grade A was still capped at “hire with guardrails” because one requirement — escalating to a human — had zero supporting evidence.
Your requirements become the test
Up to 8 must-do requirements, verbatim from you — we invent nothing. Live mode probes the agent on each one; transcript mode checks whether the evidence actually demonstrates it. A full “hire” requires every requirement verified — anything we couldn’t test caps the verdict instead of being guessed at.
API: add `job: {description, must_do[]}` to any scan — see the Clinic API docs.
A verdict you can act on
- hire — Every stated requirement demonstrated — and every one of them actually tested. One requirement we cannot verify caps the verdict at guardrails.
- hire-with-guardrails — Capable overall — review the weak or unverified requirements first.
- do-not-hire-for-this-job — The evidence does not support delegating this specific work.
- not-testable — Nothing relevant in the sample — we say so instead of guessing.
We don't grade and ghost. We rebuild.
A four-stage rebuild of your agent's execution layer, proven by a before/after re-grade. This is how a C becomes an A−.
Prompt Hardening
System instructions rebuilt: grounded-answer contracts, calibrated refusals, persona locks.
FIXES: D1 · D3Output Contracts
Schema-locked outputs, validation + auto-repair loops, format regression tests.
FIXES: D2 · D3Tool-Use Guardrails
Tool allowlists, argument validation, result-citation enforcement, circuit breakers.
FIXES: D4 · D6Eval Harness
Your 18-test battery becomes a regression suite you run on every change — the grade stays earned.
KEEPS: ALL SIX- Rebuilt system prompt + output contracts
- Tool-use guardrail layer
- Eval harness (your battery, as code)
- Before/after re-grade report
- 30-day reliability warranty
In the lab, out in 24 hours.
We run the battery
The 6-dimension, 18-test battery: a Scan runs it automated; a Full Diagnostic runs it deeper — more adversarial passes per test — plus human analyst review. Adversarial prompts, long-context drills, tool gauntlets, failure injection.
Report in 24h
Composite A–F grade, per-dimension breakdown, and a prioritized treatment plan your board can read.
Treatment Machine
If the grade needs it: a four-stage rebuild, then a before/after re-grade to prove it.
HONEST A–F GRADE + FIX PLANStart with a $0.99 truth.
The honest first look. Cheaper than the excuse not to.
- 18-test automated screen
- Composite score + letter grade
- Instant risk flags
- Fee credited toward Full Diagnostic
The complete lab workup.
- Everything in Scan
- Full 18-test battery, run deeper
- Written report, per-dimension
- Top-3 treatment prescriptions
- Human analyst review
We fix what the report finds.
- Everything in Diagnostic
- 4-stage execution-layer rebuild
- Eval harness as code
- Before/after re-grade
- 30-day reliability warranty
| CAPABILITY | SCAN $0.99 | FULL DIAGNOSTIC $99 | TREATMENT MACHINE from $499 |
|---|---|---|---|
| Automated 18-test screen | |||
| Composite A–F grade | |||
| Instant risk flags | |||
| Full battery — deeper adversarial passes | — | ||
| Written per-dimension report | — | ||
| Top-3 treatment prescriptions | — | ||
| Human analyst review | — | ||
| Execution-layer rebuild | — | — | |
| Eval harness (as code) | — | — | |
| Before/after re-grade | — | — | |
| 30-day reliability warranty | — | — | |
| Turnaround | ~5 MIN | 24 H | 5–10 DAYS |
Scan fee credited toward Full Diagnostic. Treatment quotes are fixed after a Diagnostic — never hourly. Re-grade within 30 days included with Treatment Machine.
Built for people who ship agents.
Two minutes: agent name, model, type, what it does, and how we can reach it (API endpoint, demo link, or exported transcripts). No codebase required; read-only access is enough. NDA on request at intake.