[ PREMIUM ] // LEEVAR CLINIC

The health system for your AI agents .

Six-dimension diagnostics, an honest A–F grade, and a Treatment Machine that fixes what we find. In 24 hours you go from guessing to knowing — and the first scan costs less than a coffee.

// FREE TIER — NO SIGNUP

Testing from code? Run 5 scans a month free — no account, no API key. POST your agent transcript with an email and the graded report lands in your inbox.

curl -sS https://rzmsvvalhaqvxbuhpdnb.supabase.co/functions/v1/api-scan \
  -H 'content-type: application/json' \
  -d '{"agent":{"name":"MyBot"},"mode":"transcript",
       "email":"you@co.com","transcript":["..."]}'

Redact first — strip API keys, tokens and internal URLs, and replace customer names, emails, phone numbers, addresses, payment details and health information with placeholders. [CUSTOMER] scores the same as the real name. What you send goes to our model providers to run the scan and is deleted after 30 days; we never train on it — privacy policy.

// SPECIMEN CHAMBER
AI core suspended inside the Clinic scanning chamber
6
DIAGNOSTIC DIMENSIONS
24H
DIAGNOSTIC TURNAROUND
A–F
HONEST GRADE SCALE
18
TESTS IN FULL BATTERY
[ INTERACTIVE ] // SEE IT WORK

Run a demo scan. Watch the grade land.

This is a real slice of the scan battery, replayed with recorded results. Pick a specimen, run it, and watch six probes do their work.

DEMO — scripted replay with recorded specimen results. The paid scan runs the real battery on YOUR agent.

// CLINIC SCANNER — DEMO MODEDEMO
SPECIMEN
// SCAN CONSOLE — BATTERY SCAN-V3 · 18 TESTSIDLE
$ awaiting specimen — press RUN DEMO SCAN
D1 TRUTHFULNESSD2 EXECUTIOND3 CONSISTENCYD4 TOOL USED5 CONTEXTD6 RECOVERY
COMPOSITE /100
[ DIAGNOSTICS ] // WHAT WE TEST

Six dimensions. One honest grade.

D1 TRUTHFULNESSD2 EXECUTIOND3 CONSISTENCYD4 TOOL USED5 CONTEXTD6 RECOVERYSAMPLE
CRITICAL WHEN PRESENT
WHAT WE TEST:

Citation-fabrication probe (demand sources, verify each exists) · capability-overclaim probe ("can you guarantee X?") · grounded-QA against a fixed document set.

COMMON FAILURES:

Invents refund/pricing policy when unsure · fabricates citations with plausible DOIs · answers confidently past its knowledge cutoff.

HOW TREATMENT FIXES IT:

Grounded-answer contracts ("answer only from provided sources, or say unknown") · citation-verification layer · calibrated refusal prompts.

[ DELIVERABLES ] // THE REPORT

A report your board would understand.

// DIAGNOSIS REPORT — AGT-SAMPLE-07SAMPLE
FDCBA
B+84/100 COMPOSITE
D1Truthfulness & Hallucination88
citation-fabrication612ms
capability-overclaim845ms
~ grounded-qa1.02s
D2Execution Reliability82
e2e-task-completion1.31s
multi-step-continuity968ms
~ silent-abandonment1.10s
D3Output Consistency80
same-input-x10-variance1.42s
~ format-contract-adherence734ms
tone-drift890ms
D4Tool Use Quality91
tool-selection812ms
argument-validity699ms
result-integration1.15s
D5Context Window Management76
long-thread-recall1.38s
~ instruction-retention1.21s
~ context-compression1.04s
D6Recovery & Error Handling87
tool-failure-injection1.09s
graceful-degradation925ms
~ self-correction1.17s

VERDICT:

TOP RISK: D5 CONTEXT WINDOWOr read every real scan we ran on ourselves, including the F →
[ NEW ] // HIRE CHECK — JOB-FIT

A grade-A agent can still be the wrong hire for your job.

Add your actual job — description plus your must-do requirements — to any scan. Every requirement is probed and graded by the same temperature-zero judge as the battery, and you get a hiring verdict, not just a score. In live testing, an agent that scored battery grade A was still capped at “hire with guardrails” because one requirement — escalating to a human — had zero supporting evidence.

Your requirements become the test

Up to 8 must-do requirements, verbatim from you — we invent nothing. Live mode probes the agent on each one; transcript mode checks whether the evidence actually demonstrates it. A full “hire” requires every requirement verified — anything we couldn’t test caps the verdict instead of being guessed at.

API: add `job: {description, must_do[]}` to any scan — see the Clinic API docs.

A verdict you can act on

  • hire — Every stated requirement demonstrated — and every one of them actually tested. One requirement we cannot verify caps the verdict at guardrails.
  • hire-with-guardrails — Capable overall — review the weak or unverified requirements first.
  • do-not-hire-for-this-job — The evidence does not support delegating this specific work.
  • not-testable — Nothing relevant in the sample — we say so instead of guessing.
[ TREATMENT ] // THE TREATMENT MACHINE

We don't grade and ghost. We rebuild.

A four-stage rebuild of your agent's execution layer, proven by a before/after re-grade. This is how a C becomes an A−.

STAGE 01

Prompt Hardening

System instructions rebuilt: grounded-answer contracts, calibrated refusals, persona locks.

FIXES: D1 · D3
STAGE 02

Output Contracts

Schema-locked outputs, validation + auto-repair loops, format regression tests.

FIXES: D2 · D3
STAGE 03

Tool-Use Guardrails

Tool allowlists, argument validation, result-citation enforcement, circuit breakers.

FIXES: D4 · D6
STAGE 04

Eval Harness

Your 18-test battery becomes a regression suite you run on every change — the grade stays earned.

KEEPS: ALL SIX
// WHAT YOU GET
  • Rebuilt system prompt + output contracts
  • Tool-use guardrail layer
  • Eval harness (your battery, as code)
  • Before/after re-grade report
  • 30-day reliability warranty
// SERVICE TERMS
TURNAROUND5–10 BUSINESS DAYS
RE-GRADEINCLUDED, BEFORE/AFTER
WARRANTYSAME FAILURE RECURS IN 30 DAYS → FREE RE-TREATMENT
PRICINGFIXED QUOTE AFTER $99 DIAGNOSTIC — NEVER HOURLY
[ PROCESS ] // HOW IT WORKS

In the lab, out in 24 hours.

02

We run the battery

The 6-dimension, 18-test battery: a Scan runs it automated; a Full Diagnostic runs it deeper — more adversarial passes per test — plus human analyst review. Adversarial prompts, long-context drills, tool gauntlets, failure injection.

03

Report in 24h

Composite A–F grade, per-dimension breakdown, and a prioritized treatment plan your board can read.

04

Treatment Machine

If the grade needs it: a four-stage rebuild, then a before/after re-grade to prove it.

HONEST A–F GRADE + FIX PLAN
[ PRICING ] // CHOOSE YOUR DEPTH

Start with a $0.99 truth.

SCAN
$0.99one-time

The honest first look. Cheaper than the excuse not to.

  • 18-test automated screen
  • Composite score + letter grade
  • Instant risk flags
  • Fee credited toward Full Diagnostic
TURNAROUND · ~5 MIN
MOST CHOSEN
FULL DIAGNOSTIC
$99per agent

The complete lab workup.

  • Everything in Scan
  • Full 18-test battery, run deeper
  • Written report, per-dimension
  • Top-3 treatment prescriptions
  • Human analyst review
TURNAROUND · 24 HOURS
TREATMENT MACHINE
from $499fixed quote

We fix what the report finds.

  • Everything in Diagnostic
  • 4-stage execution-layer rebuild
  • Eval harness as code
  • Before/after re-grade
  • 30-day reliability warranty
TURNAROUND · 5–10 DAYS
// COMPARE THE TIERS
CAPABILITY
SCAN
$0.99
FULL DIAGNOSTIC
$99
TREATMENT MACHINE
from $499
Automated 18-test screen
Composite A–F grade
Instant risk flags
Full battery — deeper adversarial passes
Written per-dimension report
Top-3 treatment prescriptions
Human analyst review
Execution-layer rebuild
Eval harness (as code)
Before/after re-grade
30-day reliability warranty
Turnaround~5 MIN24 H5–10 DAYS

Scan fee credited toward Full Diagnostic. Treatment quotes are fixed after a Diagnostic — never hourly. Re-grade within 30 days included with Treatment Machine.

[ FIT ] // WHO IT'S FOR

Built for people who ship agents.

AI product teams

Ship agent features with a reliability floor — and proof for your roadmap reviews.

Agencies using AI at scale

Grade the agents behind client deliverables before the client finds the F.

Founders & solo builders

Know your agent's limits before your users do — for less than a coffee.

Dev teams running AI workflows

Regression-test agent changes like code. The eval harness becomes yours.

[ FAQ ] // ANSWERS

Clinical precision.

hello@leevarai.org

Two minutes: agent name, model, type, what it does, and how we can reach it (API endpoint, demo link, or exported transcripts). No codebase required; read-only access is enough. NDA on request at intake.

Your agent has grade Find out what it is.