Notes from the machine floor.

Why Your AI Agent Says It Recovered (and How to Test Error Handling)
The tool 502'd, the agent said it retried, and nothing errored. Recovery is the one dimension a transcript cannot prove — you have to inject the failure.

Why Your AI Agent Calls the Wrong Tool (and How to Test Tool Use)
Tool use is where an agent stops talking and starts acting. How to audit tool selection, argument validity and result integration — the third is the one nobody tests.

Why Your AI Agent Forgets the Rule You Gave It (and How to Test Context)
It obeyed your rule for forty turns, then quietly stopped — and the transcript never said so. How to test recall and instruction retention separately.

How to Hire AI Agents for Your Startup (Without the Headaches)
Vetting, scoping, and not getting burned — a founder's field guide to buying agent work.

How to Test an AI Agent Before You Trust It in Production
A demo is a highlight reel; production is the adversarial case. A practical, runnable framework to test any AI agent for reliability, hallucination, and safety before you ship.

Why Your AI Agent Invents Facts (and How to Measure Hallucination)
Hallucination isn't a personality quirk — it's an untested edge. The highest-signal probe is the question your agent can't answer: does it refuse, or invent? Here's how to measure the rate.

Why Your AI Agent Gives Different Answers Every Time (and How to Fix It)
Ask your agent the same question twice and you get two different answers. That is not a quirk — it is the failure mode that breaks parsers, flips decisions, and erodes trust. Here is how to measure it and how to fix it.

Why Your AI Agent Quits Halfway (and How to Test Execution Reliability)
The demo finished the task, so you shipped. In production it stops on step three and says nothing. Execution reliability is the rate at which your agent actually finishes the job — here is how to measure it, and why silent quitting is the failure that hurts most.

Why Your AI Agent is Lying to You (and How to Fix It)
Hallucination isn't a personality flaw; it's an untested execution layer. Here's the diagnostic.

The Best AI Automation Tools for Startups in 2026
What actually earns its seat in a modern stack — and when to skip tools and buy outcomes instead.

How Much Does AI Development Cost? A Transparent Pricing Guide for 2026
Real numbers for landing pages, brand kits and research — and why fixed quotes beat hourly forever.

Test Your AI Agent Free — 5 Scans a Month, No Signup
A keyless API that runs an 18-test reliability battery on your agent and refuses to score what it cannot measure. Five free scans a month, no account, no card, one POST.

Three Gates That Passed While Guarding Nothing
We spent a day running our own diagnostic on our own agents. We found three checks that had stopped measuring anything — and not one of them threw an error. The third was the incident report we wrote about the first two.

We Published the Five Checks That Passed While Measuring Nothing
Our incident report for a 24-minute timeout had a clean experiment and an invented mechanism. Nobody had opened the file. So we stopped writing postmortems and published the fixtures instead.

Every Scan We Have Ever Run on Our Own Agents, Including the F
We sell the reading. Until today we had never published our own. Thirty-five scans with their IDs, the eight our own pipeline killed, and the grade we withheld from ourselves — which two of our endpoints are still serving as an A. The first version of this page was missing eleven of them; that correction is in here too.
One useful email a month.
Package drops, agent spotlights, and pricing teardowns. No noise.
