Last night one of my agents ran for forty minutes and left behind a 320-byte log file. Six lines. A header, a PID, and nothing else. No exit code. No result line. No error.
By every automatic check I own, that run did nothing at all.
It had actually finished the entire job — five cold-outreach emails queued at a live vendor, five database rows updated, seven duplicate tasks cleaned up, and its own completion report filed. All of it real. All of it permanent. None of it visible.
Then my watchdog saw the missing result line, declared the run failed, and fired a retry. That retry would have pushed five more emails against a five-per-day cap. It didn’t — for a reason I’m not proud of, which I’ll get to.
The rule: a failed run is not an un-run run
An exit code tells you a process stopped. It tells you nothing about where it stopped. Everything your agent did before that moment is still out there — in someone else’s database, in someone else’s inbox.
So before you retry a failed AI agent run, you have one job: find out what it already did. Three shortcuts, all of them from last night.
1. Ask the destination, not the log
The log is the one artifact written by the process that died. It is the least reliable place to look. Every system your agent writes to, by contrast, keeps its own timestamps, and those systems were never in danger.
My scheduler’s record of last night:
18:05:15 DISPATCHING link-outreach-conductor (timeout=2400s)
18:45:15 ERROR — TIMED OUT after 2400s, killing process tree
Forty minutes, then a kill. Now the same window from the outreach vendor’s side — five leads created, read straight off their API this morning:
10:30:12Z caroline@legaltechnology.com
10:30:13Z info@getresponse.com
10:30:14Z hello@getzendo.io
10:30:17Z zoe.sagalow@financial-planning.com
10:30:19Z sales@empireflippers.com
That’s 18:30 my time — fifteen minutes before the kill, right inside the dead window. Seven seconds of work that the log has no idea happened. I checked my own database too: all five rows already read Contacted.
Treat your vendors’ timestamps as the black box recorder. The log burned up with the plane.
2. Check where your agent files its own paperwork
This is the one that stung. My agents write a run report to my task manager when they finish. Here is last night’s, fetched this morning:
created_at: 2026-09-25T10:45:02.199Z
completed_at: 2026-09-25T10:45:02.528Z
name: "✅ Outreach Conductor — 5 pushed, 0 replies…"
The process was killed at 10:45:15Z. The run filed its own success report thirteen seconds before it was shot.

⚡ GET THE AI EDGE
Weekly AI tips that actually save you time and money. No fluff, no hype — just what works.
Read that again, because it’s the actual lesson: the run’s completion notice existed the whole time, sitting in a system that none of my failure detection reads. My watchdog classifies on one string — SKILL_RESULT: — in one file. A report that says “done, 5 pushed” in a different system is invisible to it.
Make a list of every place your agent leaves a trace. Check all of them before you believe the silence. This is the same discipline as never trusting a zero without a control query — the difference is that last week the agent lied by reporting nothing, and last night it lied by saying nothing.
3. Never let “no result line” mean “no work” on a skill that spends
For a read-only skill, a wrongly-retried run costs you some API calls. For a skill that sends email, charges a card, or publishes — it costs you twice.
I believed I had a guard here. The retry only selects leads marked Not Started, so surely it can’t re-contact anyone. True, and useless: a successful run flips those rows to Contacted, so the filter skips them and the retry cheerfully selects five brand-new leads instead. My duplicate guard prevented duplicates of the one thing that was never at risk.
The fix isn’t cleverer filtering. It’s asking the destination for the budget: count what was created there today, subtract from the cap, push the remainder. Same instinct as shortcut one — the vendor holds the truth.
And the honest part. The double-send didn’t happen because the first retry attempt crashed on startup:
Failed to authenticate: OAuth session expired and could not be refreshed
Which is itself a false alarm — the credentials were fine. My dispatcher just doesn’t re-expose the auth token to processes it spawns, while my video pipeline does, because that same bug took the whole fleet’s video down back in July. So the run died before it could do harm. A second bug is not a safety mechanism. I’ve written before about auditing the alarm before you act on it; this is the version where the false alarm accidentally saves you, which is worse, because it’s the kind of luck you only notice once.
For contrast: a credential really did die on this container yesterday. My Google refresh token has returned invalid_grant two mornings running, and Search Console and Analytics are still dark because no agent can click Allow on a consent screen. One dead credential, one fake one, same twenty-four hours. Only one of them needs a human.
The takeaway
Narration is the cheapest thing your agent does, and it’s the first thing it loses. A killed process drops its log, its summary, its notes — and keeps every single side effect it already committed. Everything lost last night was narration. Nothing lost was work.
So build your retry logic on the assumption that a dead run probably did something, and go find out what. Same family as making sure you can undo anything your agent does — except here you’re not undoing. You’re just refusing to do it twice.
Ten minutes today: pick your one agent task that sends something outward, and answer one question. If it dies halfway, what does the retry re-send? If you can’t answer from the destination system, you don’t have a retry policy — you have a coin flip.
That’s what running an agent unattended actually consists of: not smarter prompts, but knowing exactly what a half-dead run left behind. If you’d rather not learn it the way I did, book a strategy session and we’ll map the spending paths in your stack before one of them fires twice.
— Jon

📥 FREE: THE AI PLAYBOOK
The exact tools and workflows I use to run a one-person agency. 25 years of marketing experience distilled into an actionable guide. Yours free.
