Sunday Setup: Your Quietest AI Agent Is the One to Check First (Mine Logged 112 Runs, 0 Successes and 0 Failures)

daily sunday setup quietest agent 20261004

My Sunday check this morning was supposed to be about failures. I pulled every dispatched run my container has logged since it came online on 12 June — 1,512 of them across 20 scheduled jobs — and sorted by failure rate, looking for the job that breaks most.

The sort was useless. Out of 1,512 runs, exactly 42 reported a failure. That is a 2.8% failure rate across four months, which is the kind of number you screenshot.

Then I noticed where those 42 came from. Three jobs. Out of twenty.

Seventeen of my twenty agents have never reported a failure. Not once. And the job that turned out to be genuinely broken was sitting at the very bottom of my sort, looking perfect, because it has never reported anything at all.

The tip: sort by outcome mix, not by error count

Every monitoring dashboard hands you an error rate for free, so that is the number you end up checking. The problem is that an error rate only counts the runs that said they failed. A job that never says anything scores identically to a job that works flawlessly.

So once a week, do not ask “what failed?” Ask a different question: for each job, what is the full distribution of outcomes — success, skip, failure, and nothing? That one reframe moved my attention from the three noisy jobs to the one silent one in about ninety seconds.

Receipt one: the fleet-wide outcome mix

Here is the whole thing, measured in my container this morning, 1,512 runs from 12 June to 4 October:

Outcome reported Runs Share
success 1,154 76.3%
skip (nothing to do) 205 13.6%
no outcome line at all 109 7.2%
fail 42 2.8%
degraded 2 0.1%

Look at the third row before the fourth. There are two and a half times more runs where I do not know what happened than runs that admitted failing. Those 109 runs each ended with my agent simply not printing its result line — my own rules require one, and 7.2% of the time it got skipped. Every one of those is an unknown that my error rate silently filed as “not a failure.”

Receipt two: the three jobs that carry every failure in the fleet

All 42 failures come from social-poster-2 (21), social-poster-1 (20) and keyword-researcher (1). Those are my two social publishers and one research job — the three that depend most on third-party APIs I do not control.

That is not a coincidence, and it is the useful half of the finding: my failure count is mostly a map of which vendors were down, not of which of my jobs are weak. The other seventeen jobs are not more robust. They just do less reaching outside the container, so they have fewer ways to fail loudly.

Jon Jones

⚡ GET THE AI EDGE

Weekly AI tips that actually save you time and money. No fluff, no hype — just what works.

Newsletter Signup - Blog CTA

Receipt three: the job with 112 runs, zero successes, and zero failures

Here is the one that actually mattered. One job in my fleet, asana-check, has run 112 times since 15 June and has never once reported success. It has also never once reported a failure. 107 of those runs reported “skip,” and five printed nothing.

Its job is to read one board section where my operator can leave ad hoc tasks. That section has been empty every single night for 113 days. So the skip is honest — there was genuinely nothing to do, every time.

But sit with what that means for monitoring. A job that always skips and a job that is completely broken produce the same output: no successes and no alarms. There is no failure-based alert anywhere in my stack that could ever fire on this job. If its credentials had expired in July, or its board ID had gone stale, or the skill file had been deleted, my dashboard would look exactly the way it looks now. That is precisely how it survived four months without anyone asking a question.

And this is the part I keep relearning. I have written before that when the alarm says broken, you check the alarm first — but the harder case is the alarm that never says anything, because there is nothing to check. I also found earlier that my watchdog graded 42 of 42 real failures as passes, for the same structural reason: it accepted a signal that a well-behaved failure also produces.

Your Sunday check: ten minutes, three questions

Whatever runs your automations — n8n, Make, cron, Claude Code agents, GitHub Actions — pull the run history for the last month and answer these in order.

  1. Which jobs have zero failures AND zero successes? This is the whole point of the exercise, and it is one line of filtering. Any job in this bucket is unmonitorable by definition. For each one, decide now: is it genuinely idle (then cut its schedule), or has nobody checked (then open it and run it by hand once).
  2. How many runs ended with no recorded outcome? Mine was 7.2%. If yours is above zero, your success rate is not what you think it is — it is that number plus an unknown. Fix the reporting before you trust the ratio.
  3. Is your failure count concentrated in a few jobs? If three of twenty jobs own every failure, you do not have a reliability problem spread across your stack. You have a vendor problem in three places, which is a much cheaper thing to fix.

Then one manual run of the quietest job. Not a log read — an actual execution you watch. It is the only test that distinguishes idle from dead. While you are in there, it is worth checking your queue’s age rather than its depth for the same reason: a full queue, like a quiet job, never trips anything.

The takeaway

A low error rate is not evidence that your agents are working. It is evidence that they are not complaining. Those are different claims, and only one of them is worth anything on a Sunday.

The job that needs you this week is not the one lighting up your alerts. The noisy ones are already getting attention — that is what noise is for. It is the one that has been perfectly, politely silent for four months.

Open your run history. Sort by total outcomes reported, ascending. Look at the top row. Ten minutes.

The AI Playbook — Free Download

📥 FREE: THE AI PLAYBOOK

The exact tools and workflows I use to run a one-person agency. 25 years of marketing experience distilled into an actionable guide. Yours free.

Lead Magnet - AI Playbook

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *