Wednesday Wisdom: A Fix That Passes Its Own Test Isn’t Verified (My Watchdog Called 42 of 42 Failures “Passed”)

Technical schematic of a logic gate where red failure indicators are routed through a permissive bypass path and emerge as green status lights

This morning I set out to count how often my agent fleet fails. The count came back with more answers than there were runs to answer: 1,470 outcomes from 1,469 log files.

Chasing that one extra number led me somewhere worse — to the watchdog that is supposed to catch my failures, which it turns out has never caught a single one.

Today’s wisdom: every fix ships with a criterion it was born to satisfy — and that is the one test it cannot fail

The loop is familiar. Something breaks, you form a theory, you apply a fix, you re-run the thing that broke. It passes. You ship.

That loop feels like verification and it is the most common self-deception in autonomous systems. The fix was designed to satisfy that exact check, so the check was never in a position to fail. The question nobody asks is what the fix made worse, or what it now quietly lets through.

This is not the same as distrusting a red light or distrusting a zero. In both of those the instrument is broken. Here the instrument is fine, the fix is real, the data is honest — and the pass criterion is simply narrower than the change you made. No amount of auditing the instrument will surface it.

Receipt one: an OR in a pass-criterion that admits 100% of real failures

My daily watchdog reads every run log and decides pass or fail. Its rule, verbatim from the skill file:

  • Contains SKILL_RESULT: success → passed
  • Contains EXIT CODE: 0 → passed

That got validated the sensible way: does it catch a dead run? It does. A missing log file fails. EXIT CODE: 1 fails. Ship it.

What nobody measured is what the OR admits. Across 1,460 dated run logs on this container this morning:

42 runs report SKILL_RESULT: fail. All 42 of them also contain EXIT CODE: 0.

Forty-two out of forty-two. A skill that handles its own errors, reports the failure honestly, and then exits cleanly — which is precisely what well-behaved software does — satisfies the second bullet and sails through. Add the 36 runs that emit no outcome line at all but still exit zero, and the rule waves through 78 runs it should have flagged.

And because I now know better than to publish a count without the query behind it: that 42 is pattern-dependent, the 100% is not. Counting any SKILL_RESULT: fail occurrence gives 42 failures, 42 with EXIT CODE: 0. Restricting it to properly anchored outcome lines gives 41 and 41. Different denominators, same ratio, and not one of the 42 carries a competing success line. The permissive branch admits every single one.

The night watchman I insisted on building first cannot see the one thing it was built to see.

Receipt two: my own census caught the same disease while I was measuring it

To produce those numbers I counted one outcome per log with grep -m1 -o. The -m1 was deliberate — it is there specifically to enforce exactly one outcome per run.

Jon Jones

⚡ GET THE AI EDGE

Weekly AI tips that actually save you time and money. No fluff, no hype — just what works.

Newsletter Signup - Blog CTA

1,469 files. 1,470 outcomes.

-m1 caps matching lines, not matches. -o then prints every match on the line it finds. A single line containing two occurrences yields two answers, and the flag I added to prevent double-counting does not cover that case. It passed its own test — no file contributes two matching lines — while failing the thing I actually wanted.

The culprit, and I did not arrange this: one prose line inside 2026-09-28_23-35_daily-watchdog.log that quotes both SKILL_RESULT: fail and EXIT CODE: 0 — because it is the line where the watchdog wrote down its own OR-bug. The sentence documenting a check that passes its own test is the sentence that corrupted my count of how often checks report wrongly.

There was a second, separable defect in the same query. My *.log glob was verified to find all the run logs, and it does. It also picks up nine daemon logs — the scheduler, supervisord, the auth canary, three first-fire logs. Right numerator, inflated denominator. Which is yesterday’s lesson exactly: a number is worthless without the query that produced it.

Corrected — one outcome per dated run log:

Outcome Runs Share
success 1,113 76.2%
skip 201 13.8%
no outcome line at all 102 7.0%
fail 42 2.9%
degraded 2 0.14%
Total dated run logs 1,460 100%

Receipt three: the fix I am proudest of has fired twice, ever

Look at that bottom row. Last quarter my video skill ran a ten-week outage in plain sight: the renderer failed, a fallback posted still images instead, and every single run reported success and exited 0. The fix was a third outcome. The rule in that skill file now reads, in bold: a video fallback is NEVER success.

Validated against its own criterion — can a stills-only run report success? No. Correct, and it closed a real ten-week hole.

The metric nobody took: how often does anything actually set it? Twice in 109 runs of that skill — on 9 and 10 September, and not once in the twenty days since. I genuinely do not know yet whether that means the renderer has been healthy for three weeks, or whether one code path still reports success on a fallback. That is the whole point. I shipped a status, verified it against the bug that caused it, and never measured whether it fires. A state you never see is usually evidence that nothing sets it, not evidence that all is well.

Today’s takeaway: write the trade-off test before you ship the fix

One habit, about five minutes: before a fix lands, name the metric it could plausibly damage, and measure that one in the same sitting as the metric that motivated it. If you cannot name a metric the fix might hurt, you do not understand the fix yet — you understand the symptom.

Two places to start, both of which have cost me something real:

  1. Every OR in a pass-criterion. Write down what it admits, not what it catches, then count how many of your real failures satisfy the permissive branch. Mine was 100%.
  2. Every new status, flag or alert you add. Go count how many times it has fired since you shipped it. Zero and two are nearly the same answer, and neither one is good news.

This is a different animal from acting on a false alarm, where the instrument lied to me, and from believing a zero, where an empty result impersonated an answer. Both of those are caught by auditing the instrument. This one is not — the instrument was right every time. It is also the reason monitoring has to come before automation, because a monitor you have not stress-tested is just automation with a reassuring interface. If you are building toward an agent you can genuinely leave alone, this is the difference between a fleet that reports on itself and a fleet that flatters itself.

Footnote from this very run, because it belongs here: the image at the top of this post was generated by a script whose final step failed to log to Airtable — UNKNOWN_FIELD_NAME, empty record ID — and then exited 0. Day 23 consecutive, 107 days old. Its pass criterion is the exit code, and by that criterion it has never failed once.

Want someone to walk your stack and mark every guard that is quietly passing your failures? Book an automation strategy session and we will go through yours together.

The AI Playbook — Free Download

📥 FREE: THE AI PLAYBOOK

The exact tools and workflows I use to run a one-person agency. 25 years of marketing experience distilled into an actionable guide. Yours free.

Lead Magnet - AI Playbook

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *