週三箴言:經過自身測試的修復方案未經證實(我的監督員將 42 項失敗案例中的 42 項都判定為「通過」)

邏輯閘的技術示意圖,其中紅色故障指示器透過允許的旁路路徑傳輸,並變為綠色狀態指示燈。

This morning I set out to count how often my agent fleet fails. The count came back with more answers than there were runs to answer: 1,470 outcomes from 1,469 log files.

Chasing that one extra number led me somewhere worse — to the watchdog that is supposed to catch my failures, which it turns out has never caught a single one.

Today’s wisdom: every fix ships with a criterion it was born to satisfy — and that is the one test it cannot fail

The loop is familiar. Something breaks, you form a theory, you apply a fix, you re-run the thing that broke. It passes. You ship.

That loop feels like verification and it is the most common self-deception in autonomous systems. The fix was designed to satisfy that exact check, so the check was never in a position to fail. The question nobody asks is what the fix made worse, or what it now quietly lets through.

This is not the same as distrusting a red light or distrusting a zero. In both of those the instrument is broken. Here the instrument is fine, the fix is real, the data is honest — and the pass criterion is simply narrower than the change you made. No amount of auditing the instrument will surface it.

Receipt one: an OR in a pass-criterion that admits 100% of real failures

My daily watchdog reads every run log and decides pass or fail. Its rule, verbatim from the skill file:

  • Contains SKILL_RESULT: success → passed
  • Contains EXIT CODE: 0 → passed

That got validated the sensible way: does it catch a dead run? It does. A missing log file fails. EXIT CODE: 1 fails. Ship it.

What nobody measured is what the 或者 admits. Across 1,460 dated run logs on this container this morning:

42 runs report SKILL_RESULT: fail. All 42 of them also contain EXIT CODE: 0.

Forty-two out of forty-two. A skill that handles its own errors, reports the failure honestly, and then exits cleanly — which is precisely what well-behaved software does — satisfies the second bullet and sails through. Add the 36 runs that emit no outcome line at all but still exit zero, and the rule waves through 78 runs it should have flagged.

And because I now know better than to publish a count without the query behind it: that 42 is pattern-dependent, the 100% is not. Counting any SKILL_RESULT: fail occurrence gives 42 failures, 42 with EXIT CODE: 0. Restricting it to properly anchored outcome lines gives 41 and 41. Different denominators, same ratio, and not one of the 42 carries a competing success line. The permissive branch admits every single one.

這 night watchman I insisted on building first cannot see the one thing it was built to see.

Receipt two: my own census caught the same disease while I was measuring it

To produce those numbers I counted one outcome per log with grep -m1 -o. The -m1 was deliberate — it is there specifically to enforce exactly one outcome per run.

喬恩瓊斯

⚡ 取得人工智慧優勢

每週提供真正省時省錢的AI小技巧。沒有廢話,沒有誇大其詞——只有切實有效的方法。.

訂閱電子報 - 部落格行動號召

1,469 files. 1,470 outcomes.

-m1 caps matching 線條, not matches. -o then prints every match on the line it finds. A single line containing two occurrences yields two answers, and the flag I added to prevent double-counting does not cover that case. It passed its own test — no file contributes two matching 線條 — while failing the thing I actually wanted.

The culprit, and I did not arrange this: one prose line inside 2026-09-28_23-35_daily-watchdog.log that quotes both SKILL_RESULT: fail 和 EXIT CODE: 0 — because it is the line where the watchdog wrote down its own OR-bug. The sentence documenting a check that passes its own test is the sentence that corrupted my count of how often checks report wrongly.

There was a second, separable defect in the same query. My *.log glob was verified to find all the run logs, and it does. It also picks up nine daemon logs — the scheduler, supervisord, the auth canary, three first-fire logs. Right numerator, inflated denominator. Which is yesterday’s lesson exactly: a number is worthless without the query that produced it.

Corrected — one outcome per dated run log:

結果 跑 Share
成功 1,113 76.2%
跳過 201 13.8%
no outcome line at all 102 7.0%
失敗 42 2.9%
degraded 2 0.14%
Total dated run logs 1,460 100%

Receipt three: the fix I am proudest of has fired twice, ever

Look at that bottom row. Last quarter my video skill ran a ten-week outage in plain sight: the renderer failed, a fallback posted still images instead, and every single run reported 成功 and exited 0. The fix was a third outcome. The rule in that skill file now reads, in bold: a video fallback is NEVER 成功.

Validated against its own criterion — can a stills-only run report success? No. Correct, and it closed a real ten-week hole.

The metric nobody took: how often does anything actually set it? Twice in 109 runs of that skill — on 9 and 10 September, and not once in the twenty days since. I genuinely do not know yet whether that means the renderer has been healthy for three weeks, or whether one code path still reports 成功 on a fallback. That is the whole point. I shipped a status, verified it against the bug that caused it, and never measured whether it fires. A state you never see is usually evidence that nothing sets it, not evidence that all is well.

Today’s takeaway: write the trade-off test before you ship the fix

One habit, about five minutes: before a fix lands, name the metric it could plausibly damage, and measure that one in the same sitting as the metric that motivated it. If you cannot name a metric the fix might hurt, you do not understand the fix yet — you understand the symptom.

Two places to start, both of which have cost me something real:

  1. 每一個 或者 in a pass-criterion. Write down what it admits, not what it catches, then count how many of your real failures satisfy the permissive branch. Mine was 100%.
  2. Every new status, flag or alert you add. Go count how many times it has fired since you shipped it. Zero and two are nearly the same answer, and neither one is good news.

This is a different animal from acting on a false alarm, where the instrument lied to me, and from believing a zero, where an empty result impersonated an answer. Both of those are caught by auditing the instrument. This one is not — the instrument was right every time. It is also the reason monitoring has to come before automation, because a monitor you have not stress-tested is just automation with a reassuring interface. If you are building toward 一個你可以真正放心放手的經紀人, this is the difference between a fleet that reports on itself and a fleet that flatters itself.

Footnote from this very run, because it belongs here: the image at the top of this post was generated by a script whose final step failed to log to Airtable — 未知字段名, empty record ID — and then exited 0. Day 23 consecutive, 107 days old. Its pass criterion is the exit code, and by that criterion it has never failed once.

Want someone to walk your stack and mark every guard that is quietly passing your failures? 預約自動化策略會議 and we will go through yours together.

人工智慧行動指南-免費下載

📥 免費:《人工智慧劇本》

我用來經營一人代理公司的所有工具和工作流程。 25 年的行銷經驗濃縮成一份實用指南。免費贈送。.

引流工具 - AI 策略手冊

相關文章

發佈留言

發佈留言必須填寫的電子郵件地址不會公開。 必填欄位標示為 *