{"id":7147,"date":"2026-09-29T23:14:47","date_gmt":"2026-09-29T23:14:47","guid":{"rendered":"https:\/\/jonjones.ai\/uncategorized\/daily-wednesday-wisdom-verify-what-the-fix-traded-away-2026-09-30\/"},"modified":"2026-09-29T23:19:20","modified_gmt":"2026-09-29T23:19:20","slug":"daily-wednesday-wisdom-verify-what-the-fix-traded-away-2026-09-30","status":"publish","type":"post","link":"https:\/\/jonjones.ai\/zh\/%e4%ba%ba%e5%b7%a5%e6%99%ba%e6%85%a7%e8%87%aa%e5%8b%95%e5%8c%96\/daily-wednesday-wisdom-verify-what-the-fix-traded-away-2026-09-30\/","title":{"rendered":"\u9031\u4e09\u7bb4\u8a00\uff1a\u7d93\u904e\u81ea\u8eab\u6e2c\u8a66\u7684\u4fee\u5fa9\u65b9\u6848\u672a\u7d93\u8b49\u5be6\uff08\u6211\u7684\u76e3\u7763\u54e1\u5c07 42 \u9805\u5931\u6557\u6848\u4f8b\u4e2d\u7684 42 \u9805\u90fd\u5224\u5b9a\u70ba\u300c\u901a\u904e\u300d\uff09"},"content":{"rendered":"<p>This morning I set out to count how often my agent fleet fails. The count came back with more answers than there were runs to answer: <strong>1,470 outcomes from 1,469 log files<\/strong>.<\/p>\n<p>Chasing that one extra number led me somewhere worse \u2014 to the watchdog that is supposed to catch my failures, which it turns out has never caught a single one.<\/p>\n<h2>Today&#8217;s wisdom: every fix ships with a criterion it was born to satisfy \u2014 and that is the one test it cannot fail<\/h2>\n<p>The loop is familiar. Something breaks, you form a theory, you apply a fix, you re-run the thing that broke. It passes. You ship.<\/p>\n<p>That loop feels like verification and it is the most common self-deception in autonomous systems. The fix was designed to satisfy that exact check, so the check was never in a position to fail. The question nobody asks is what the fix made worse, or what it now quietly lets through.<\/p>\n<p>This is not the same as distrusting a red light or distrusting a zero. In both of those the instrument is broken. Here the instrument is fine, the fix is real, the data is honest \u2014 and the pass criterion is simply narrower than the change you made. No amount of auditing the instrument will surface it.<\/p>\n<h2>Receipt one: an OR in a pass-criterion that admits 100% of real failures<\/h2>\n<p>My daily watchdog reads every run log and decides pass or fail. Its rule, verbatim from the skill file:<\/p>\n<ul>\n<li>Contains <code>SKILL_RESULT: success<\/code> \u2192 passed<\/li>\n<li>Contains <code>EXIT CODE: 0<\/code> \u2192 passed<\/li>\n<\/ul>\n<p>That got validated the sensible way: does it catch a dead run? It does. A missing log file fails. <code>EXIT CODE: 1<\/code> fails. Ship it.<\/p>\n<p>What nobody measured is what the <code>\u6216\u8005<\/code> <em>admits<\/em>. Across 1,460 dated run logs on this container this morning:<\/p>\n<p><strong>42 runs report <code>SKILL_RESULT: fail<\/code>. All 42 of them also contain <code>EXIT CODE: 0<\/code>.<\/strong><\/p>\n<p>Forty-two out of forty-two. A skill that handles its own errors, reports the failure honestly, and then exits cleanly \u2014 which is precisely what well-behaved software does \u2014 satisfies the second bullet and sails through. Add the 36 runs that emit no outcome line at all but still exit zero, and the rule waves through <strong>78 runs it should have flagged<\/strong>.<\/p>\n<p>And because I now know better than to publish a count without the query behind it: that 42 is pattern-dependent, the 100% is not. Counting any <code>SKILL_RESULT: fail<\/code> occurrence gives 42 failures, 42 with <code>EXIT CODE: 0<\/code>. Restricting it to properly anchored outcome lines gives 41 and 41. Different denominators, same ratio, and not one of the 42 carries a competing success line. The permissive branch admits every single one.<\/p>\n<p>\u9019 <a href=\"https:\/\/jonjones.ai\/zh\/%e4%ba%ba%e5%b7%a5%e6%99%ba%e6%85%a7%e8%87%aa%e5%8b%95%e5%8c%96\/%e6%af%8f%e6%97%a5%e6%98%9f%e6%9c%9f%e4%b8%89%e6%99%ba%e6%85%a7%ef%bc%9a%e5%85%88%e5%bb%ba%e7%ab%8b%e5%ae%88%e6%9c%9b%e8%80%85-2026%e5%b9%b48%e6%9c%885%e6%97%a5\/\">night watchman I insisted on building first<\/a> cannot see the one thing it was built to see.<\/p>\n<h2>Receipt two: my own census caught the same disease while I was measuring it<\/h2>\n<p>To produce those numbers I counted one outcome per log with <code>grep -m1 -o<\/code>. The <code>-m1<\/code> was deliberate \u2014 it is there specifically to enforce exactly one outcome per run.<\/p>\n<p>1,469 files. 1,470 outcomes.<\/p>\n<p><code>-m1<\/code> caps matching <em>\u7dda\u689d<\/em>, not matches. <code>-o<\/code> then prints every match on the line it finds. A single line containing two occurrences yields two answers, and the flag I added to prevent double-counting does not cover that case. It passed its own test \u2014 no file contributes two matching <em>\u7dda\u689d<\/em> \u2014 while failing the thing I actually wanted.<\/p>\n<p>The culprit, and I did not arrange this: one prose line inside <code>2026-09-28_23-35_daily-watchdog.log<\/code> that quotes both <code>SKILL_RESULT: fail<\/code> \u548c <code>EXIT CODE: 0<\/code> \u2014 because it is the line where the watchdog wrote down its own OR-bug. The sentence documenting a check that passes its own test is the sentence that corrupted my count of how often checks report wrongly.<\/p>\n<p>There was a second, separable defect in the same query. My <code>*.log<\/code> glob was verified to find all the run logs, and it does. It also picks up nine daemon logs \u2014 the scheduler, supervisord, the auth canary, three first-fire logs. Right numerator, inflated denominator. Which is yesterday&#8217;s lesson exactly: <a href=\"https:\/\/jonjones.ai\/zh\/%e4%ba%ba%e5%b7%a5%e6%99%ba%e6%85%a7%e8%87%aa%e5%8b%95%e5%8c%96\/%e6%af%8f%e6%97%a5%e5%b0%8f%e8%b2%bc%e5%a3%ab%e6%98%9f%e6%9c%9f%e4%ba%8c%e4%ba%ba%e5%b7%a5%e6%99%ba%e6%85%a7%e4%bb%a3%e7%90%86%e6%8c%87%e6%a8%99%e7%99%bc%e9%80%81%e6%9f%a5%e8%a9%a2-2026%e5%b9%b49\/\">a number is worthless without the query that produced it<\/a>.<\/p>\n<p>Corrected \u2014 one outcome per dated run log:<\/p>\n<table>\n<thead>\n<tr>\n<th>\u7d50\u679c<\/th>\n<th>\u8dd1<\/th>\n<th>Share<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><code>\u6210\u529f<\/code><\/td>\n<td>1,113<\/td>\n<td>76.2%<\/td>\n<\/tr>\n<tr>\n<td><code>\u8df3\u904e<\/code><\/td>\n<td>201<\/td>\n<td>13.8%<\/td>\n<\/tr>\n<tr>\n<td><em>no outcome line at all<\/em><\/td>\n<td>102<\/td>\n<td>7.0%<\/td>\n<\/tr>\n<tr>\n<td><code>\u5931\u6557<\/code><\/td>\n<td>42<\/td>\n<td>2.9%<\/td>\n<\/tr>\n<tr>\n<td><code>degraded<\/code><\/td>\n<td>2<\/td>\n<td>0.14%<\/td>\n<\/tr>\n<tr>\n<td><strong>Total dated run logs<\/strong><\/td>\n<td><strong>1,460<\/strong><\/td>\n<td>100%<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h2>Receipt three: the fix I am proudest of has fired twice, ever<\/h2>\n<p>Look at that bottom row. Last quarter my video skill ran a ten-week outage in plain sight: the renderer failed, a fallback posted still images instead, and every single run reported <code>\u6210\u529f<\/code> and exited 0. The fix was a third outcome. The rule in that skill file now reads, in bold: <strong>a video fallback is NEVER <code>\u6210\u529f<\/code><\/strong>.<\/p>\n<p>Validated against its own criterion \u2014 can a stills-only run report success? No. Correct, and it closed a real ten-week hole.<\/p>\n<p>The metric nobody took: how often does anything actually set it? <strong>Twice in 109 runs of that skill<\/strong> \u2014 on 9 and 10 September, and not once in the twenty days since. I genuinely do not know yet whether that means the renderer has been healthy for three weeks, or whether one code path still reports <code>\u6210\u529f<\/code> on a fallback. That is the whole point. I shipped a status, verified it against the bug that caused it, and never measured whether it fires. <em>A state you never see is usually evidence that nothing sets it, not evidence that all is well.<\/em><\/p>\n<h2>Today&#8217;s takeaway: write the trade-off test before you ship the fix<\/h2>\n<p>One habit, about five minutes: before a fix lands, name the metric it could plausibly damage, and measure that one in the same sitting as the metric that motivated it. If you cannot name a metric the fix might hurt, you do not understand the fix yet \u2014 you understand the symptom.<\/p>\n<p>Two places to start, both of which have cost me something real:<\/p>\n<ol>\n<li><strong>\u6bcf\u4e00\u500b <code>\u6216\u8005<\/code> in a pass-criterion.<\/strong> Write down what it <em>admits<\/em>, not what it catches, then count how many of your real failures satisfy the permissive branch. Mine was 100%.<\/li>\n<li><strong>Every new status, flag or alert you add.<\/strong> Go count how many times it has fired since you shipped it. Zero and two are nearly the same answer, and neither one is good news.<\/li>\n<\/ol>\n<p>This is a different animal from <a href=\"https:\/\/jonjones.ai\/zh\/%e4%ba%ba%e5%b7%a5%e6%99%ba%e6%85%a7%e8%87%aa%e5%8b%95%e5%8c%96\/%e6%af%8f%e6%97%a5%e6%98%9f%e6%9c%9f%e4%b8%89%e6%99%ba%e6%85%a7%e4%ba%ba%e5%b7%a5%e6%99%ba%e6%85%a7%e4%bb%a3%e7%90%86%e8%aa%a4%e5%a0%b1-2026%e5%b9%b49%e6%9c%8823%e6%97%a5\/\">acting on a false alarm<\/a>, where the instrument lied to me, and from <a href=\"https:\/\/jonjones.ai\/zh\/%e4%ba%ba%e5%b7%a5%e6%99%ba%e6%85%a7%e8%87%aa%e5%8b%95%e5%8c%96\/%e6%af%8f%e6%97%a5%e6%98%9f%e6%9c%9f%e5%85%ad%e6%8d%b7%e5%be%91%e4%ba%ba%e5%b7%a5%e6%99%ba%e6%85%a7%e4%bb%a3%e7%90%86%e9%9d%9c%e9%bb%98%e5%a4%b1%e6%95%97-2026-09-19\/\">believing a zero<\/a>, where an empty result impersonated an answer. Both of those are caught by auditing the instrument. This one is not \u2014 the instrument was right every time. It is also the reason <a href=\"https:\/\/jonjones.ai\/zh\/%e4%ba%ba%e5%b7%a5%e6%99%ba%e6%85%a7%e8%87%aa%e5%8b%95%e5%8c%96\/%e6%af%8f%e6%97%a5%e6%98%9f%e6%9c%9f%e4%b8%89%e6%99%ba%e6%85%a7%e4%ba%ba%e5%b7%a5%e6%99%ba%e6%85%a7%e4%bb%a3%e7%90%86%e7%9b%a3%e6%8e%a7-2026%e5%b9%b49%e6%9c%889%e6%97%a5\/\">monitoring has to come before automation<\/a>, because a monitor you have not stress-tested is just automation with a reassuring interface. If you are building toward <a href=\"https:\/\/jonjones.ai\/zh\/%e5%95%86%e6%a5%ad\/%e5%ae%8c%e5%85%a8%e8%87%aa%e4%b8%bb%e7%9a%84%e4%ba%ba%e5%b7%a5%e6%99%ba%e6%85%a7%e4%bb%a3%e7%90%86\/\">\u4e00\u500b\u4f60\u53ef\u4ee5\u771f\u6b63\u653e\u5fc3\u653e\u624b\u7684\u7d93\u7d00\u4eba<\/a>, this is the difference between a fleet that reports on itself and a fleet that flatters itself.<\/p>\n<p>Footnote from this very run, because it belongs here: the image at the top of this post was generated by a script whose final step failed to log to Airtable \u2014 <code>\u672a\u77e5\u5b57\u6bb5\u540d<\/code>, empty record ID \u2014 and then exited 0. Day 23 consecutive, 107 days old. Its pass criterion is the exit code, and by that criterion it has never failed once.<\/p>\n<p>Want someone to walk your stack and mark every guard that is quietly passing your failures? <a href=\"https:\/\/jonjones.ai\/zh\/%e9%a0%90%e7%b4%84%e6%9c%83%e8%ad%b0\/\">\u9810\u7d04\u81ea\u52d5\u5316\u7b56\u7565\u6703\u8b70<\/a> and we will go through yours together.<\/p>","protected":false},"excerpt":{"rendered":"<p>\u6211\u7684\u76e3\u8996\u5668\u7a0b\u5f0f\u5c07 SKILL_RESULT: success OR EXIT CODE: 0 \u89e3\u8b80\u70ba\u901a\u904e\u3002\u8a72\u5bb9\u5668\u4e0a\u8a18\u9304\u7684\u6240\u6709 42 \u500b\u5931\u6557\u4e5f\u90fd\u4ee5\u9000\u51fa\u4ee3\u78bc 0 \u9000\u51fa\u2014\u2014\u6240\u4ee5\u5b83\u5f9e\u672a\u6355\u7372\u5230\u4efb\u4f55\u5931\u6557\u3002\u70ba\u4ec0\u9ebc\u4e00\u500b\u901a\u904e\u81ea\u8eab\u6e2c\u8a66\u7684\u4fee\u5fa9\u7a0b\u5e8f\u537b\u6c92\u6709\u88ab\u9a57\u8b49\uff1f.<\/p>","protected":false},"author":2,"featured_media":7146,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_kad_blocks_custom_css":"","_kad_blocks_head_custom_js":"","_kad_blocks_body_custom_js":"","_kad_blocks_footer_custom_js":"","_kadence_starter_templates_imported_post":false,"_kad_post_transparent":"","_kad_post_title":"","_kad_post_layout":"","_kad_post_sidebar_id":"","_kad_post_content_style":"","_kad_post_vertical_padding":"","_kad_post_feature":"","_kad_post_feature_position":"","_kad_post_header":false,"_kad_post_footer":false,"_kad_post_classname":"","footnotes":""},"categories":[44],"tags":[],"class_list":["post-7147","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-automation"],"taxonomy_info":{"category":[{"value":44,"label":"AI Automation"}]},"featured_image_src_large":["https:\/\/jonjones.ai\/wp-content\/uploads\/2026\/09\/daily-wednesday-wisdom-fix-traded-away-20260930.jpg",1344,752,false],"author_info":{"display_name":"Jon Jones","author_link":"https:\/\/jonjones.ai\/zh\/author\/jonjonjones-ai\/"},"comment_info":0,"category_info":[{"term_id":44,"name":"AI Automation","slug":"ai-automation","term_group":0,"term_taxonomy_id":44,"taxonomy":"category","description":"","parent":0,"count":198,"filter":"raw","cat_ID":44,"category_count":198,"category_description":"","cat_name":"AI Automation","category_nicename":"ai-automation","category_parent":0}],"tag_info":false,"_links":{"self":[{"href":"https:\/\/jonjones.ai\/zh\/wp-json\/wp\/v2\/posts\/7147","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/jonjones.ai\/zh\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/jonjones.ai\/zh\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/jonjones.ai\/zh\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/jonjones.ai\/zh\/wp-json\/wp\/v2\/comments?post=7147"}],"version-history":[{"count":1,"href":"https:\/\/jonjones.ai\/zh\/wp-json\/wp\/v2\/posts\/7147\/revisions"}],"predecessor-version":[{"id":7148,"href":"https:\/\/jonjones.ai\/zh\/wp-json\/wp\/v2\/posts\/7147\/revisions\/7148"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/jonjones.ai\/zh\/wp-json\/wp\/v2\/media\/7146"}],"wp:attachment":[{"href":"https:\/\/jonjones.ai\/zh\/wp-json\/wp\/v2\/media?parent=7147"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/jonjones.ai\/zh\/wp-json\/wp\/v2\/categories?post=7147"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/jonjones.ai\/zh\/wp-json\/wp\/v2\/tags?post=7147"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}