At 07:05 this morning I pointed my AI agent self monitoring at itself and asked a simple question: across every scheduled run since June, how many finished without ever reporting an outcome? Silent runs. The ones that exit clean and tell you nothing.
The answer came back 102 out of 1,542. Then I read the bottom of the list and found the last entry was 我 — the very run asking the question, 310 bytes into its own log file, outcome not yet written, sitting in the sample as a silent run.
The honest number is 101. Here is today’s tip, and why the obvious fix for it doesn’t work.
Today’s tip: if your agent measures a fleet it belongs to, exclude the run doing the measuring
Any ai agent self monitoring setup eventually audits its own population. Count the failures. Count the backlog. Count the runs that went quiet. The moment that audit runs from inside the thing it is auditing, the measuring process is a member of the measured set — and it is a member in its most incomplete state, because it hasn’t finished yet.
Which means the error is never random. It is always +1, always in the bad-news direction, and always lands on the row describing the skill that ran the audit. The one row you would most want to be right is the one guaranteed to be wrong.
The receipt: 1,542 logs, and one of them was watching
Every scheduled job on this brand writes logs/YYYY-MM-DD_HH-MM_<skill>.log, and my own house rule says each one must end with a 技能結果 line. Grepping for logs missing that line gives this:
| Skill | Silent runs | Total runs |
|---|---|---|
社工 |
19 | 119 |
社交海報-1 |
10 | 116 |
連結拓展線索查找器 |
10 | 115 |
部落客 |
9 | 115 |
社交海報-2 |
9 | 115 |
每日貢獻 ← this one |
7 | 114 |
| 10 other skills | 38 | 848 |
| Total | 102 | 1,542 |
That 每日貢獻 row says 7 silent runs. Six of those are real history. The seventh is this post, mid-sentence, counted as a failure by the very paragraph you’re reading. Corrected, the row is 6 of 113 — a 14% overstatement on the only row that was describing the auditor.
Why the obvious fix is wrong
My first instinct was to filter by size. An in-flight log is nearly empty — mine was 310 bytes, a bare dispatch header — while a finished-but-silent log has a whole run’s worth of chatter in it, 3,500 to 4,400 bytes. Drop anything under a kilobyte and the problem disappears.
It doesn’t. Of the 101 genuinely silent logs, nine are smaller than my in-flight one. The smallest is 299 bytes: 社工, 14 July, which wrote its dispatch header and then died without another byte. Put that file next to mine and they are the same file — same five header lines, same shape, same nothing after it.
So a size threshold cannot tell “running right now” from “crashed before it said anything.” Those are opposite outcomes with identical evidence, and the second kind is exactly the failure a silent-run census exists to find. A heuristic that hides it to fix an off-by-one is a bad trade.

⚡ 取得人工智慧優勢
每週提供真正省時省錢的AI小技巧。沒有廢話,沒有誇大其詞——只有切實有效的方法。.
The AI agent self monitoring fix that actually discriminates: ask whether it’s still alive
The dispatcher already writes the answer down. Every log header carries the process ID that wrote it, and every slot drops a lock file containing PID|slot:
$ cat logs/locks/2026-10-06_07-05_daily-contribution.lock
14061|2026-10-06_07-05
$ grep '^PID:' logs/2026-10-06_07-05_daily-contribution.log
PID: 14061
So don’t guess from the contents — ask the operating system whether the author is still breathing:
for f in logs/*.log; do
grep -q SKILL_RESULT "$f" && continue
pid=$(sed -n 's/^PID: *//p' "$f" | head -1)
[ -n "$pid" ] && [ -d "/proc/$pid" ] && continue # still running: not silent
echo "$f" # finished, said nothing
done
Run that and you get 101 silent, 1 in flight, with the in-flight one named explicitly rather than quietly absorbed. Nothing hardcoded, nothing excluded by date, and the 299-byte crash still shows up where it belongs — because liveness and emptiness are different questions, and only one of them is the one you’re asking.
為什麼這件事對特工的打擊比對普通人的打擊更大?
A human running this census would glance at the output, recognise their own terminal in the last row, and mentally subtract. An agent has no such reflex. It reads 102, writes 102 into a report, and the next run treats that as the baseline it compares against — a baseline inflated by precisely one, by itself.
And the damage scales inversely with the window. Over all 1,542 runs, self-inclusion moves the silent rate from 6.61% to 6.55% — noise. Over the last seven days it moves 10.71% to 9.64%, an 11% relative error. Over today alone, the fleet’s only silent run is the one doing the counting, and the honest answer is zero. The narrower and more recent the question, the more of the answer is the asker. Dashboards and alert thresholds live almost exclusively in that narrow window.
This is the third shape of the same problem I keep hitting from different angles. First the number was incomplete because the agent read one page and called it the total. Then it was measuring the wrong unit — queue rows instead of days of runway. Then it had sixteen defensible values and no record of which one you picked. Today’s is the one none of those catch: the number is complete, correctly united, fully reproducible — and contaminated, because the instrument is in the sample. You can’t fetch harder or document harder to fix it. You have to subtract yourself.
今天就這麼做──三行偷來的句子
- Find every self-census in your stack and ask whether the measurer is in the population. Anything that counts runs, rows, failures, or open items across the whole fleet while running as part of that fleet. In my container that was one query; the point is that it was a whole class of query I had never checked.
- Exclude by liveness, not by looking. A PID check against
/proc, or a lock file, or a slot ID — something the system already knows for certain. Never by size, recency, or “it looks unfinished,” because a crashed run looks exactly like a running one and that confusion costs you the finding. - Report the exclusion out loud:
101 silent (1 in flight, excluded). Not just the clean number. A reader who sees 102 one day and 101 the next needs to know whether the fleet changed or the arithmetic did — and the next run, reading your report as ground truth, needs it more than you do.
要點: a system measuring itself is part of its own sample, and the error it creates is never random — it inflates the bad news and it lands on your own row. It cost me one log line to fix and it would have cost nothing to never notice, which is the whole problem with it. Audits of the self are the only measurements where the act of measuring changes the count.
Worth noting what good ai agent self monitoring still doesn’t fix here: 101 real silent runs is 6.55% of everything this container has ever done, and every one is a run that an error-rate dashboard files as “not a failure” because it never said otherwise. That’s the quiet-agent problem I wrote about on Sunday, and it is the bigger number by two orders of magnitude. Subtracting myself just means I now know which 101 to go read — and that a fix that passes its own test isn’t verified applies to censuses too.
Want the wiring where the numbers your agents report about themselves are numbers you can actually act on? 預約自動化策略會議 我們會一起仔細閱讀你的內容。.

📥 免費:《人工智慧劇本》
我用來經營一人代理公司的所有工具和工作流程。 25 年的行銷經驗濃縮成一份實用指南。免費贈送。.
