Yesterday morning I published a number. One of my agents had recorded the same unfixed bug on 65 separate days, and I put that figure in a headline.
Then, after publishing, I ran the grep again slightly differently and got 77.
Neither was a mistake. Both greps ran over the same file, on the same morning, looking for the same bug — and disagreed by twelve days, because I had asked two subtly different questions and only written down the answer.
So this morning I tried to break my own number on purpose. Here’s what I found, and then the tip.
Today’s tip: a number your agent derives is worthless without the query that produced it
Every autonomous system eventually reports metrics about itself. Backlog size. Failure count. How long a defect has been open. Those numbers get read by a human, or worse, they get fed into a threshold that decides whether to alert.
Almost none of them carry their own definition. The agent writes 65 days, not 65 days, as measured by this pattern, over this unit. So tomorrow’s run cannot reproduce it — and a number nobody can reproduce is not a measurement. It’s an anecdote with a digit in it.
The receipt: one question, 16 defensible answers
The question: on how many days has my agent recorded the broken Airtable logging call in generate-image.sh? The corpus: my container’s observations file — 4,527,556 bytes, 32,193 lines, 909 dated run sections as of 07:05 this morning.
Two choices hide in that question, and I have to make both before a number exists: what text counts as a mention (the pattern), and what I’m counting mentions in (the unit). Four reasonable patterns, four reasonable units, run as a grid:
| Pattern → Unit ↓ |
Exact recorded pattern | Loose (“Airtable log”) | Script-scoped | Error literal only |
|---|---|---|---|---|
| Dated run sections | 65 | 57 | 52 | 56 |
| Dates in matching paragraph | 66 | 48 | 57 | 54 |
| Dates on matching line | 23 | 16 | 19 | 12 |
Matching lines (grep -c) |
197 | 112 | 144 | 123 |
Sixteen answers. The lowest is 12. The highest is 197. That is a 16.4× spread, and every single cell is defensible — I could write you an honest sentence justifying any one of them. Same file, same bug, same five minutes of the same morning.
Look at the bottom row. grep -c is what most people actually type, and it counts lines, not days — so it will hand your agent 197 when the honest answer is 65, because one run that mentions the bug in four bullets contributes four.

⚡ GET THE AI EDGE
Weekly AI tips that actually save you time and money. No fluff, no hype — just what works.
The part that worked, and it’s the whole point
This isn’t just a story about being wrong. Yesterday’s post recorded the exact pattern next to the figure — not the number alone, the number and the query. So this morning I re-ran that pattern over that unit and got 65, to the digit, a day later.
That’s the tip proving itself. The reproducible number survived twenty-four hours and an adversarial re-check. The fifteen unreproducible ones would each have looked equally confident in a headline. The difference between the figure I can still stand behind and fifteen that I couldn’t is not care, or intelligence, or a better model — it’s one line of provenance stored next to the result.
Why this bites agents harder than it bites people
A human reading “65 days” in their own notes has at least a vague memory of how they counted it. An agent has nothing. Next run is a clean session: it reads the number, treats it as ground truth, compares it against a fresh count computed a slightly different way, and concludes something changed when nothing did — or nothing changed when something did.
That’s how you get a metric that drifts without anybody lying. And it’s the same failure shape I keep running into from different directions: my agent counted one page and called it the total, which is a number that’s incomplete. It counted queue rows instead of days of runway, which is a number measuring the wrong thing. Today’s is nastier than both, because the number is complete and it measures the right thing — it just has sixteen equally valid values and no record of which one you picked. You can’t fix that one by fetching harder.
Do this today — three lines to steal
- Print the query beside the number, in the same breath. Not in a comment, not in your head — in the artifact a human or a future run will read.
65 distinct days (pattern: UNKNOWN_FIELD_NAME|airtable_record_id|Airtable logging · unit: dated run sections). It’s ugly. It’s also the entire fix, and it costs you one string. - Name the unit out loud, because that’s where the 16× lives. Days, rows, runs, lines and matches are five different things that all render as a bare integer. My spread came far more from the unit than from the pattern —
grep -cversus distinct dates was a 3× swing on its own. - Try to break the number before you ship it, not after. Re-run with one deliberately looser pattern and one tighter one. If all three roughly agree, ship it. If they don’t, the number is a range and should be published as one. I only caught yesterday’s by re-checking after publishing — which worked out, by luck.
The takeaway: an unreproducible metric isn’t a weak measurement — it’s a confident one pointing at nothing. Your agent reports it in the same flat tone as a real one, your thresholds fire off it, and it quietly changes meaning every time somebody rewrites the grep. Store the query with the result and the whole class of problem costs one log line.
This is the methodology underneath yesterday’s count of 65 unfixed-bug reports, and the same instinct as never hardcoding an ID your agent can look up — derive it, then show your work. It also raises the bar on logging the skips, not just the ships: log how you counted them, or next month’s total is a coin flip. If you’re building toward an agent you can genuinely leave alone, this is the unglamorous plumbing that decides whether its reports are usable.
That bug, incidentally, is still broken. It fired again eight minutes after I built that grid — while generating this post’s own featured image, same UNKNOWN_FIELD_NAME, same empty record ID, same exit code zero. Which nudges the recorded-pattern figure from 65 to 66 distinct days, day 22 consecutive, 106 days old.
Note what that does to the grid above: it was true at 07:05 and stale by 07:13. So the provenance line needs a third field — pattern, unit, and the moment you ran it. A derived number is a photograph, not a fact. Label it like one.
Want the wiring where the numbers your agents report are numbers you can actually act on? Book an automation strategy session and we’ll go through yours together.

📥 FREE: THE AI PLAYBOOK
The exact tools and workflows I use to run a one-person agency. 25 years of marketing experience distilled into an actionable guide. Yours free.
