Monday Myth-Busting: “AI Agents Do What You Tell Them” (Mine Followed a Rule That Never Existed — 15 Mondays Running)

An open manual on a desk with one paragraph glowing, in front of a wall of status lights all showing green — documentation and system state disagreeing silently.

Every week someone tells me their automation went sideways because “the AI didn’t do what I told it.” I have run 23 scheduled agent jobs on this brand since June — 1,536 logged runs — and I almost never see that failure. My agents are relentlessly obedient.

That is the problem.

The myth is not that agents follow instructions. They do. The myth is the sentence people tack onto the end: “…so if it’s written in the config, it’s happening in the system.” That second half does not follow from the first, and this morning I proved it on myself with a line that had been wrong in my own file for 111 straight days.

The rule my agent never followed, because nobody ever wrote it

The config file every one of my agents reads at the start of every session contains this:

On Mondays, social-miner is overridden by link-outreach-query-researcher at 12:00 (weekly override). Runs 6/7 days.

Specific. Operational. The kind of line you skim past because it reads like something a careful person wrote. I have read it dozens of times and never questioned it once.

It has never been true. Not for a single day. Here is the dispatch record, pulled an hour ago:

  • social-miner has run 118 times across 111 consecutive days. Every day since 16 June. 111 of 111.
  • It ran on all 15 Mondays in that window. Monday isn’t even its lightest day — 16 runs, against 16 to 18 for every other weekday. Statistically invisible.
  • The job documented as overriding it runs at 10:28, not 12:00. It has fired 15 times, all Mondays, and has never collided with anything.

So the line is wrong twice: wrong about the time, and wrong about the outcome. “Runs 6/7 days” is actually 7/7. And in 111 days, not one run, not one log line and not one alert ever contradicted it.

Why no alarm could ever have fired on this

I went and read my own scheduler instead of assuming. It resolves an override by building a key out of the day and the time slot:

weekly_key = f"{day_name}_{slot}"

An override only fires when the weekly job sits in the exact same slot as the daily one. social-miner is at 12:07. The job documented as overriding it is at 10:37. Different slots.

The override isn’t broken. It was never wired. Someone described an intention in prose, nobody implemented it, and prose doesn’t execute.

That generalises well past my stack. Your agent can be wrong in two completely different ways. It can disobey an instruction — loud, testable, your logs will tell you. Or your instructions can describe a system that doesn’t exist — silent, untestable, permanent. Every stack I’ve seen tests whether code does what code says. None test whether the documentation is telling the truth. There is no linter for a false claim.

So I audited the rest of the file: 14 claims, 9 of them wrong

Once I found one, I checked every factual claim in that file I could verify in one sitting. Fourteen of them. Nine did not survive contact with the live system:

Jon Jones

⚡ GET THE AI EDGE

Weekly AI tips that actually save you time and money. No fluff, no hype — just what works.

Newsletter Signup - Blog CTA
  • The Monday override above — false, 15 Mondays running.
  • Two integrations documented as “MCP server connected” — neither is in the connection config at all. That file holds exactly one server, and it returns a 401.
  • My task board is documented as having 9 sections. It has 10. The undocumented one is the default section every stray task silently lands in — which is to say, the one nothing I own has ever read.
  • Three credentials listed as available are unset or empty. One of them powers a music pipeline that the same file, 160 lines earlier, says was deleted five months ago.
  • The authentication expiry is a literal TODO in one file and a hard date in another.
  • One agent is documented as running Thursdays and Sundays. It runs Thursdays and Saturdays.

Two more measurements, because they say something the list doesn’t. My config still contains 14 unrendered template tokens — the setup wizard’s own placeholder syntax, sitting in the live source of truth four months later. And 23 of 23 skill files declare a schedule that disagrees with the actual scheduler.

Across all of it: zero errors. Not one of these contributed a single failure to 1,536 runs.

Three tiers of wrong, and only one of them matters

This is where config audits go bad — you find twelve discrepancies, declare the system rotten, and rewrite things that were fine. Sort them instead:

Tier 1 — cosmetic. Twenty of my twenty-three schedule mismatches are a flat seven-minute stagger. The template tokens. The TODO markers. Untidy, not wrong. Fix them when you’re already in the file; otherwise leave them.

Tier 2 — stale enough to mislead. A credential listed as live for a pipeline deleted in May. Three env vars that don’t exist. An expiry date that’s a TODO in one place and a date in another. These cost nothing today and cost you an hour on the day you debug the wrong thing.

Tier 3 — actively false descriptions of behaviour. The Monday override. The agent that runs Saturday while the file says Sunday. This is the only tier that changes a decision. I was reasoning, budgeting and scheduling against a fleet where one job stood down weekly. That fleet does not exist. Everything I concluded downstream of that line inherited the error.

The 15-minute audit

Whatever runs your automations — n8n, Make, cron, Claude Code agents, GitHub Actions — three questions, and the first is the whole exercise.

  1. Grep your config for “override”, “unless”, “except” and “instead of”. Every hit is a conditional somebody wrote in prose. For each one, go find the person or the line of code that implemented it. In my file, that single grep found the defect. These words are where intentions get recorded and never built.
  2. Separate settings from descriptions, then only check the descriptions. A setting (“the API key is X”) is usually right, because something reads it and breaks loudly when it’s wrong. A description (“job A overrides job B on Mondays”) is usually unverified, because nothing reads it but you. Pick your three load-bearing descriptions and check each against the actual run record.
  3. Count what your config claims is connected, then open the connection file and count what’s actually there. Mine claimed three integrations. The file holds one. That took forty seconds.

If you want a deeper version of this, the same blind spot shows up in what your agents are instructed to read at startup, and in the jobs that never report anything at all. Different symptoms, one cause: our tooling reads code and ignores prose.

The takeaway

AI agents do what you tell them. That is precisely what makes this dangerous — obedience makes a false claim look verified. Nothing disobeyed. Nothing broke. Nothing even got slower. A sentence was wrong, every system downstream behaved perfectly, and four months of green lights quietly certified it.

This is the same family as my agent logging the same bug 65 times without fixing it and my watchdog passing a fix that never worked: writing is automatic, checking is not, and nothing in an automated stack ever subtracts a claim. It is also why an agent is not a cheaper VA — a person would have mentioned, somewhere around week three, that the Monday thing wasn’t happening.

So stop asking whether your agent is following instructions. It is. Ask when anybody last checked that your instructions were true.

Open your config. Search it for the word “override”. Fifteen minutes.

And if your stack has outgrown the point where you can hold its real behaviour in your head — that’s the actual problem, and it’s fixable. Book an automation strategy session and I’ll walk your config against your run record with you.

The AI Playbook — Free Download

📥 FREE: THE AI PLAYBOOK

The exact tools and workflows I use to run a one-person agency. 25 years of marketing experience distilled into an actionable guide. Yours free.

Lead Magnet - AI Playbook

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *