Rule one in my container’s config is a single line: always read MEMORY.md at the start of every session. Twelve scheduled sessions run in that container every day, and every one of them is told to do exactly that.
This morning MEMORY.md is 1,680,212 bytes. At the usual four-characters-per-token rule of thumb, that is roughly 420,000 tokens. The sidecar notes file beside it, observations.md, is 5,128,705 bytes — about 1.28 million tokens.
The biggest context window you can currently buy on a frontier Claude model is one million tokens. So rule one asks each session to spend 42% of its window before it does a single piece of work, and the file next to it no longer fits in the window at all.
Nothing errored. Nothing will.
First, a correction to my own post
Two weeks ago I ran the 10-minute memory check and told you the file had passed 235,000 tokens, which was “larger than the context window it is supposed to be loaded into.”
That was wrong. 235,000 tokens is larger than a 200K window — that is Haiku 4.5 today. It is nowhere near the 1M window on the Opus and Sonnet tier. I measured honestly, then compared against the wrong ceiling.
The part that is worse than the mistake: that post prescribed a weekly ten-minute prune. Thirteen days later the file has gone from 920 KB to 1,680 KB. It nearly doubled. Zero prunes happened.
I grepped every shell script, Python file and config in the stack for rotation, truncation or pruning logic. There is none. Not one line, anywhere. I published a chore and never built a mechanism, so the chore lost to real work every day for thirteen days.
That changes the fix. Pruning is a habit you have to keep. A read bound is an edit you make once.
Shortcut 1: convert bytes to context windows before you decide anything
“My notes file is 1.6 megabytes” sounds like nothing. You have photos bigger than that. It is the single most misleading sentence in this whole subject, because the unit is wrong.
Convert it to the only unit that matters — how much of the window it eats:
wc -c MEMORY.md | awk '{print $1/4 " tokens, " ($1/4)/1000000*100 "% of a 1M window"}'
Same number. Completely different decision. 1.6 MB reads as fine; 42% of the window reads as broken.
Then price it. Opus-tier input runs $5 per million tokens, so one full read of my memory file costs about $2.10. Twelve sessions a day makes that roughly $25 a day, or $756 a month — to re-read the same notes, most of which are a diary of what happened on a Tuesday in August. Caching does not rescue you: twelve sessions spread across twenty-four hours, and the cache TTL is minutes, not days. Every read is full price.
Shortcut 2: grep your own instructions, not just your files
This is the one nobody runs, and it is where my actual bug was hiding. The file size is the symptom; the instruction is the defect. Find every place in your setup that tells an agent to read something without saying how much of it to read:
grep -rniE "read .*(MEMORY|observations|notes)\.md" .claude/skills/ CLAUDE.md
My results, and I did not enjoy them. Twelve of my skill files reference the memory files. Only four bound the read — they say things like “tail ~30 lines” or “tail ~50 lines”, and those four are fine forever regardless of how fat the file gets.

⚡ GET THE AI EDGE
Weekly AI tips that actually save you time and money. No fluff, no hype — just what works.
The other eight do not. And the instruction that governs all twelve daily sessions — rule one, the one at the top of the config, the one I have never once questioned — has no bound on it at all.
So the only reason this stack still functions is that four skills happened to be written with a tail in them. That is not design. That is luck, and I had been reading it as health.
I hit the same blind spot last week from the other direction: when I audited which model IDs my stack pins, 10 of the 17 files doing the pinning were instruction files, not code. Same cause — your tooling reads your code and ignores your prose. No linter will ever flag a markdown file that tells an agent to swallow a 420,000-token document. Only you will, and only if you look.
Shortcut 3: put the bound in the instruction, not on your calendar
Thirteen days of evidence say a recurring prune does not survive contact with a working week. So stop making the file the thing you manage, and make the read the thing you manage.
Two edits, both one-time:
Bound every read. Change “read observations.md” to “read the last 50 lines of observations.md”. The file can then grow to any size it likes and your startup cost stays flat. My four bounded skills have been immune this whole time without anyone noticing.
Split conclusions from events. Memory holds conclusions. Logs hold events. When they live in one file, the append-only log always buries the durable rules, because events outnumber lessons about a hundred to one. Keep a small rules file you read in full, and an append-only log you read bounded. That is the grown-up version of the file-based memory trick — still the cheapest long-term memory you will ever build, it just needs a shape.
Then measure the slope, because it turns “someday” into a date. My backups are a free time series: MEMORY.md gains about 53,800 bytes a day, roughly 13,451 tokens. Divide that into the headroom and it hits the 1M ceiling around 15 November 2026. Six weeks. A date I can act before.
The decoy: the thing that looks expensive is free
While measuring, I found 251 stale backup copies of those two files in the project directory. 580 megabytes — eighty-nine times the live files they back up. Eighteen different skills independently invented the same .bak-<skill>-<date> convention, which appears in no config and no skill file anywhere. Pure emergent habit, copied agent to agent.
It looks like the problem. It is not. That disk is 27% full of 301 GB — those 580 MB cost nothing, slow nothing, break nothing. The 1.6 MB text file is the one that breaks the agent.
You will instinctively clean up the big number, because big numbers feel like problems. Meanwhile the small number sits at 42% of your window, growing, with no alarm attached. A dead API key throws a 401. An oversized memory file throws nothing at all — it just quietly makes every session dumber and more expensive than the last.
Run this today
Three commands, ten minutes:
- Convert.
wc -cevery file your agent reads at startup, divide by four, express it as a percentage of your model’s window. Decide on that number, not the megabytes. - Grep. Search your skills and config for read instructions with no line bound. If fewer than all of them are bounded, you found it.
- Bound. Add
tail -n 50to the instruction. It cannot lapse, cannot be skipped on a busy Tuesday, and does not care how big the file gets next year.
The takeaway: your agent’s memory file will never tell you it has grown too big. Nothing in the stack is watching it, and the advice to prune it weekly is advice I personally ignored for thirteen straight days. Do not manage the file. Bound the read.
Same failure shape as an agent that logs the same bug sixty-five times without learning from it — writing is automatic, reading is not, and nothing in between ever subtracts. It is also why giving your agent a memory is still the best text file you will ever create, and still the one that needs a shape. Last Saturday was about what a failed run already did; this week is about the next run still thinking straight when it starts.
Want a fleet whose memory gets sharper instead of just heavier? That is the unglamorous plumbing behind agents you can actually leave alone. Book an automation strategy session and I will show you exactly where mine is bounded, and where it still is not.

📥 FREE: THE AI PLAYBOOK
The exact tools and workflows I use to run a one-person agency. 25 years of marketing experience distilled into an actionable guide. Yours free.
