This is the question I get most from people who are almost ready to build one. And they always expect the same shape of answer — a list of tasks that are too hard. Too creative. Too strategic. Too human.
That is not where the wall is. I audited my own fleet this morning to answer it properly: 1,399 logged runs since 12 June. 1,071 successes, 197 clean skips, 2 degraded, and 49 outright failures. Not one of those 49 was the agent being too stupid to do the job. Here is what they actually were.
Start with this morning, because it failed live
At 07:05 the agent that wrote this post went to pull its own search traffic. Google’s token endpoint answered:
{"error": "invalid_grant", "error_description": "Bad Request"}
HTTP 400. The refresh token is present — 103 characters of it. Client ID present. Secret present. Google simply does not accept the token anymore. Search Console and Analytics both ride that single credential, so this morning the agent is blind on both.
Now notice what it could do. It found the failure in one call. It identified the exact cause. It is writing about it in public. What it cannot do is fix it — because the fix is a human being in a browser, on a Google consent screen, clicking Allow.
There is no API for consent. That is the whole answer to today’s question, and everything below is just the shape of it.
AI agent limitations come in exactly three buckets
Across 105 days of running this thing, every genuine “my agent can’t” has been one of these:
- Credentials it can’t re-issue.
- Connections it can’t re-link.
- Consent it can’t give.
Bucket one is almost the entire failure log. Of those 49 failures, 48 name a single vendor — one social scheduler, HTTP 401, token revoked. The 49th was an image-compression quota hitting 429 alongside a CLI returning 401. So all 49 were a credential or quota surface. Zero were architectural. Zero were the agent reasoning badly. And they touched 2 of the 29 skills that have ever run on this container — the two that talk to social platforms.
Bucket two is live right now. I read my scheduler’s connection record this morning: Facebook, Instagram, Threads, X, Pinterest, YouTube and Google Business Profile all connected. TikTok, Bluesky and LinkedIn all null. Three platforms where the agent will compose a perfectly good post and have nowhere to put it — which is exactly the kind of silent failure that looks like success from the outside. Re-linking them means a human logging into TikTok. In a browser. As themselves.
Bucket three is different, and people always mistake it for the same thing. My Drafts-Awaiting-Approval section holds 138 open items. Last night a PayPal receipt for $11 came in; the agent read it, labelled it, filed it and escalated it — and deliberately did not reply. Not because it couldn’t write the reply. Because the rule says it never makes a commitment or a pricing decision on my behalf.
Two of these you should fix. One you shouldn’t.
This is the distinction that actually answers the question.
Buckets one and two are engineering limits. They are accidents of how the plumbing was wired, and you can shrink them. Bucket three is a design limit. It is there on purpose, and removing it is how people end up with an agent that agrees to a discount they never approved.

⚡ GET THE AI EDGE
Weekly AI tips that actually save you time and money. No fluff, no hype — just what works.
Most “AI agents can’t really be trusted” takes are someone looking at bucket three and drawing conclusions about capability. It isn’t a capability gap. It’s a review queue doing its job.
Proof that bucket one is shrinkable
Here is the best receipt I have, and it happened in the same container on the same morning as the failure above.
Gmail didn’t break. Same Google. Same dead token sitting in the same environment. On 8 September I moved this brand’s email off that human-held refresh token and onto a service account with domain-wide delegation. So when the token died overnight, Search Console and Analytics went dark — and email carried on completely untouched. It processed 6 messages last night and 13 the morning before.
One vendor. Two auth models. One survived a credential death; the other didn’t. The difference was never intelligence or model quality. It was whether a human identity sat in the path.
That is the lever. Every credential you move from “a person consented to this once” to “a machine holds this” removes a 3am phone call from your future.
The audit worth doing this week
Write down every credential your automation holds. For each one, ask a single question: if this expires at 3am, who has to be awake?
Anything that answers “a human in a browser” is your real ceiling — not the model, not the prompting, not the box it runs on. Then split the list. The ones that can move to a machine credential, move them. The ones that genuinely can’t — most social platforms, by design — need an expiry alarm and a defined owner, not optimism. That is a different problem from what happens when a tool breaks mid-run: a broken tool is an incident, an expired credential is a standing appointment you haven’t scheduled yet.
One more thing, and it cost me nothing but very nearly cost me accuracy. My own notes have said for days that my voice API key was dead. I checked it properly this morning instead of trusting the note: HTTP 200, active subscription, 163,126 of 300,000 characters used. Alive the whole time. A line in your own documentation is an instrument too, and stale notes about what your agent can’t do will shrink your agent faster than any real limit.
So — what can’t it do?
It can’t prove it’s me.
It can’t stand in front of a consent screen and be a person. It can’t log into a platform as a human being. It can’t accept an obligation in my name. Everything else on the list is downstream of that one sentence, and most of it is a plumbing decision you get to make rather than a law you have to accept.
If you have never traced which of your credentials has a human on the critical path, that is this weekend’s job. It is most of the unglamorous work behind an agent you can genuinely leave alone, and it is usually the difference between a system that runs for 105 days and one that dies quietly on day 9. Book an automation strategy session and we will map yours.

📥 FREE: THE AI PLAYBOOK
The exact tools and workflows I use to run a one-person agency. 25 years of marketing experience distilled into an actionable guide. Yours free.
