The featured image on this post was made by a model I do not own, running on a GPU I do not rent, in a data centre I could not point to on a map. It ran for a few seconds, produced one file, and switched itself off. I paid for the seconds.
That is Replicate, and it is the quietest load-bearing tool in my whole stack. It is also the tool that billed me double what my own documentation said it would this morning, which turns out to be the more useful half of this post.
What Replicate actually is
Think of it as a rental counter for AI models. One API key, thousands of open models behind it — image generators, video models, speech synthesis, transcription, music. You send a POST with your inputs, you get a file URL back, and you are charged for the compute the run consumed. Nothing is installed. Nothing idles.
For a one-person business, that middle option is the whole point. The usual alternatives are a per-seat SaaS that wraps a single model and marks it up, or renting a GPU by the hour and eating the 23 hours you were asleep. Replicate is neither: it is per-run.
My container calls it for four different jobs off one key — images for every blog post and social card, word-level transcription for video captions, original background music, and the occasional one-off experiment. That is four vendors I do not have to hold accounts with.
The receipt from this morning
Here is the full path of the image at the top of this page, start to finish, from a run that happened about twenty minutes before I wrote this sentence:
- One POST to Replicate. One JPEG back.
- Compression: 146,699 bytes down to 40,053 bytes — 72.7% saved. I checked the compressor’s own quota counter first, at 4,798, before spending anything.
- Uploaded to the CDN, then to WordPress as media ID 7047.
Four hops, no human, and the only line item with real money attached to it is the first one.
The part that cost me double
My own project documentation states, in writing, that the default image model is flux-pro at roughly four cents a run, and it names this exact daily post as one of the jobs that uses it. This morning I ran the pipeline with no model flag at all, deliberately, to see what it picked.

⚡ GET THE AI EDGE
Weekly AI tips that actually save you time and money. No fluff, no hype — just what works.
It printed: Step 1/6: Generating image via Replicate (flux-2-pro). That is roughly eight cents, not four.
So I read the script instead of the doc. Line 26 sets flux-2-pro as the default, and the comment next to it carries a dated decision from 2026-05-17: the cheaper model produced weak output for blog use, so the floor was raised on purpose. The script is not broken. The script is correct and four months newer than the document describing it.
That is the trap in per-run pricing, and it has nothing to do with the vendor. A stale line in your own config notes is not a typo — it is a standing order you are still paying. Nobody gets an alert, because nothing failed.
Now scale it. Yesterday’s social run generated seven images in a single pass. At the documented price that pass costs $0.28. At the real price it costs $0.56. Every day, from one skill, and the invoice will never tell you which line of which file did it.
Three lines to steal
- Ask the tool, not the doc. Run the thing with no flags and read the first line of its own output. It will tell you what it actually chose.
- Grep the default, then check its date. One command —
grep -n "MODEL=" your-script.sh— beat my entire written spec. A default with a dated comment explaining itself is worth ten pages of documentation. - Multiply by your loudest job, not your quietest. Per-run pricing is harmless per run. It bills at the volume of your highest-frequency task, so audit that one first.
Why this is a tool recommendation anyway
I am still recommending Replicate, and the doubled cost is the reason rather than the objection. Per-run billing is the only pricing model where a four-month-old decision shows up as a number you can actually find. On a subscription it would have shown up as nothing at all, which is precisely how the true cost of running an agent around the clock hides from the people paying it.
It is also one link in a chain that only works because each hop is cheap and inspectable: Replicate makes the file, Tinify shrinks it and tells you when it is about to quit, Cloudinary serves it in a format social platforms will not reject, and WordPress gets the finished thing. That is four vendors, four receipts, and one pipeline that publishes while I sleep. Yesterday I argued that a red light is a claim and not a result. A line in your own documentation is an instrument too, and this one was quietly wrong in the direction that costs money.
If your stack has a per-run bill you have never traced back to the line of code that causes it, that is the audit worth doing this week. It is most of the unglamorous work behind an agent you can genuinely leave alone. Book an automation strategy session and we will go find out what your pipeline is really charging you.

📥 FREE: THE AI PLAYBOOK
The exact tools and workflows I use to run a one-person agency. 25 years of marketing experience distilled into an actionable guide. Yours free.
