I had a run of projects to finish. The chat interface was getting in my way, so I moved onto Anthropic's API and wired the models straight into my workflow: faster, cleaner, and without the copy-pasting.
Fast forward a week: I'd spent £3,000.
It was not an ‘accident’, like a runaway loop, a bug or a leaked key. In fact, I got exactly what I paid for and completed said projects. But let’s just say I got a little carried away and thought of the associated AI costs more in terms of electricity rather than tokens used.
Here's what I learned, including the part where I was wrong.
First, the Number
£3,000 is roughly $3,900. Claude Opus 5 is $5 per million input tokens and $25 per million output. Mind, for ease I’ll be using USD throughout this article.
Agentic work is input-heavy, as every tool call resends the accumulated context. A realistic blend puts that week somewhere in the region of 400 million tokens. A token is about three-quarters of a word. Call it 300 million words, in seven days. The complete works of Shakespeare, several hundred times over.
For a sense of the gap between the brochure and the reality: Anthropic's own pricing documentation gives an illustrative one-hour session at 50,000 input tokens and 15,000 output — about seventy cents. I'd got through the equivalent of five and a half thousand of them.
Real agentic work looks nothing like what Anthropic pictures above. Instead, imagine a machine reading everything relevant, “thinking”, trying something, reading it all again, and repeating — several thousand times a day, without lunch breaks.
My Misconception
I went into this convinced that AI tokens are sold below cost with venture money covering the difference, expecting the real bill to land in 2028.
That's half right, but the wrong half is what matters.
On paid API tokens, margins are positive and improving. SemiAnalysis reporting puts Anthropic's gross margin at roughly -94% in 2024, around 40% in 2025, and near 60% in 2026 — driven mainly by inference efficiency rather than price increases. OpenAI's adjusted gross margin went the other way across 2024–25, from about 40% to 33%, largely because of the enormous free tier it carries.
So, where is the subsidy? Three places.
1. The free tier. OpenAI serves something like 900 million weekly users of whom only about 7% pay anything. Nobody outside the company knows what that free tier costs to run, but it is the largest cross-subsidy in software — and it subsidises consumers, not your API bill.
2. Flat-rate subscriptions. A £20 or £100 monthly plan absorbs variance: heavy weeks and quiet weeks average out, and the provider eats the difference. That’s where I got caught out, as the API averages nothing and all of it counts. My usage barely changed when I switched, but my visibility changed completely.
4. Gross margin excludes training. Your tokens pay to serve the model, yet don't pay to build the next one. That remains investor capital, and it is the part of the equation that has never once closed on its own.
Worth knowing there's a dissenting view: Ed Zitron argues that prepaid enterprise token commitments flatter these revenue and margin numbers, because the cash arrives before the compute is delivered. I don't think it changes the direction of travel, but it’s definitely a fair argument.
The Real Culprit: the Volume
Per-token prices have been falling roughly tenfold a year — a16z christened it "LLMflation." GPT-4-class output went from about $30 per million tokens in 2023 to under $0.50 by 2026. By every measure of unit cost, AI got dramatically cheaper while I was using it.
My bill went up anyway, because the shape of the work changed underneath me:
Ramp, which sees the actual card and invoice data, found median monthly business AI spend grew 4x between February 2025 and February 2026.
Agentic workflows consume 5 to 30 times more tokens per task than a chatbot exchange. Research from Microsoft and Stanford's Digital Economy Lab puts the extreme end nearer 1,000x.
EY costed a single customer-service interaction at about $0.04 in 2023 and roughly $1.20 in 2026 — a thirtyfold rise driven entirely by orchestration.
Ever heard of Jevons paradox? It’s when a resource gets cheaper to use, so we find far more uses for it — and end up consuming more of it, not less. Tokens are the resource here.
So, while the tokens got cheaper, the workflows got greedier. Multi-step agentic architectures were always buildable – just not at 2023 prices. At 2026 prices they are, and they consume tokens at a rate no chatbot exchange ever did.
Something sneaky has also happened: Anthropic's docs note that Claude 4.7 and later use a new tokeniser producing roughly 30% more tokens for the same text. Need I say more?
I'm in decent company. Uber's CTO reported burning the company's entire 2026 AI coding budget in four months. And in June, Sam Altman told CNBC that the ROI question is the fairest criticism of AI right now.
When the vendor concedes the point, the point is conceded.
The Next 12–18 Months: From Cost-per-Seat to Cost-per-Action
Here's my hypothesis, feel free to challenge me.
Every agentic product available now— Cowork, managed agents, background workers, scheduled tasks — moves AI from something a person uses to something that runs. That's a category change in how it costs money, and that’ll come as a surprise for finance departments.
Look at how the metering is already evolving. Anthropic's Managed Agents bill $0.08 per session-hour on top of tokens. Code execution gets 1,550 free container-hours a month, then five cents an hour. Web search is a penny a query. These are small numbers designed for a world where agents run continuously — but small numbers in large volumes – as the cloud has already shown us – add up.
Three things I expect:
A sky-high surprise bill caused by someone in operations scheduling an agent against a mailbox or a data feed which runs every fifteen minutes for a quarter. No-one accountable for the spend – as a threshold was never set, and no-one was watching the meter go up. Cloud had this moment around 2015 with AI's coming very soon.
Per-seat pricing will stop. It doesn’t make business sense anymore to offer a £30/seat licence for a product where one enthusiastic user generates £3,000 of compute in a week. Expect credit systems, hard caps, throttling, and some difficult renewal conversations. It's already starting.
FinOps for AI becomes a real job. We’re talking token budgets, routing policies, per-agent cost attribution, and step limits. I’m looking forward to seeing an uplift in the job market where the people who built cloud cost governance between 2018 and 2022 are extremely employable again.
Goldman Sachs projects a 24-fold increase in token consumption by 2030, while semiconductor cost-per-token falls 60–70% a year. Read those together: unit prices collapse, volumes explode faster, total spend rises = don’t expect AI costs to fall.
Have a Think About This
I don’t mean to argue against AI in this article. I got the projects done, and for that specific work, £3,000 was good value.
It's an argument against reflexive AI. So, here's what I now ask before anything touches a model:
Does this task require actual thinking?
We have collectively started routing everything through the most expensive reasoning engines ever built — Fable 5, Opus 5, GPT-5.6 — and a startling proportion of what we send them isn't reasoning at all. It's deterministic, specifiable (and boring) logic.
Three tests:
Can you state the rule? If you can specify in advance what your output should look like, you don't need something that follows the rule rather than reasons. A fixed rule costs almost nothing, runs instantly, and is right every single time.
Does it need judgement, or a specification? Ambiguity, nuance, synthesis across messy sources, edge cases requiring interpretation — it’s model work worth paying for. If we’re talking about simple things like reformatting dates, validating postcodes, pulling a known field out of a standard form, or adding up a column, we’re talking about ordinary automation. Sending it to a maths-olympiad-grade reasoner is hiring a barrister to alphabetise your filing cabinet: extraordinarily capable, wildly overqualified, billing by the hour.
Does the same input need the same output every time? This is extremely important. Models are non-deterministic. Ask the same question twice, get two defensible and slightly different answers. While this is considered a feature in a creative brief, in a compliance pipeline processing ten thousand records it's an unreconciled variance nobody can explain. Unfortunately, it’s often only discovered during an audit.
When a task genuinely does need a model, it very often doesn't need the biggest one. Within a single vendor the spread is roughly tenfold: Haiku 4.5 at $1/$5 per million tokens against Fable 5 at $10/$50. Anthropic's own documentation costs processing ten thousand support tickets on Haiku at about $37. Routing routine work down the ladder and escalating only the failures is the single largest cost lever available — and it's an architecture decision.
My Conclusion
My £3,000 week wasn't a warning that AI is too expensive. Tokens are cheap and getting cheaper.
However, it was a warning about inattention. About the gap between a flat monthly fee and a live meter. And about how much of what we hand to frontier models is work that a simple, fixed set of rules would do faster, cheaper, and identically every single time.
The question isn't whether AI is right for your business.
It's whether this task, right now, is right for AI (and which provider/model at that).
I'd like to hear from people who've had the same wake-up call — or who think I've called it wrong.
---