AI Token Spend: Why It's an AI Governance Failure, Not a Pricing One

Ask a board what they spend on software and they'll tell you to the pound. Ask them what a token is, or why last month's AI bill jumped forty percent with no change in usage, and the room goes quiet. That gap, between “we've adopted AI” and “we understand what it costs to run”, is where AI adoption quietly stalls, or quietly bleeds money. Usually both.
Every number in what follows is sourced as I go, because a stat with nothing behind it is just a rumour with a decimal point. That's the same standard I'd hold a client to if they brought me a market figure with no source attached.
Where your business actually sits
Every organisation I've worked with sits somewhere on the same four-stage curve with AI spend.
The blind spend. Flat subscription, nobody looks at the meter. This is most SMEs on day one.
The badge spend. Usage becomes a status symbol. Meta built an internal leaderboard for employee AI usage; over one thirty-day period, total usage on it passed 60 trillion tokens, with the top-ranked individual averaging 281 billion tokens a month, a sum that could have cost Meta over $1.4 million at standard frontier-model pricing for one person alone. Uber pushed the same culture, encouraging staff to lean on agentic coding tools and tracking usage competitively.
The frozen spend. Fear shuts the tap. Meta quietly killed the dashboard. Uber burned through its entire annual AI coding budget in about four months and responded with a hard $1,500-a-month cap per employee, per coding tool.
The governed spend. Deliberate, measured, tied to value rather than volume. This is where you want to be, and almost nobody starts here.
Here's the problem with lifting that curve straight into SME governance work: most of my clients never touch the badge spend stage. They go from blind to frozen in a single invoice, with no phase in between where anyone actually learned what the technology could do. A Meta engineer burning through billions of tokens on an internal leader-board is a rounding error against hyperscaler revenue.
A construction firm turning over £30 million, or a legal practice turning over £900,000, discovering an unbudgeted five-figure AI bill is a very different conversation, and it happens at board level, not in an engineering Slack channel. Big tech's lesson isn't wrong. It's just calibrated for a business that can absorb the learning curve as a rounding error. Yours probably can't, which is exactly why the governance has to be built in from day one rather than bolted on after the first bad invoice.
What a token actually is, in language a board will sit still for
A token is a chunk of text, smaller than a word. In English, that's roughly three-quarters of a word per token, so a page of writing runs to about a thousand tokens. Ask the same question in Hindi, Thai or Greek and you can pay two to five times as many tokens for the same content. Code and numbers tokenise unevenly too. None of that is controversial, and it's the first thing I walk clients through before we get anywhere near strategy, because you cannot govern a cost you don't understand mechanically.
The sharper point, and the one most boards have never heard, is that tokens aren't standardised across providers. Every model lab runs its own tokenizer. Identical text can cost 10 to 20 percent more on one platform than another before a single word of output changes. When Anthropic shipped a new tokenizer with its Opus 4.7 model, its own migration guide put the increase at up to 1.35 times more tokens for the same content, and independent testing found the real-world hit ran from roughly 20 to 37 percent more, all at an unchanged headline price per million tokens. That's a price rise wearing a technical explanation, and it's exactly why “price per token” is close to useless as a comparison metric between providers. Test it on your own documents before you take a vendor's sticker price at face value. That's due diligence, not paranoia.
I've written before about the hidden costs of context in AI and it's worth reading alongside this, because context and tokens are two sides of the same invoice. Then there's the layer most people don't know exists: reasoning tokens. Every request carries input tokens, which are cheap, output tokens, which run several times the input price, and reasoning tokens, the model's internal working-out before it answers, billed at output rates and invisible unless you deliberately switch on visibility.
A four-line answer can sit on top of thousands of reasoning tokens underneath it, and dialling up “reasoning effort” can multiply the bill tenfold for the same question. This is the Automation leg of the PAO diagnostic I run with clients, Productivity, Automation, Opportunity, and the point is simple: automation is not a fixed cost once it's switched on. It flexes with how hard you ask the model to think. That needs to be part of the sign-off, not a surprise three weeks later.
Cost per task, not cost per token
Price per million tokens is the wrong number to compare. Cost per finished task is the right one, and almost nobody measures it. Databricks tested this properly on real engineering work from its own multi-million-line codebase: Claude Sonnet 5 was roughly 1.7 times cheaper per token than Claude Opus 4.8, yet it cost $2.09 per completed task against Opus's $1.94, and still scored six points lower on task completion. Cheap tokens and an expensive result. Remember that next time a vendor leads with their headline rate.
The harder question is how a business without an API console or an admin dashboard is meant to measure cost per task in practice, and that's most of the businesses I work with. This is precisely why Competence sits as its own pillar in my 4 Cs of AI Governance framework, alongside Culture, Compliance and Control. Cost-per-task thinking doesn't arrive with the subscription. It's a discipline you build deliberately, with a handful of representative tasks tested properly and someone accountable for reviewing the results.
I've watched this play out for real.
One anonymised case from a South of Scotland Enterprise mentoring engagement involved a creative-industries business that had built a working AI agent handling around twenty different task types. The technology worked, no question. The harder conversation, and the more valuable one, was which of those twenty tasks actually earned their AI token spend and which were running on autopilot because nobody had thought to ask. That question is the difference between AI adoption that pays for itself and AI adoption that just runs.
Build spend, ship spend, bleed spend
Every pound of AI token spend falls into one of three buckets.
Build spend. Experimentation, failed attempts, the context and documentation you feed a model about your own business. This is tuition. Defend it.
Ship spend. The spend that produces actual, used output. Easy to justify, easy to defend.
Bleed spend. Automation nobody's reading. The wrong model doing a simple job. Agents running on a schedule long after anyone checked whether the output still matters. Kill it.
Bleed spend maps almost exactly onto Control, another of my 4 Cs: knowing what's running, why, and whether anyone is still accountable for what it produces. Ask any SME leader whether they could name every automated AI process currently running in their business and watch the pause before they answer.
I'd push the point further. Bleed isn't just a cost problem, it's a governance failure wearing a different coat. On a care-technology engagement I advised on, the AI feature itself wasn't the first thing to fix; the client needed their data flows, access controls, audit trails and DPIA sorted before any AI functionality went anywhere near sensitive children's data. Different domain, same underlying discipline: don't let something run just because it's technically possible, ask first whether anyone is accountable for what comes out the other end. An unread automation and an ungoverned AI feature are the same failure, dressed differently.
Big tech's numbers are not your benchmark
Meta's 60 trillion tokens and Uber's four-month budget blowout are real, well-reported figures, not rumours. Use them to prove the point that unmanaged spend is a genuine risk at any scale, not as a benchmark for your own business. A hyperscaler burning through a leaderboard is a rounding error against its revenue. Your token bill isn't. Your numbers are the only ones that matter, and most businesses I meet haven't actually pulled them yet, which is precisely the gap 360 Strategy provides AI consulting in Scotland to close, alongside the wider cost discipline I set out in AI cost management as the new line on your P&L.
Ready to get control of your AI token spend? Book a Clarity Call.
What I'd actually tell you to do for better AI Governance
Understand the three layers of pricing before you sign anything. Stop comparing headline token prices between providers and start asking what a completed task actually costs. Audit for bleed spend before you touch build spend, because the learning budget is the one that pays off later. Put this inside your AI governance from the outset, not after the invoice that makes your finance director's eyebrows go up.
By the time a token bill reaches board level, it's already a Control failure. It was never really a pricing one.
Mark Evans is founder of 360 Strategy, growth strategy and AI consultancy based in Scotland.
Comments