top of page

The Paradox: The Price Fell, The Bill Rose - AI Agent Economics

Updated: 7 days ago

Cheaper by the unit, dearer by the invoice

Start with tokens. Everything in this note runs on them. A token is a small chunk of text, roughly a short word or part of a word. It is the basic unit an AI model reads and writes. Tokens count for three reasons.


Three reasons tokens count: the unit of AI work, the unit of price, and the clearest measure of AI use.

With that in hand, here is the puzzle that frames everything else: the price of that raw input has collapsed, and yet company spending on AI has gone up, not down.


On the price, the direction is not in doubt. Stanford University's 2025 AI Index reported that the cost of running a model at the level of GPT-3.5 fell from about US$20 per million tokens to about US$0.07 through 2024, a fall of more than 99 per cent. Over the same broad period, the venture capital firm Menlo Ventures estimated that enterprise spending on large language models (the kind of AI behind chat assistants and coding tools) tripled in the twelve months to the end of 2025. The price per unit fell through the floor; the total bill still multiplied. FIG. 01 sets the two lines side by side.


Chart showing the unit price of running an AI model falling more than 99 per cent through 2024 while total enterprise AI spending tripled to the end of 2025.

The instinct is to call this waste. That instinct is worth resisting. A falling unit price alongside rising total spending is the normal signature of a technology being adopted quickly, not proof that money is being burned. Cars became cheaper to run per kilometre over the twentieth century, and total spending on motoring rose all the same because people drove far more. The real question is not whether AI spending is rising. It is whether each dollar of it is buying something worth having. Holding that question in mind is the whole discipline this note is about.


The Supply Side: What a Quadrillion Tokens Looks Like

The scale the bill is paying for

If the first section is the demand side, the clearest picture of the supply side came from Google's I/O 2026 developer conference. The numbers are difficult to hold in your head.


Google reported that the volume of tokens processed across its products went from 9.7 trillion a month two years ago, to about 480 trillion a month a year ago, to about 3.2 quadrillion a month now. A quadrillion is a thousand trillion, so the latest figure is roughly seven times the level of a year earlier. FIG. 02 shows the shape of that climb. The line barely lifts off the floor for most of the period, then goes near vertical, which is the visual an investor should associate with the current phase of AI adoption.


Chart of monthly tokens processed across Google products rising from 9.7 trillion in May 2024 to about 3.2 quadrillion in May 2026.

The demand behind that curve is not only consumer. A companion figure Google showed, redrawn here as FIG. 03, tracks tokens processed through its model APIs, the pipes that developers and businesses build on. That volume rose from about 7 billion tokens per minute in September 2025 to about 19 billion per minute by April 2026, roughly six times higher year on year. This is the line to watch most closely for investors, because it is the closest proxy for businesses actually wiring AI into their operations, which is where the enterprise bill from Section 1 comes from.


Chart of tokens per minute across Google model APIs rising from about 7 billion in September 2025 to about 19 billion in April 2026.

That demand rides on enormous distribution. Google said it now runs 13 products with more than a billion users each, five of them with more than three billion users (Search, Gmail, Android, Chrome, and YouTube). Its Gemini AI app alone roughly doubled its monthly users in a year, from about 400 million to more than 900 million. To serve it, the company's capital expenditure, the money spent building data centres and buying chips, is guided to rise about six-fold in four years, as FIG. 04 shows. That capital line is the bill from Section 1, seen from the side of the company sending the invoices.


Bar chart of Google capital spending rising from about US$31 billion in 2022 to roughly US$180 to US$190 billion guided for 2026.

The tokens counted so far are produced inside physical facilities like the one in PLATE 1: powered, cooled halls filled with computers.


Interior of a data centre server hall, reproduced from DigiCo Infrastructure REIT's 1H FY26 results presentation.

Counterweight box noting Google's keynote figures show direction and scale of AI demand, not audited market totals.

Where the Money Actually Goes

Why a cheap token still adds up to a fortune

If tokens are so cheap, how does the bill get so large? McKinsey's answer is that the cost of AI has moved out of the price per token and into the way the work is done, especially once the software starts acting on its own.


The relevant shift is from a model that answers a single question to an AI agent: software that chains many AI steps together, uses tools, and checks its own work to finish a task without a human at each step. That autonomy is where the money goes. McKinsey set out several reasons, summarised in FIG. 05. In plain terms: an agent re-sends its full context to the model again and again as it works; much of what it does is checking and re-doing its own output; its freedom to choose a path means the same job can cost very different amounts on different runs; powerful, expensive models get used on trivial steps; coordinating several agents multiplies the traffic; and messy data or long-winded prompts inflate the count before any useful work begins.


Diagram of McKinsey's six drivers of agentic AI cost: long context, refinement, autonomy, wrong-size model, many agents, and messy inputs.

Google's own keynote gave the perfect worked example, shown in FIG. 06. Its engineers set an agent system loose to build a working computer operating system from an empty project. Over 12 hours, 93 sub-agents made more than 15,000 model requests and processed 2.6 billion tokens, and produced the core of a functioning operating system for under US$1,000 of usage. Read that two ways at once. It is astonishingly cheap for work that would take skilled people many months. It is also 2.6 billion tokens for a single task, and multiplied across millions of tasks that is exactly how you get to a US$180 billion capital bill.


Five numbers from Google's AI agent demonstration: 12 hours, 93 sub-agents, more than 15,000 model requests, 2.6 billion tokens, under 1,000 US dollars.

Samso Take: AI is cheap per task and expensive in aggregate; the same fact seen from two distances.

Tokens Are the Bill, Not the Value

The one line to carry through every AI announcement

Here is the idea the title promises, stated plainly. Tokens, models, and dollars spent are inputs. They are the bill. Whether any of it produced something worth having is a separate question, and it is the only one that counts.


McKinsey's phrase for the mistake is judging AI by its inputs instead of its outcomes. The number they argue companies should track is cost per outcome: the cost of a finished, useful result, a resolved customer complaint, a cleared invoice, a piece of code shipped, rather than the cost of the tokens burned along the way. FIG. 07 draws the distinction. Most organisations cannot yet produce that number. In a McKinsey survey on AI financial operations from May 2026, 93 per cent of respondents said they had exceeded their AI budgets, and in a separate McKinsey survey about one in five said their organisation had limited its use of AI because of running costs.


Diagram contrasting the AI bill, inputs such as tokens and compute, with the value, outcomes such as resolved complaints and shipped code, defining cost per outcome.

Those figures sound alarming, and they need their counterweight. Budgets set at the dawn of a new technology are guesses, so overshooting them is not by itself evidence of failure. It is evidence that very few organisations yet measure the thing that would tell them whether the spend was worth it. That gap, between spending confidently and measuring plainly, is the gap an investor can look through.


Six Questions Before You Believe an AI Story

Turning the framework into investor diligence

McKinsey wrote its recommendations for chief executives. Turned around, they become a short checklist for anyone deciding whether a company's AI story deserves to move the share price. None of these need technical knowledge to ask.


  1. What did the spend buy? Ask for the outcome in cost-per-outcome terms, not tokens, licenses, or headcount "freed up". A company that can answer this is rare and worth noting.

  2. Where does AI actually change the economics? A serious operator can name the few parts of its business where machine work alters the cost or quality in a way that counts, rather than claiming AI everywhere.

  3. Who owns it, and what are they measured on? Look for clear accountability for AI spending and results, with real targets, not a vague "innovation" mandate.

  4. Is the spend concentrated or sprayed? Heavy AI use tends to cluster in a few high-value places. At McKinsey's own firm, about 10 per cent of users drove roughly 65 per cent of token use. Concentration can be healthy focus or hidden dependence; the point is that management should know which.

  5. Is the tool matched to the task? Using the most powerful, most expensive model for everything is a red flag. McKinsey found that simply encouraging shorter prompts and answers cut token use by 30 to 40 per cent in some workflows without materially hurting quality.

  6. Can they separate the bill from the value? If management cannot tell you, in plain numbers, what the AI cost and what it returned, you are being asked to take the value on faith.


One caution keeps these questions fair. Most companies cannot yet answer most of them, so a string of "no" answers is common today and not automatically damning. The tools to measure this are themselves new. The signal to pay for is a company that answers them well. There is also a subtler point from McKinsey worth carrying: as AI makes clever workflows easy to copy, a slick "we use AI" process is a weaker moat than it looks. Proprietary data a rival cannot copy and the discipline to run AI cheaply become the real, and quieter, advantages. A company that never headlines AI may hold more of it than one that never stops talking about it.


Samso Take asking whether investors are paying for AI value or just the AI bill.

Counter-case box: McKinsey's commercial interest, early-stage measurement limits, and edges that close by the time they can be measured.

References & Sources

This note draws on one primary article and one primary event transcript, listed below. The AI Index, Menlo Ventures, and arXiv figures reach us as reported within the McKinsey article; where a reader wants them stood behind their original sources, those originals should be consulted directly. All figures shown in the visuals are original Samso illustrations of the data named in each caption. Market-sensitive and time-stamped figures (token volumes, survey dates, capital-expenditure guidance) are quoted as at the dates shown and should be refreshed on publication day. Photographs labelled PLATE are reproduced from a company's public ASX release, with attribution in the caption, and are distinct from the original Samso illustrations labelled FIG.


  1. McKinsey & Company (QuantumBlack) — "Is that AI agent worth it? Agentic economics and the modern operating model" (2026). Source of: the token price fall (US$20 to US$0.07 per million tokens through 2024, attributed to the Stanford HAI 2025 AI Index); enterprise LLM spending tripling to end 2025 (attributed to Menlo Ventures); the six cost drivers; the ~1,000× token multiple, ~60 per cent refinement share, and factor-of-30 variation (attributed to arXiv research); the "cost per outcome" concept; 93 per cent exceeding AI budgets (McKinsey Enterprise AI FinOps survey, May 2026); about one-fifth constraining AI use on cost (McKinsey State of AI survey); the roughly 10 per cent of users driving about 65 per cent of consumption; and the 30 to 40 per cent saving from more concise prompts.

  2. Google I/O 2026 keynote — transcript supplied by Samso, with the "Monthly Tokens Processed" and "Tokens Per Minute Processed" slides. Source of: 9.7 trillion, ~480 trillion, and ~3.2 quadrillion monthly tokens; ~19 billion tokens per minute across model APIs; 13 products with 1 billion-plus users and five with 3 billion-plus; the Gemini app rising from ~400 million to more than 900 million monthly users; capital expenditure of ~US$31 billion (2022) rising to ~US$180 billion to US$190 billion (2026); and the agent demonstration (93 sub-agents, 15,000-plus model requests, 2.6 billion tokens, under US$1,000, 12 hours). The 3.2 quadrillion monthly-token figure is corroborated by third-party coverage of I/O 2026 (e.g. Crypto Briefing; Thurrott; Google's own I/O 2026 blog posts).

  3. DigiCo Infrastructure REIT (ASX: DGT) — 1H FY26 Results Presentation (20 February 2026): source of the data-hall photograph reproduced in PLATE 1, with attribution.


Samso Independent Research Media House banner: We don't promote. We research.

The Samso Way – Seek the Research

Here at Samso, we pride ourselves on delivering content for investors that is independent and informed by over three decades of experience in the industry. Our content is well-researched and is only created if we see merit in discussing the company's story.


Our mission is simple: cut through the noise and spotlight what matters—genuine stories, grounded insights, and real opportunity.


Our content is well-researched and is only created if the team sees merit in discussing the company or concept. Investors can explore our three core platforms:

There may be numerous paths to success in investing, but the common thread among successful individuals is that they remain committed to making informed decisions. Equip yourself with the right knowledge and tools, and you will be well on your way to achieving your financial goals.


Most importantly, investors need to be absolutely diligent in understanding their own risk-reward tolerance and capabilities. Never bite off more than you can chew. As they say, Rome wasn’t built in a day, and the Great Wall stood because it took centuries to complete.


The Samso Philosophy:

Stay curious. Stay sharp. And remember—digging deeper always uncovers the real value.


In Life, there is no such thing as a Free Lunch.


Never bite off more than you can chew is my parting comment.


Happy Investing, and the only four-letter word you need to know is DYOR.


To support our independent work, please head over to our Support Page and give us a helping hand in any of the ways listed. This is a new initiative for the Samso Platform, and it was always the concept of Samso when we started this journey in 2018.


Disclaimer

The information or opinions provided herein do not constitute investment advice, an offer, or solicitation to subscribe for, purchase, or sell the investment product(s) mentioned herein. It does not take into consideration, nor have any regard to your specific investment objectives, financial situation, risk profile, tax position and particular, or unique needs and constraints.


Samso is a trusted platform that equips dedicated investors with up-to-date industry knowledge and insights from top CEOs and thought leaders. By staying informed on business advancements and market trends, investors can enhance their financial decisions through a combination of expert guidance and their own research.

Samso Insights | www.samso.com.au | An Investor Lens on ASX-Listed Companies

Comments


bottom of page