The token, the invisible billing unit of AI, escapes control: the drop in unit price masks the explosion in volumes, and the ROI remains uncertain…
L’artificial intelligence generative is taking hold in organizations at an unprecedented speed. However, one question remains largely absent from management committees: how much does AI through tokens really cost? This billing unit, invisible to the end user, has become the nerve of the economic war of AI systems. And it still escapes most control systems.
The illusion of zero marginal cost
The dominant discourse has long maintained confusion: since the unit cost of the token continues to fall, AI would be more and more affordable. The reasoning is misleading. What falls is the price of a token; what explodes is the number of tokens consumed. So-called “reasoning” models generate internal thought chains that multiply consumption by ten, twenty, sometimes more, for the same user request. Agentic architectures, which chain tool calls and iterations, further amplify the phenomenon. The cost per unit decreases, the overall bill increases. This is the Jevons paradox applied to calculation.
Added to this are peripheral costs that are rarely consolidated: document ingestion and vector indexing, storage of embeddings, human supervision of outputs, GDPR and IA Act compliance, contractual reversibility. An organization that only budgets for an API subscription structurally underestimates its actual expense. In higher education as in industry, I have seen projects whose twelve-month operating cost exceeded the initial estimate by three to five times, without anyone making a calculation error: simply, the measured scope was the wrong one.
An ROI that remains to be demonstrated
The MIT Media Lab report published in 2025, The GenAI Divide: State of AI in Business, made an impact by establishing that approximately 95% of generative AI pilot projects in business produced no measurable return on the P&L. The figure has been discussed, its methodology contested. There remains nonetheless a signal: the gap is widening between organizations that experiment and those that industrialize.
The difficulty lies in the very nature of the gains. The time saved by an employee only becomes an economic value if it is reallocated to a productive activity. Without redefining processes, the gain dissipates. This is why the strongest ROIs are observed where AI has been backed by an organizational overhaul, and not tacked on to what already exists.
A changing business model
The economic model of LLM providers is being rebuilt before our eyes. Three converging movements deserve management attention.
First, pricing segmentation: prompt caching, batch processing, distilled models, differentiated pricing based on latency. The same task can see its cost vary by a factor of five to ten depending on the architecture chosen.
Then, the shift towards billing for business use rather than for tokens. Several publishers are experimenting with pricing indexed to the task accomplished or the result obtained. This shift transfers the technical risk to the supplier, but makes budgetary anticipation more opaque.
Finally, commonality through open source. The availability of efficient open models exerts continuous deflationary pressure and opens up the option of sovereign execution, on controlled infrastructure. Arbitration is no longer binary: it becomes a portfolio of models to be orchestrated according to the criticality and sensitivity of uses.
To lead, or to suffer
These developments call for a governance response, not just a technical one. Three principles seem structuring to me.
Establish observability of the token. Account for consumption by use case, by direction, by user. What we don’t measure, we can’t control. FinOps systems applied to the cloud offer a directly transposable framework.
Define thresholds and guardrails. Capped budget per project, deviation alerts, quarterly review. A use case whose unit cost exceeds the value produced must be able to be stopped without drama.
Route intelligently. Not every query deserves the most powerful model. A routing policy that reserves advanced models for complex tasks and entrusts the volume to lightweight models commonly divides the bill by three.
Generative AI will be no exception to the rule that applies to all technology: its value lies not in its adoption, but in its management. Organizations that treat the token as a strategic, measured, arbitrated, optimized resource will gain a lasting lead over those that continue to consider it as an implementation detail.
References
- MIT Media Lab, The GenAI Divide: State of AI in Business 2025, NANDA Project, July 2025.
- McKinsey & Company, The State of AI: How organizations are rewiring to capture value, March 2025.