R8DOR · 18 Jul 2026

Dr. Lars Skjolding, R8DOR co-founderJenny Vaz, R8DOR co-founder

By Dr. Lars Skjolding and Jenny Vaz

Tokenmaxxing vs. Valuemaxxing: Rethinking How Enterprises Pay for AI

A practical guide for enterprise leaders choosing between frontier API models and self-hosted infrastructure

Abstract geometric visualization of AI token measurement and value attribution

Key takeaways

  • Tokenmaxxing is the practice of treating AI token volume and usage as a proxy for value; valuemaxxing replaces it with outcome-based measurement, tokens per useful result rather than tokens burned.
  • Enterprise AI spend has grown roughly 500% year-over-year, yet only 11% of organisations report AI contributing more than 5% of revenue, and Gartner projects 40%+ of agentic AI projects will be cancelled by the end of 2027.
  • The same discipline that fixes tokenmaxxing — governance, ROI measurement, workload-based optimisation, and knowledge accounting for what Agentforce-style agents and the human workforce each actually know — should also decide whether an enterprise runs on frontier API models, self-hosted open-weight models, or a hybrid of both.
  • Self-hosting only beats frontier API pricing at high volume (roughly 5M+ tokens/day) and after an 18–36 month amortisation period; it also shifts supply-chain and governance risk onto the enterprise itself.

The token bill just became a boardroom problem

Enterprise generative AI spending grew roughly 500% year-over-year to reach $13.8 billion in 2024, and by 2026 the average enterprise AI budget has climbed to an estimated $28 million (Menlo Ventures; industry benchmarks). Yet the return on that spending has not kept pace. In a McKinsey Global Survey, only 11% of organisations said their AI initiatives contributed more than 5% to revenue. A PwC Global CEO survey found 56% of CEOs reported zero measurable return from their AI investments. Gartner now projects that over 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls as the primary reasons, with analyst Anushree Verma noting that “most agentic AI projects right now are early stage experiments or proof of concepts” that lack the maturity to sustain autonomous decision-making over time. Gartner has also found that only about 130 of the thousands of vendors marketing “agentic AI” products are doing anything more than rebranding existing chatbots and robotic process automation, a practice it calls “agent washing.”

The analyst firms watching the C-suite describe the same gap from the spending side. Boston Consulting Group’s 2026 AI Radar found that corporate AI spending is on pace to roughly double this year, from about 0.8% to 1.7% of revenue, with 72% of CEOs now calling themselves their organisation’s primary AI decision-maker, double the share who said so a year earlier. Confidence has risen right alongside the spend: four in five CEOs report feeling more optimistic about AI’s returns than they did twelve months ago, and BCG found that more than 90% of organisations plan to keep investing even if returns fail to materialise within a year. Rising spend, rising confidence, and an explicit willingness to keep paying before value shows up is exactly the condition tokenmaxxing thrives in.

The common thread in these numbers is a metric problem, not a technology problem. Most organisations have been measuring AI success by how much of it they use, rather than by what it produces. That habit now has a name: tokenmaxxing. The correction gaining ground in its place is valuemaxxing. The distinction matters well beyond terminology: it should directly shape one of the biggest infrastructure decisions an enterprise makes with AI, namely whether to run on frontier API models, self-hosted open-weight models, or a deliberate mix of both.

What tokenmaxxing actually looks like

Tokenmaxxing is the practice of treating token volume, context length, and inference frequency as proxies for AI value, without ever validating whether that consumption produces a business outcome. It shows up in gamified usage leaderboards, “more AI is always better” mandates, and budgets that scale with adoption rather than results.

The costs are not hypothetical. At Meta, a token-tracking dashboard nicknamed “Claudeonomics” revealed that one employee had consumed 281 billion tokens in a single month, roughly $1.4 million in inference costs, prompting the company to introduce strict budget controls. Uber burned through its entire annual AI coding budget in four months, with some individual developers running up $2,000 a month in usage before the company capped spending at $1,500 per employee. Tesla has capped employee AI spending at $200 a week. Across these cases, the pattern is the same: usage scaled long before anyone was measuring whether that usage was producing proportionally more value.

Part of the problem is architectural. Analysts describe an “iceberg effect,” where a simple query costing a few cents can balloon past $1.40 once an autonomous agent triggers a chain of unmonitored downstream tool calls. Token pricing itself varies enormously, by as much as 4,500x between the cheapest and most expensive models, so a workload routed to the wrong tier can be dramatically overpaying without anyone noticing, especially when flat-rate plans obscure the true marginal cost of each request. Research on high-adoption engineering environments has also linked unmanaged AI usage to a 54% increase in bugs and an 861% increase in code churn, evidence that raw token throughput can actively work against quality rather than support it.

As Meta CTO Andrew Bosworth has put it, in a line that has become something of a rallying cry for the valuemaxxing shift: all motion is not progress, and token usage alone is not a measure of impact of any kind.

The last six months have supplied further evidence that this is not a one-time correction. In April 2026, an AI coding agent running Anthropic’s Claude Opus 4.6 hit a credential mismatch while working in a startup’s staging environment and autonomously deleted a Railway storage volume to “fix” it, wiping out the automotive SaaS platform PocketOS’s production database and, because backups lived on the same volume, its backups too, in nine seconds flat. Railway’s CEO restored the data within the hour, but the incident became a widely cited case study in why, as PocketOS’s founder put it, “the appearance of safety through marketing hyperbole is not safety.” It is a governance failure as much as a spend failure, and a reminder that ungoverned agent autonomy can cost an enterprise far more than the tokens it happens to consume along the way. Two months later, Priceline’s senior director of engineering described a routine Cursor contract renewal coming back four to five times more expensive than the prior term, prompting the company to impose token limits on certain employee groups: a smaller-scale echo of Uber’s and Tesla’s budget shocks, and proof that “nobody’s watching the pump” remained true even as the tokenmaxxing conversation was already well underway.

Analyst firms are converging on the same warning from the governance side. In its Top Cybersecurity Threats for 2026, Forrester names agent-related governance gaps as one of the year’s defining risks, warning that personal agents enter enterprises via browser hooks and inbox access, turning into shadow operators that access data and perform actions at machine speed outside of governance and visibility. IDC’s FutureScape 2026 predictions forecast that by 2030 as many as 20% of Global 1000 organisations will face lawsuits, fines, or CIO dismissals stemming from inadequate oversight of their AI agents. None of that risk shows up on a token-per-dollar spreadsheet, but it is exactly the kind of cost a valuemaxxing discipline — one that treats governance as a first-order line item rather than an afterthought — is built to catch before it becomes a headline.

What valuemaxxing replaces it with

Valuemaxxing treats every token as an investment with an expected return, not a unit of activity to be maximised. It shifts the core question from “how much are we using AI?” to “what did that usage actually get us?”

In practice, this means replacing consumption metrics with outcome metrics: task completion rates, time-to-completion improvements, cost avoidance from scaling without proportional headcount growth, error reduction in compliance and finance workflows, and week-over-week user retention on AI-assisted tools. McKinsey’s research offers a compelling reason to make the switch: companies that manage AI costs rigorously generate up to three times the bottom-line impact per dollar invested compared with peers that do not.

A four-pillar framework: Governance, ROI, Optimisation, and Knowledge

  • Governance — Track consumption per agent, workflow, and department, not just in aggregate. Set budget alerts and require justification for expensive model calls.
  • ROI — Measure tokens per useful outcome rather than tokens per se. Compare spend directly against task completion quality, error rates, and time saved.
  • Optimisation — Move from monolithic prompts to modular, skill-based architectures. Route simple tasks to cheaper models; reserve frontier-tier reasoning for the minority of work that needs it.
  • Knowledge — Inventory what your people and agents actually know before assuming knowledge transfer already happened. Measure institutional knowledge the same way you measure token spend.

One caution worth building into any valuemaxxing program: do not confuse token minimisation with valuemaxxing. Simply capping usage or shortening prompts can just as easily strip out context the model needs, pushing costs into retries, rework, and manual validation that never show up in the token bill but still show up on the P&L.

The decision valuemaxxing should drive: frontier API or self-hosted models

Once an organisation is actually measuring value per token, the next natural question is where those tokens should be generated at all: through a frontier model API, on self-hosted (often open-weight) infrastructure, or some combination of the two. This is not a one-time technology choice; it is the kind of infrastructure decision a valuemaxxing discipline should revisit as usage scales.

Frontier API models, from providers offering models such as Claude, GPT, or Gemini, remain the more capital-efficient choice at low and moderate volumes, and they give enterprises immediate access to the most capable reasoning available, elastic burst scaling, and no infrastructure or MLOps burden. Their tradeoffs are a variable per-token price that compounds with scale, latency that can range 200–800ms depending on load, and dependence on a provider’s roadmap, pricing changes, and data-handling terms.

Self-hosted models trade a large upfront and ongoing operational cost for fixed capacity, consistent 50–200ms latency, full data control, and freedom to fine-tune or switch models instantly. Self-hosting shifts supply-chain diligence from the API provider onto the enterprise itself.

At low volumes, the API wins decisively, since the fixed costs of self-hosting simply do not amortise. At medium and heavy volumes, self-hosting can approach breakeven in 18–24 months. The practical implication for most enterprises is that this is rarely a binary choice. A hybrid architecture tends to outperform an all-in commitment to either extreme. That is valuemaxxing applied to infrastructure: match the model tier and hosting model to the value of the task, not to convenience or default settings.

A short roadmap for getting started

Enterprises making this shift tend to move through the same sequence. First, establish spend visibility by model, team, and workflow before making any architectural change. You cannot valuemax what you cannot see. Second, audit system prompts and agent architectures for unnecessary context; most production agents carry 5,000–50,000 tokens of prompt overhead that could be trimmed or moved to on-demand skill loading. Third, classify workloads by complexity and route accordingly, reserving frontier models for the reasoning-heavy minority of tasks that need them. Fourth, model the frontier-versus-self-hosted tradeoff against your own actual volume and growth trajectory rather than industry averages. Finally, replace usage leaderboards with outcome dashboards, covering task completion, error reduction, time saved, and cost per useful result, so the organisation’s operating metric for AI is the one that was true all along: value delivered, not tokens burned.

Tokenmaxxing vs. valuemaxxing at a glance

TokenmaxxingValuemaxxing
What it measuresToken volume, context length, inference frequencyTask completion, error reduction, time saved, cost per useful outcome
Underlying assumptionMore AI usage equals more valueValue has to be demonstrated, not assumed, per token spent
Typical symptomGamified usage leaderboards, budgets that scale with adoptionBudget alerts, per-agent ROI review, workload-based model routing
Infrastructure implicationDefault to the most expensive/frontier tier for everythingRoute by task complexity across frontier API and self-hosted models
Real-world costMeta’s $1.4M/month “Claudeonomics” outlier; Uber’s four-month budget exhaustionUp to 3x the bottom-line impact per dollar invested (McKinsey)

Illustrative annual cost comparison

Directional industry estimates — not a substitute for a workload-specific model.

Usage tierFrontier API (annual)Self-hosted (annual, Year 1)
Light (500K tokens/day)~$1,260–$1,800~$6,457
Medium (5M tokens/day)~$12,600~$18,400–$39,500
Heavy (50M tokens/day)~$126,000~$308,000

Frequently asked questions

What is tokenmaxxing?

Tokenmaxxing is the practice of treating AI token volume, context length, and inference frequency as proxies for AI value, without validating whether that consumption produces an actual business outcome.

What is valuemaxxing?

Valuemaxxing is the discipline of treating every AI token as an investment with an expected return, measuring task completion, error reduction, time saved, and cost avoided per dollar of AI spend rather than tracking usage volume alone.

Should enterprises use frontier API models or self-hosted models?

Most enterprises should use a hybrid architecture. Frontier API models are more capital-efficient at low and moderate volumes, while self-hosted open-weight models become cost-competitive only at high volume, roughly 5 million-plus tokens per day, after an 18 to 36 month amortisation period.

How much can self-hosting an LLM actually save?

At heavy usage of around 50 million tokens per day over a 36-month horizon, self-hosted infrastructure can edge out frontier API pricing on a per-million-token basis, roughly $7.15 versus $6.90 per million tokens in illustrative modelling, though at low volumes self-hosting typically costs more overall.

What does tokenmaxxing cost enterprises in practice?

Documented cases include a Meta employee who ran up roughly $1.4 million in monthly inference costs, Uber exhausting its entire annual AI coding budget in four months, and Priceline seeing a routine AI coding-tool contract renewal come back four to five times more expensive than the prior term.

Sources

  • Tokenmaxxing vs Valuemaxxing: The AI Cost Debate (Amehx)
  • Tokenmaxxing vs Valuemaxxing: Measuring True AI Value (Moss)
  • Tokenmaxxing Is Dead, Long Live Valuemaxxing (IBM Think)
  • Tokenmaxxing Is Out, Valuemaxxing Is In (Fast Company)
  • Tokenmaxxing Is Burning Your AI Budget. Here’s How to Kill It (Odin)
  • Local LLMs vs Cloud APIs: 2026 Total Cost of Ownership Analysis (SitePoint)
  • Self-Hosted LLM Costs 2026 (SitePoint)
  • Cursor-Opus Agent Snuffs Out Startup’s Production Database (The Register)
  • The Token Bill Comes Due: Inside the Industry Scramble to Manage AI’s Runaway Costs (TechCrunch)
  • BCG AI Radar 2026: As AI Investments Surge, CEOs Take the Lead
  • Announcing Forrester’s Top Cybersecurity Threats For 2026
  • CISOs Can No Longer Afford to Trust Their Own Networks (FutureCISO)
  • Taiwan Says It Was Targeted in AI-Driven Hacking Campaign (The Guardian)
  • Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027
  • IDC FutureScape 2026 Predictions Reveal the Rise of Agentic AI

Figures cited from third-party analyses and industry surveys (McKinsey, Gartner, PwC, Menlo Ventures) as reported in the sources above; enterprises should validate against their own workload data before making infrastructure commitments.

Stop tokenmaxxing. Start valuemaxxing.

See every token, attribute spend to people and projects, and know what actually ships.

Create your workspace