Companies are slashing AI spending, triggering a rapid collapse in token prices that is squeezing margins across the generative‑AI stack and forcing model providers into their first real pricing reckoning — with major implications for investors watching the sector’s profitability evaporate.

The most powerful and expensive AI models aren’t necessary for relatively mundane tasks. “It’s like driving a Lamborghini to go to the grocery store to pick up milk when that was designed to be raced around a track,” Cursor field chief technology officer Mike Saeks told The Wall Street Journal .

In response, companies are using lower-priced models, including some built in China. The potential cost savings — for example, saving 87% on the cost of building a web browser — are considerable.

Doing that from scratch costs more than $10,000 on OpenAI’s GPT-5.5; whereas a combination of Cursor’s Composer and Anthropic’s Opus 4.8 gets the job done for $1,339, according to Cursor research featured by the Journal.

Businesses that use such chatbots are getting big discount offers. Pylon, a customer support platform for high voltage towers, has received around $1.6 million in free tokens from one vendor, $65,000 from another and $10,000 from a third, co-founder Marty Kausas told the Journal.

Why Tokenomics Suddenly Matters

The fierce price competition is bringing to life a new field of study for decision-makers: tokenomics. This discipline helps businesses make the best use of limited AI budgets by tracking the rapidly changing price of tokens — chunks of data on which AI chatbot prices are set — as well as the cost of providing those tokens to users.

By applying tokenomics, AI users are inadvertently creating winners and losers across the generative AI value network.

The winners include Chinese suppliers of less expensive large language models — including DeepSeek, Qwen, GLM, Kimi and MiniMax — which by mid-2026 won 46% market share . Other beneficiaries are GPU and memory chip designers and fabricators such as Nvidia, TSMC, Micron, SanDisk and Western Digital.

Meanwhile, more expensive LLM providers — such as Cursor, Uber and Salesforce — are adjusting their prices downward.

The Forces Cutting Into AI Profit Potential

My work with Harvard Business School professor Michael Porter — whose books in industry analysis and competitive advantage transformed the field of business strategy — is making me acutely sensitive to the forces driving down the AI industry’s profit potential.

I previously mapped out this industry — identifying segments that vary in their inherent profitability. In a nutshell, the AI chip design and manufacturing segment is highly profitable while the AI chatbot suppliers lose huge amount of money.

How so? OpenAI and Anthropic face attacks from all sides. Suppliers have huge bargaining power; price competition is fierce as is pressure to innovate; the cost for new entrants to participate is dropping; and customers face low switching costs and a strong incentive to find the best value.

  • Bargaining power of suppliers. AI chatbot companies buy tokens from AI cloud services providers, which buy GPUs and memory chips from chip makers that license designs from the likes of Nvidia. Although the cost of tokens is dropping about 10-fold per year — falling from $60 per million tokens in 2021 to roughly $0.06, according to EpochAI — agentic AI is using far more tokens. Google’s token consumption has soared 330-fold from 9.7 trillion tokens a month in 2023 to more than 3.2 quadrillion today, Google CEO Sundar Pichai said at Google I/O in May . Moreover, demand for the chips that do all this work is greater than the supply — with prices of direct random access memory chips expected to rise 335% in 2026 .
  • Rivalry among existing competitors. As noted above, OpenAI and Anthropic face competition from so-called open weight models, which can do the work companies need for far less money. While companies do not pay a fee to license them, they must pay to operate them, per layer3labs . However, the availability of these lower-cost models puts pressure on the incumbents.
  • Threat of new entrants. While OpenAI and Anthropic incurred enormous expenses to build their models — $100 million for Claude 3 — rivals such as DeepSeek paid much less: $6 million to train (excluding the $500 million spent on the hardware and other upfront costs). These lower barriers to entry make continued price competition inevitable.
  • Bargaining power of buyers. Finally, companies are fed up with spending so much money on AI — especially in light of the difficult-to-measure economic payoff from the investment.

With IPOs looming for OpenAI and Anthropic, the diminishing profit potential of the AI chatbot market does not bode well for retail investors.

Profit Potential Across The AI Stack

The chip and cloud services providers enjoy the highest profit potential while the model providers are being squeezed the most.

Model providers serving consumers suffer lower margins than business-focused companies. OpenAI, at roughly $25 billion of annualized revenue, is estimated to run about a 33% gross margin — dragged down by consumer ChatGPT, which charges a monthly rate regardless of how much is consumed.

Anthropic — which gets about 80% of its $30 billion run-rate from enterprise use — saw inference margins climb from 38% to over 70% during 2026 due to the popularity of Claude Code, according to SemiAnalysis .

AI software costs more than cloud-based software. Providers have responded with price changes. For example, heavy agentic use forced Cursor to replace its $20 flat plan with credits-based repricing . Uber used up its entire 2026 budget in four months . Salesforce has changed its Agentforce pricing models three times in eighteen months.

AI cloud services providers — with the exception of CoreWeave — are profitable. CoreWeave has deeply negative free cash flow, a roughly $99 billion backlog and about $24.9 billion of debt .

Other hyperscalers are part of larger companies — such as Amazon, Microsoft and Google — which are expected to spend $805 billion on capital expenditures in 2026 and as much as $1.1 trillion in 2027, per Morgan Stanley.

AWS is profitable ( 37.7% operating margin) as is Google Cloud ( 35.6% operating margin). It is difficult to tell whether Microsoft’s AI cloud services are making money since it bundles Azure and AI infrastructure into its broader Intelligent Cloud segment.

Chips are highly profitable. Nvidia’s net margin is 63% . Other chip makers are also highly profitable — such as TSMC ( 46.5% ), SanDisk ( 34.2% ) and Micron ( 56% ). With demand vastly exceeding supply and new production capacity not coming online for at least a year, prices and margins are rising.

Who’s Positioned To Benefit — And Who Isn’t

The winners are companies that own the token-cost floor or capture the deflation. Nvidia, TSMC and the high-bandwidth memory companies such as SanDisk benefit as cheaper tokens drive more demand. Google is uniquely benefiting from both custom silicon — helped by a deal with Anthropic to use the company’s Tensor Processing Units — and distribution.

Anthropic benefits from its enterprise focus, and users of AI are getting more bang for their buck by routing routine work to a $2 model while reserving a $25 model for the most difficult 20% of their computing tasks..

The losers suffer from the gap between flat-rate pricing and unhedged token costs. Indebted neoclouds with single-customer concentration are at risk, as are leading U.S. frontier model providers, which are just 2.7% ahead of China’s best, noted Stanford’s 2026 AI Index .

What Investors Should Watch

Investors should track the following trends:

  • Price per million tokens using Artificial Analysis and OpenRouter , where the open-versus-closed and China-versus-U.S. token share is also visible.
  • Disclosed token volumes and whether growth still outpaces price declines from sources such as Yipit Data.
  • Inference gross margins from ICONIQ Capital, Bessemer Venture Partners and specialized publications like SemiAnalysis.
  • AI cloud capex and depreciation assumptions using sources such as Goldman Sachs Insights and company reports.
  • Neocloud backlog versus free cash flow from financial reports produced by vendors like CoreWeave.
  • HBM and packaging capacity from the SemiAnalysis Accelerator & HBM Model and the Silicon Analysts Foundry Allocation Dashboard .
  • AI data center power prices best found at International Energy Agency regional grid operators like PJM Interconnection and specialized energy research groups such as the MIT Energy Initiative.

Ultimately the future of the AI industry depends on whether the value of AI exceeds its price. For now, companies are working on lowering the price.

A case in point is Telnyx, whose AI agent costs were $100,000 a day using Anthropic. This cost sent the company to open models. “They worked,” CEO David Casem told the Journal. “It’s not like we don’t use OpenAI or Anthropic models, we still do. They just don’t do everything anymore,” he added.

The company’s 1,400 agents use Chinese startup Z.AI’s product which costs around $100 per agent per day. Anthropic’s most powerful model, Fable, acts as a conductor that plans out work while open-weight models do the implementation. OpenAI’s Sol handles a review of what the open-weight models produce, added the Journal.

Investors should consider placing their bets on the companies with the most profit potential in light of the forces driving Telnyx’s AI consumption recipe.