In late 2024, Heidi cofounder Yu Liu was shocked to discover that he was spending up to $3 million every month on the latest models from OpenAI, Anthropic and Google to power his note-taking startup for doctors. “We were making good revenues, but if I [subtracted] our AI bills, we barely had any margins,” Liu says.

At first, Liu ignored the outrageous costs, expecting the AI labs to lower their prices as intelligence became cheaper to serve and more abundantly available. That didn’t happen. Instead, the bills kept piling up. By mid-2025, he’d had enough. He turned to another startup, San Mateo, California–based Fireworks, which helped Heidi build its own, cheaper AI using the 280 open models hosted on Fireworks’ site. Unlike with pricey, off-the-shelf models from the big American labs, developers can adjust how these models perform by feeding them data, creating customized models that bring down costs.

By moving 95% of his startup’s workload to specialized, open-weight models, Liu says he slashed his monthly AI bill by almost 90%. That’s significant savings for his business, which brought in $4.1 million in April. There were other benefits, too: While a standard model took 25 seconds to spit out transcriptions during peak hours, Heidi’s own bespoke model could do it almost instantly.

AI sticker shock has become the topic du jour for corporate America. For the last couple years, companies gave employees one mandate: Use as much AI as possible. Incentives became perverse. Performance reviews rewarded workers who burned through expensive tokens—the standard unit of AI usage. On internal leaderboards, bonus-obsessed “tokenmaxxing” employees competed against each other for titles like “token legend” and “AI god.”

Then the bill arrived. Uber famously blew through its entire 2026 budget for AI in four months using Anthropic’s coding tool Claude Code. Amazon reportedly spent $1.8 million on a Claude project that never launched. Salesforce expects to rack up a $300 million Anthropic bill this year. “It’s not sustainable,” says Kay Zhu, CTO of Genspark, which builds AI tools for knowledge workers.

But it is understandable. Alarmed by the pace of progress, companies are racing to adopt AI—both by arming their employees with it and infusing it into their products. “They don’t want to hold back,” Menlo Ventures partner Matt Murphy says. But they must also confront the steep costs tied to it. AI is inherently expensive because it gobbles up huge amounts of resources—GPUs, electricity, water, real estate—all of which cost gobs of money. A single gigawatt data center can cost upward of $40 billion. America is currently building more than 1,500 data centers, at an aggregate cost of $2.3 trillion . “It significantly shifts the business cost structure,” says Fireworks CEO and cofoun­der Lin Qiao. For companies building products that use AI, that can become a glaring problem as they scale from hundreds to thousands to millions of users. “When you hit product-market fit, one scenario is you scale into bankruptcy,” Qiao says.

Gartner forecasts global AI spending to reach $2.6 trillion this year, or 2% of planetary GDP. In August, Nvidia struck a deal with six Wall Street giants, including Blackstone and Goldman Sachs, to secure $500 billion in AI financing. It hinges on one assumption: that there will be near-unlimited demand, enough to justify it all. Between them, OpenAI and Anthropic have raised $310 billion in venture funding. They need it because their AI products are ludicrously expensive to build and run. They’re growing fast, with trailing 12-month revenue estimated at around $25 billion each, but that’s not even close to the amount they’re spending.

“Businesses don’t even know how much they’re spending on every query. And if they did, they’d be disgus­ted,” says Anastasios Angelopoulos, CEO of Arena, a site that rates models on performance and price. Now, with trillions pouring into the infrastructure buildout, the mad scramble among companies to cut AI costs raises the question: Will there ultimately be enough demand for the math to work out for top-tier AI labs?

After all, there are drastically cheaper alternatives to the pricey frontier models from OpenAI and Anthropic. Chinese firms like DeepSeek, Alibaba, Mini­Max and Moonshot used to trail the big U.S. labs, but they’ve recently caught up. And they’re dirt cheap. Moonshot’s Kimi K3 can churn out 1,500 pages (750,000 words) of text for $15, less than a third of the $50 that Anthropic charges for Fable 5 to do the same job. Alibaba’s Qwen 3.8-Max can design a website’s landing page for 12 cents. OpenAI’s GPT 5.6 Sol will do the same chore—for more than three times the price.

The open-source shift is in full swing. Coinbase cut its AI spending roughly in half by routing simple tasks like basic programming to cheaper Chinese models. DoorDash now delegates “lower level” jobs like code review to Moonshot’s Kimi K2.6, among others, reserving Anthropic’s frontier model for tougher tasks. Airbnb CEO Brian Chesky has called Alibaba’s Qwen model “fast and cheap,” and uses it for the company’s customer service chatbot. In August, New York City–based OpenRouter, whose software gives developers access to hundreds of AI models, reported that 55% of tokens passing through its software went to Chinese models, up from 4.5% in early 2025 . “The reality is right now, the best open models are from China,” says Matan Grinberg, CEO of AI coding startup Factory, though he adds homegrown open-weight models like Nvidia’s Nemotron are “pretty good” too.

The U.S. labs aren’t taking such fierce competition lying down. In July, OpenAI cut prices on its mid-tier and smaller models by up to 80%, and Anthropic released models that cost half as much as its most capable ones. In general, the price of tokens has plummeted some 70% since the start of the year. “We should all be cheering when­ever there’s more competition between the . . . models, because that means we are all going to get the best experience for the cheapest cost,” Grinberg says. “If the open models didn’t exist, I don’t think OpenAI would have cut the cost as much.”

Even with cheaper tokens costs can balloon as AI gets more capable. Today’s best models aren’t just chatbots answering questions. Tools like Claude Code spin up dozens if not hundreds of agents that all work together to solve problems, spending more time “thinking”—and using many more tokens. “The token number actually increased faster than the token cost decreased,” says Genspark’s Zhu, who is spending 50% less on tasks like PowerPoint slide generation after his Palo Alto–based startup switched to Chinese open-source. “People are saying hello and they’re being charged $5 by Fable,” Angelopoulos adds. “It’s crazy.”

In response, a group of startups is helping businesses get smarter about how they use AI. The fixes range from routers that intelligently move AI traffic to cheaper models to agents that sit silently on employee laptops, monitoring and tracking their AI spend. One researcher is even working on an entirely new AI architecture that uses fewer resources, inspired by a tiny roundworm’s brain.

“Companies are starting to open their eyes to the very clear future in which intelligence becomes the number one line item in everyone’s operating expenses for knowledge work,” says OpenRouter CEO and cofoun­der Alex Atallah, 34.

O f these upstarts, four-year-old Fireworks has perhaps capitalized the most on American businesses’ desperate need to rein in their artificial-intelligence spen­ding. The company’s software helps its roughly 10,000 enterprise custo­mers—including Uber, Shopify, Upwork and DoorDash, as well as marquee startups like Cursor, Cognition and Harvey—build specialized AI using their own proprietary data on top of cheaper, often Chinese models. That doesn’t just cut costs. It gives them more granular control over how their models work.

With just a few lines of code and a credit card, developers can experiment, fine-tune and run any open model in the world. There’s no need to worry about getting your hands on expensive, in-demand GPUs or reengineering hardware to run new models efficiently. Fireworks takes care of it all.

The company booked about $250 million in 2025 revenue, according to Forbes estimates. Its valuation has grown fourfold to $17.5 billion in the past nine months, making the 50-year-old Qiao a new billionaire, with an estimated 6% stake worth $1 billion.

The child of an accountant mother and a father who worked as a mechanical engineer designing cargo ships, Qiao grew up in Chengdu, China, which is mostly known for its giant pandas. She came to the U.S. in 2000 for a Ph.D. in computer science at the University of California, Santa Barbara. After stints at IBM and LinkedIn, she spent seven years at Meta building PyTorch, a popular open-source tool with 2 billion downloads that was used to build ChatGPT.

In September 2022, she left Meta with five of her colleagues to start Fireworks (a seventh cofounder joined from Google). Benchmark partner Eric Vishria, who wrote the company’s $25 million Series A check, decided to invest 10 minutes into their meeting. Now, growth is exploding. At its July run rate, Fireworks would generate about $1 billion in revenue in the 12 months ending July 2027.

Open-weight models’ appealing economics are also propelling Open­Router. Its software breaks a task into a series of separate steps and auto­matically sends each one to the most cost-efficient of 500 open-weight and frontier models. Less important jobs, such as fetching a file from a directory, are sent to a cheaper DeepSeek model, while it might be worth it to spend the money to have Anthropic plan out the software architecture for a new product.

Every week 100 trillion tokens pass through OpenRouter’s system thanks to customers like Zoom and Nvidia. That data gives the startup a bird’s-eye view of how employees are “voting with their dollars,” says Atallah, who previously cofounded NFT marketplace OpenSea in 2017. His startup’s router looks at spending patterns for the last seven days to decide which model to send a request to. “An exciting part about this router is it leverages the wisdom of the crowds to work,” he says.

Back in 2023, Atallah realized there would soon be thousands of cheaper, smaller models when a group of Stanford students trained one called Alpaca (on top of Meta’s Llama model) for just $600: “It seemed like we really needed a good way of exploring all of these models.”

In August Stripe agreed to acquire OpenRouter for $8 billion, a person familiar with the terms says. It would make Atallah, who owns more than 50% of the business, worth an estimated $4 billion. The company was on track to make over $100 million in revenue this year with 90% gross margins, the person added.

The Trump administration has considered banning American companies from using Chinese models under the hood—hard to enforce, when they’re free to download from the open web. Hundreds of companies have opposed such a ban, arguing they would fall behind as the rest of the world builds on cheaper open-source AI. “It’s not just a risk for our business,” says Fireworks’ Qiao. “It’s a risk for the whole U.S. economy.”

A key point: Fireworks and OpenRouter keep American data in America. The startups act as a buffer between Chinese models and U.S. companies. “Wherever the model is created, it effectively becomes a Fireworks model because we’re serving it, we’re running it, not anyone else. This is not data that’s going anywhere else,” George Hu, president at Fireworks, says. Atallah says any work that is routed to Chinese open-weight models can be sent to non-Chinese data centers, keeping U.S. businesses’ data secure.

E ven when everything goes right, AI is expensive. But accidents are common. Agents left running in the background can spin out of control, go off on a tangent and rack up hundreds of thousands of dollars in bills. Fireworks’ own internal recruiting agent blew $24,000 in just a few days because it was misconfigured to “max mode,” where it spent more time “thinking” and burned more tokens. Shopify has built so-called circuit breakers to catch unattended agents that were left running by mistake. “Our concern is the spend we can’t account for, because waste is usually a signal that something is broken,” Shopify’s head of engineering, Farhan Thawar, says.

When a C-suite executive at an advertising company started using Genspark’s AI-powered meeting note taker, he frittered away an entire month of credits in just a few days. The reason? The AI tool was recording and summarizing the eight hours’ worth of meetings he sat through every day—by default. After he flagged the problem, Genspark replaced the model with an alternative that was a third of the price.

Such waste occurs when enterprises have no way to track where their money is going and what they’re getting in return, says Jyoti Bansal, CEO and cofounder of Harness. The San Francisco–based startup sells software that scans for quality issues, safety bugs and compliance loopholes in code. In May, Bansal’s team launched an AI cost-management agent that sits on employees’ laptops and tracks every token spent, the cost per token and whether employees actually accomplished the task they set out to do with AI. “Just by having visibility, people can start fixing things,” says Bansal, who is worth $2.3 billion.

United Airlines is building its own internal tracker to manage AI expenses, with human approvals to keep costs in check. Ratna Devarapalli, the airline’s IT director, says the rising prices of tools like Microsoft Copilot and overuse of AI as people learn how to use it has led to cost spikes. “It’s also addictive,” he says. “I’m sitting at home using it for hours trying to do something.”

Developer tools startup Vercel, a Fireworks customer headquartered in San Francisco, still wants its engineers to use the strongest frontier models but has placed caps on usage. But open models have no restrictions. Most engineers now use frontier to plan and open-weight to execute, CTO Malte Ubl says. Employees at Coinbase and their managers get a regular report breaking down how much they spent on models. “We’re not blocking you from using a lot, but we are making clear you’re going to have to explain why you spent so much money last week on a model,” says Coinbase CTO Rob Witoff.

AI is expensive because of its very nature. The “transformer” architecture, which undergirds ChatGPT, Claude and all those Chinese open models, is basically “an act of multiplication,” explains Liquid AI CEO and cofounder Ramin Hasani. The more information that passes through it, the more expensive it becomes.

But what if there were a more efficient way to train models? Hasani spent 11 years as a researcher, first at Vienna University of Technology and then at MIT, trying to find it. His answer, surprisingly, lay in a tiny roundworm called C. elegans, whose brain processes information similarly to a human’s. Learning from the worm’s neural signals, Hasani discovered a mathematical process that relies on addition instead of multiplication.

Using that insight, his Boston-based startup, which spun out of MIT in 2023, is building “liquid neural networks” that process data more efficiently and don’t need expensive chips stored in data centers. The startup is worth $4.5 billion; Hasani owns an estimated 20%.

Liquid AI’s models are small enough that they can run directly on laptops and phones. There’s no cost to host them. Mercedes-Benz, one of its customers, plans to use Liquid’s models for the voice assistant inside the auto­maker’s North American cars. “Enterprises are thinking, can I actually have a ChatGPT-level quality model that runs directly inside my car on a chip that is $60?” Hasani says. Compare that to Nvidia chips that cost $50,000 each.

We’re still in the early innings of the AI rollout. An expensive period of trial and error lies ahead. AI’s economics are bound to change as the big labs compete in fierce pricing wars. But one thing is clear: Intelligence has become a permanent expense on the balance sheet—and every competent CFO in America is going to be looking for ways to manage it.