Frontier model companies are leaning more heavily into high-powered models that are faster and cheaper, but sometimes trading one for the other. OpenAI’s latest push to make AI faster and more economical are two new offerings that address different parts of that AI performance equation. The company is rolling out Ultrafast for GPT-6.1 Sol, promising up to eight times faster token generation in Codex, alongside a new Decisions API that handles the quick judgments AI applications and autonomous agents make throughout their work.

The announcements reflect a growing industry focus on how much time and money AI consumes to get useful work done, more so than simple performance on benchmarks.

The question of cost and speed is becoming increasingly important as companies move from prompt and chat-based interactions to longer-lived automated workflows involving hundreds or thousands of model calls.

Saving a few cents or milliseconds on an individual request might seem insignificant, but those savings add up when businesses process millions of transactions. The economics become even more compelling when developers can use inexpensive specialized models for tasks that don’t require sophisticated reasoning.

Another factor driving the demand for faster, cheaper model access are open weight and foreign models that are steadily gaining market and developer share. Chinese AI models, including DeepSeek, Moonshot, Kimi and Alibaba’s Qwen, are competing aggressively on price and performance. Open-weight models provide alternatives to expensive proprietary APIs.

Entering the market are startups such as TypeSafe AI that are building specialized systems that provide focused models that address parts of what developers use AI models for, which question whether expensive foundation generative language models are necessary for every task. The markets are maturing and competition is emerging over which providers can deliver useful intelligence at prices that make large-scale automation commercially practical.

OpenAI Is Offering More Speed, But At A Price

OpenAI recently introduced GPT-6.1 Sol provides what OpenAI describes as near-Astra intelligence for agentic coding, computer use and professional tasks at approximately one-fifth of GPT-6 Astra’s standard token prices.

GPT-6.1 Sol costs $2 per million input tokens and $10 per million output tokens, compared with $10 and $50 for Astra. OpenAI reports that Sol approaches Astra’s performance on several demanding benchmarks, providing developers with a less expensive option for complex work. These comparisons are based on OpenAI’s reported evaluations, and the actual performance differences will depend on the applications involved.

The company’s new Ultrafast mode focuses on speed. OpenAI says Ultrafast delivers up to eight times faster token generation than Standard mode in Codex. Developers can use it through the API, with access also available in Codex and ChatGPT Work on eligible plans.

But there’s a significant catch. Ultrafast costs six times the standard API price , or $12 per million input tokens and $60 per million output tokens. So the speed comes as a tradeoff in price. Developers have to weigh their priorities.

For a coding agent that would otherwise spend several minutes generating, debugging and testing software, that premium might be worthwhile. A developer could save enough time to justify the additional expense. For a background process analyzing documents overnight, paying six times more for faster responses might accomplish very little.

With this added dimension, speed, performance and price can now be adjusted separately, at least for certain workloads. Developers will need to determine which combination actually improves their application economics.

AI Doesn’t Need To Think So Hard About Everything

OpenAI’s new Decisions API addresses builds upon the wave of new cheaper, faster and more focused decision models that have taken off since the recent announcement of Jev by Typesafe.

Consider the everyday decisions an autonomous agent might make. Is this email spam? Does this webpage contain useful information? Which tool should the agent use next? Should a transaction be flagged for review?

These questions often require a simple classification, score or selection among predefined options. But developers frequently use general-purpose language models designed to generate lots of text and perform many expensive reasoning steps to perform these relatively narrow tasks.

That’s an expensive way to answer yes or no or classify things into buckets.

The Decisions API, powered by GPT-6 Luna, accepts text and image inputs and produces structured predictions, including classifications, numerical scores and probability estimates. OpenAI says it can respond up to ten times faster than comparable requests through its conventional Responses API. The service is priced at $0.10 per million input tokens, without a separate output-token charge.

With this approach, OpenAI has touted a number of early users are reporting substantial improvements.

HubSpot has been testing the Decisions API for pull request triage, where its Sidekick system must determine which software changes require further evaluation. According to results shared by OpenAI, early testing found that the new approach was six to fifteen times faster than HubSpot’s existing agent-based decision system, at approximately half the cost and with half the false-positive rate.

Parallel Web Systems reported near-perfect filtering of low-quality webpages and an 88.8% area under the precision-recall curve for identifying changes to webpages. A Mapbox developer demonstrated using the API to screen image tiles before passing selected images to more computationally demanding models for feature extraction.

While these are early tests that haven’t had enough production time to show long term results, they illustrate an opportunity that enterprise AI developers have often overlooked. Many steps in an AI workflow never required generative language capabilities in the first place.

Chinese And Open-Weight Models Are Driving Price Competition

Open-weight models and models from Chinese AI developers have increasingly shown better performance at prices that are becoming significantly competitive to more expensive systems from American AI companies.

DeepSeek’s published API pricing lists V4.1 Flash at $0.30 per million input tokens and $1.20 per million output tokens during peak hours, with rates falling by half during off-peak periods. Alibaba’s Qwen3.8 Flash offers rates starting at $0.113 per million input tokens and $0.382 per million output tokens for certain deployments.

These prices are substantially below those of many frontier proprietary models. These models also perform competitively on benchmarks, and while none of them have yet overtaken any of the frontier models, their capabilities match those of the best models from earlier in 2026. For many companies, this is already sufficient capability.

While different models have different reasoning capabilities, token requirements, response times and failure rates, the price differences are large enough to force customers to ask what they’re getting for the premium.

Open-weight models add another complication for proprietary AI providers. Companies can deploy qualifying models on their own infrastructure or purchase hosted access from competing providers. Organizations with sufficient technical expertise gain more control over their deployments, more ownership over their data inputs and outputs and aren’t necessarily locked into one vendor’s API pricing.

New entrants are responding to that competition. On October 5, Nvidia-backed Reflection AI announced Beam , a 501-billion-parameter open-weight model designed for coding, reasoning and agentic tasks. Its mixture-of-experts architecture activates just 23 billion parameters for a given request, reducing the amount of computation needed compared with activating the entire model.

If a company can operate a capable model on its own infrastructure, or obtain hosted access from several competing providers at much lower prices, how much extra will it pay for OpenAI or Anthropic? Better reasoning performance is one answer, but only when customers can measure the value of that advantage in the work they’re actually doing.

On the other hand, a cheaper AI model doesn’t necessarily make an AI application cheaper to operate.

A model charging half as much per token might require more tokens, take longer to finish a task or make mistakes that trigger expensive retries. A premium model that resolves a problem correctly on the first attempt could be the better bargain.

This gets especially complicated with autonomous agents. A single request can trigger dozens of model calls, database queries, decisions and validation steps. The initial prompt may account for only a small portion of the total expense.

Consider a customer service agent processing a disputed payment. It might use a cheap classification model to identify the complaint, another model to retrieve transaction records and a reasoning model to decide whether the case requires human intervention. Running every step through the most expensive frontier model would waste money. Running everything through the cheapest model could be equally costly if errors require additional work.

This is actually a familiar tradeoff for many developers and those experienced with enterprise IT, especially anyone who manages cloud infrastructure. Faster isn’t always worth the premium. Cheaper isn’t always more economical. The challenge is allocating computing resources where they produce the greatest return. For AI providers, that means factoring in the total cost of getting a customer’s work done correctly.

None of this means the race for more powerful AI is over. The most demanding coding, research and reasoning tasks will continue to reward superior capabilities.

But the question for AI developers is becoming increasingly practical. If a ten-cent model can reliably accomplish the same task as a ten-dollar model, why pay the difference? Companies will increasingly have to decide how much intelligence they actually need, how quickly they need it, and what they’re willing to pay.

For AI providers investing billions in ever more powerful models, they are starting to realize that customers may care less about who has the smartest AI and more about who delivers the best results for the money.