Paying By The Token: Why The Unit Economics Of Clinical AI Matters
Hospitals have now begun buying artificial intelligence by the token, converting an obscure artifact of large language-model (LLM) architecture into a line item that finance departments now track alongside vendor contracts and compliance costs .
For clinicians and clinical investigators, this is not an information-technology curiosity. The token is becoming the cost denominator of every generative tool we are being asked to adopt, pilot, and publish on--and we are currently reporting “efficacy” without reporting “dose.”
The trajectory is not hypothetical, and it is not confined to medicine. In June 2026, Gartner forecast that by 2028, the cost of AI coding agents will exceed the salary of the average developer using them, driven by rising token consumption and an industry-wide migration from seat-based to consumption-based licensing .
Software engineering is about two years ahead of medicine on this adoption curve. It has arrived at the point where the tool costs more than the professional it was purchased to make more productive.
There is no structural reason medicine is exempt, and several reasons--encounter length, context breadth, regulatory retention requirements--to think our per-task consumption will be higher.
What An ‘AI token’ Is And Why The Unit Is Unstable
LLMs do not process words. They process tokens: sub-word fragments, whole words, or punctuation, converted to numeric representations and handled sequentially. To provide better framing, roughly four characters--about three-quarters of a word--correspond to one token .
Vendors monitor both directions of the exchange: the context fed into the model and the text it generates. Output tokens typically carry three to five times the price of input tokens, because generation is sequential while input processing can run in parallel.
The unit is not one price. Published list prices in mid-2026 span roughly $0.02 to $15 per million input tokens and $0.14 to $75 per million output tokens across model tiers — a spread of two to three orders of magnitude for a nominally identical unit. A tool that silently routes an unchanged clinical task to a higher tier can multiply its cost without any change to the contract, the workflow, or the note that appears in the chart. This is why the model tier is not a technical footnote. It is the formulation of the intervention.
The clinically important consequence is that consumption is a function of behavior, not of headcount. A longer encounter, a wider context window that pulls in prior notes and problem lists, a more capable model tier, or a workflow that resends accumulated context at every step all increase token volume for the same clinical task. A single query can generate thousands to millions of tokens before anyone in the organization notices.
Unit prices have significantly collapsed. The Stanford AI Index documented that querying a model performing at GPT-3.5 level on MMLU fell from $20.00 per million tokens in November 2022 to $0.07 per million tokens by October 2024--a more than 280-fold reduction in roughly 18 months, with task-specific declines ranging from ninefold to 900-fold per year . And yet aggregate spending is still increasing. For example, Uber exhausted an annual agentic AI budget in three months; Google has reported token consumption seven times higher year over year.
Three mechanisms explain the paradox, and each is architectural rather than incidental. The first mechanism is capability inflation: as frontier tiers reset upward, organizations migrate to the expensive ceiling instead of banking savings at the collapsing floor. The second mechanism is context retransmission: a model is stateless, so a multi-step workflow must resend its entire accumulated history on every subsequent call.
A twenty-step task pays for the original prompt twenty times, which means the cost of a task grows with roughly the square of its length rather than in proportion to it. And finally, the third mechanism is reasoning expansion: current models spend billed output tokens on internal deliberation the clinician never sees, so the note that appears in the chart represents only a fraction of what was purchased to produce it.
A fourth consideration is that the floor itself may not hold. Current pricing reflects a market-share phase in which vendors price aggressively against enormous capital expenditure on computing power.
Gartner's forecast rests partly on the expectation that model pricing will rise as infrastructure investment and unresolved profitability push costs back onto buyers. A health system building a five-year AI budget on straight-line extrapolation of the 2022–2024 deflation curve is extrapolating from a subsidized period and should model at least one scenario in which the per-unit price rises while consumption also rises.
None of these mechanisms appears in a hospital budget, and all of them multiply consumption for clinically identical workflows. That is the essential point for anyone evaluating these tools: unit price and total cost have decoupled. A 280-fold collapse in the cost of a token is fully compatible with a tenfold increase in what an organization pays, because the same architectural advances that made tokens cheap also made systems that consume vastly more of them. This is Jevons’Paradox , and it is the reason a favorable per-token quote at contract signing carries almost no information about the bill two renewals from now.
The Failure Modes Are Already Documented Outside Medicine
Gartner's analysis names the specific governance failures that drive token overspending: ungoverned autonomy in agent-driven workflows, bloated context windows, and the absence of structured feedback on usage. It also notes that many vendors do not disclose how consumption is calculated and billed, which limits an enterprise's ability to forecast or control it, and that mature built-in cost- optimization tooling has not yet arrived. Each of these approaches translates without modification into a clinical deployment. In practical terms, a wider chart-context setting is a bloated context window. An agent that summarizes a patient's entire longitudinal record before drafting an inbox reply is ungoverned autonomy. A department that receives one aggregated monthly invoice has no structured feedback.
The most portable observation is behavioral in nature. Gartner’s analysis states in clearly: token discipline will not emerge through developer choice alone , because users optimize for speed and convenience over cost. If you substitute "clinician" for "developer" then we have created the conditions under which ambient documentation, inbox drafting, and chart summarization are being deployed. A physician at the bedside optimizes for speed and cognitive accuracy, and the reality is that a single overnight shift is not going to narrow a context window to protect an operating margin. Consumption discipline, if it is going to exist, has to be implemented and made part of the configuration, not requested of the user.
Healthcare's exposure is currently masked
Most health systems do not yet see a token meter. Ambient documentation is sold on flat per-clinician subscriptions , commonly $200 to $600 per clinician per month. That structure insulates the buyer from consumption variance and transfers it to the vendor — who must eventually reprice, tier, or meter. The denominator is not small: Permanente reported more than 2.5 million ambient scribe uses in the first year of deployment, and Epic has stated that over 85% of its customers now use its AI tools, which run on a hyperscaler's hosted model service. Any migration from subscription to consumption pricing at that scale is a material financial event, and it will not be visible in a clinical outcome.
The measurement infrastructure that would detect such a migration is not yet in place. A 2026 study from UPMC's Center for Connected Medicine and KLAS Research found that 93% of surveyed health system and ambulatory leaders reported deploying third-party AI solutions, but only 44% had a dedicated data platform or environment in which to test them. Clinical documentation was the most commonly deployed category at 52%, followed by revenue cycle, coding, and billing applications at 36%. Return on investment and the financial case were named among the top internal barriers to systemwide AI strategy, alongside budget and talent constraints, governance, and change management. Respondents described their data pain points as manual workarounds, spreadsheets, and inconsistent definitions across teams.
The sample is small--27 leaders--and qualitative, and should be read as directional rather than representative. But the direction is unambiguous and consistent with what clinical leaders describe informally: adoption has outrun measurement. An organization that lacks an environment in which to test a tool also lacks the instrument with which to meter it. That gap is tolerable under a flat subscription. It is not tolerable under consumption pricing, and it takes longer to build than a contract takes to renew.
Run The Arithmetic On The Available Evidence
The best available direct-productivity data come from a cohort study of nearly 1.2 million ambulatory encounters involving roughly 1,565 physicians at a single academic center. Ambient scribe access was associated with a 5.8% increase in weekly work relative value units--about 1.81 additional wRVUs per week--and a 2.8% increase in encounters, without an increase in claim denials.
Take those figures at face value. Roughly 1.8 incremental wRVUs per week across about 46 clinical weeks yields on the order of 83 additional wRVUs per physician per year. Even at a generous, realized professional-revenue conversion, the direct return lands in the low thousands of dollars annually, against a subscription cost of $2,400 to $7,200 per clinician per year. The margin is real but narrow, derived from one site, and critically asymmetric in its risk profile: the numerator is fixed by a fee schedule that does not care how many tokens were consumed, while the denominator floats with note length, context breadth, and model tier.
That asymmetry is the clinical translation of the Gartner benchmark. The relevant question is not whether ambient AI produces value; the cohort data suggest it does. The question is what happens to a narrow positive margin when the cost side is indexed to consumption and the revenue side is indexed to a fee schedule. In software engineering, that arithmetic is projected to invert within two years. Medicine's version of the same calculation has a numerator that is centrally administered and, in most years, does not grow.
The larger return Is The Harder One To Measure
The stronger financial case is indirect. Reductions in documentation burden and burnout of 20% to 40% have been reported across multiple systems and vendors; burnout doubles or triples the likelihood of physician turnover, and each departing physician is estimated to cost $500,000 to $1 million in recruitment, onboarding, and lost productivity. Even modest retention gains overwhelm the subscription line. But this is precisely the class of return that is easy to assert, difficult to attribute, and impossible to audit from an invoice--which is why the cost side deserves the same methodological rigor we demand of the benefit side.
The bill arrives Even Where The Meter Does Not
There is a second channel of exposure that no ambient scribe contract will reveal, and it is arguably the larger one. Vizient’s mid-2026 spend forecast projects overall healthcare supply chain price inflation of 3.39% for 2027, with indirect spend and purchased services rising 4.73%--faster than pharmaceuticals at 3.54%. Within that, IT hardware and software is projected at 6.29%, including 8.5% for IT security, 8% for IT hardware and accessories, and 7.5% for IT software and licensing .
Vizient's analysts identify AI demand as a primary driver, alongside energy and logistics costs: as demand for computing infrastructure grows across industries, health systems compete for the same storage, servers, and IT services as everyone else. These increases land against hospital margins that already trail 2025 levels, ahead of reimbursement headwinds.
The implication for clinical readers deserves emphasis, because it is counterintuitive. A health system that never signs a consumption-priced contract, never deploys an agent, and never buys a single token is still paying for the AI build-out through the price of the servers, licenses, and security infrastructure underneath everything else it does. That cost is real, recurring, and attributable to no clinical tool, department, and certainly no outcome. It will not appear in any evaluation of an ambient scribe. It is the part of the drain that no ROI analysis is structured to capture, and it is why AI cost containment cannot be delegated entirely to whoever negotiates the documentation contract.
A Reporting Standard For The Literature
We would not accept an interventional trial that reported effect size without reporting dose. Token consumption is the dose of a generative clinical intervention, and the model tier is its formulation. Studies of clinical LLM deployment should report, at minimum: mean input and output tokens per encounter or task, including reasoning tokens; the model tier and context configuration in use, since vendors update both silently; cost per completed task rather than cost per prompt; and, where clinical benefit is claimed, cost per unit of that benefit.
One practical obstacle should be named. If vendors do not consistently disclose how consumption is calculated and billed, investigators will not be able to extract this from a dashboard after the fact.
Computational dose disclosure therefore belongs in the deployment agreement, negotiated before a study begins, in the same way that access to source data is negotiated before an industry-sponsored trial. Research offices and contracting offices need to be in the same conversation.
The reproducibility argument is straightforward. The same ambient tool running on a different model tier with a different context window is not the same intervention--it may differ severalfold in cost and measurably in output. Without disclosure of the computational dose, a favorable single-site result cannot be replicated, priced, or generalized.
Questions To Ask Before Signing
Health systems that are ahead of this have adopted concrete controls: restricting agent-building to evaluated users and shutting off any agent whose cost is untethered from a demonstrated return, as Houston Methodist has done; negotiating price protection and defining who absorbs token inflation at renewal; demanding per-department consumption telemetry rather than a single aggregated bill; and requiring model routing and prompt caching, which materially reduce spend when configured deliberately.
The engineering-sector governance recommendations translate cleanly and are worth adopting before the meter becomes visible: define in advance which tasks warrant an agent and at what level of autonomy; match model tier to task complexity rather than defaulting to the frontier for every note; treat context configuration as a deliberate engineering decision rather than a vendor default; and embed consumption review into operational cycles instead of discovering it at renewal. Add to these the gap the UPMC and KLAS data identify--a dedicated environment in which tools can be tested and measured-because a control you cannot measure is an assertion. Clinical leadership belongs in each of these conversations, because every one of them is a decision about which model writes the note.
The Digitization Precedent
Medicine has already struggled through a version of this “ experiment ”. Electronic record adoption succeeded on its own terms and was paid for over fifteen years in recurring cost, vendor consolidation, practice closures, clinician burnout, and a degraded clinical narrative. Digitization was procured on capital-and-license logic. Generative AI is being procured on consumption logic, and the meter runs in a unit almost no clinician has been trained to read.
The difference this time is that the warning arrived early and from an adjacent industry rather than in retrospect from our own profession. Software engineering is already reporting that the tool may cost more than physicians; hospital supply chain analysts are already pricing AI demand into annual budgets for software and security line items; and governance surveys already show adoption outpacing the capacity to test and measure its outcomes and clinical applicability.
While physician leaders in the C- suite have certainly learned to interpret a confidence interval for academic purposes, the stark reality is that they must now become equally fluent in cost per million tokens--because that number will shape what tools reach the bedside, and what happens to them when any vendor AI contract is up for renewal.
Dr. Peter Papadakos , Professor of Anesthesiology and Critical Care at University of Rochester, contributed to this article.
Loading article...