Gateway & Ground

LLM Token Cost Benchmarking Across Major Providers

Advertised per-token prices hide the real costs that determine your LLM bill.

Correspondent · · 9 min read
Cover illustration for “LLM Token Cost Benchmarking Across Major Providers”
AI Spend Management · October 2, 2026 · 9 min read · 1,937 words

A chatbot handling customer support can cost anywhere from $500 to $15,000 a month to run, depending on which model answers the tickets, how the prompts are built, and whether caching is set up at all. No input-price table predicts that spread. The whole exercise of comparing providers by their advertised dollars-per-million-input-tokens number misses the line items that actually decide the bill: output tokens, cache writes, cache reads, and in some models, reasoning tokens the caller never even sees. Two providers can post nearly identical input prices and still produce wildly different invoices once real traffic hits the API. The rest of this piece works through why, and how to calculate the number that actually matters before a model gets chosen rather than after the first invoice lands.

The four token types that compose your bill

Every major provider bills across at least three distinct token categories, input, cached input, and output, each priced on its own schedule, and treating them as one blended "price" produces a cost estimate that's wrong from the start.

Input tokens, the prompt, the system instructions, the conversation history, and the retrieved context, are the cheapest category per token. Teams over-provision this category the most as a result. A bloated system prompt sent on every single request adds up fast across thousands of calls, even at a low per-token rate.

Output tokens, the generated response, cost more to produce because generating text is more compute-intensive than reading it, and that's reflected in the pricing: output is billed at a multiple of the input rate across every major provider. This multiple is the single biggest lever in the whole cost structure, and it gets its own section below.

Cache read tokens offer the steepest discount in the entire pricing structure. When the same input text gets reused within a provider's cache window, the portion that's re-read bills at a fraction of the standard input rate, which makes cache hit rate one of the most powerful cost controls available to any team running repeated prompts.

Cache write tokens work against that logic at first glance: writing content into cache can cost more than paying standard input rates for it. Anthropic charges double the standard rate for a one-hour cache write. That extra cost only pays for itself if the cached content gets read back multiple times, so a cache write on a prompt used once is pure waste, while the same write on a prompt reused hundreds of times turns into the cheapest tokens on the bill.

Reasoning tokens apply to a subset of models, where the model generates internal "thinking" tokens that get billed on the invoice without ever appearing in the visible response. A team estimating cost by counting the words in the prompt and the words in the answer will miss this category completely, because it appears only on the bill, never in the prompt or response text.

The output multiplier and provider comparisons

Once a workload's real traffic ratio is applied, the actual split between input tokens sent and output tokens generated, providers that look close on headline price can diverge sharply, and the ranking between them can flip depending on how that workload is shaped. The formula itself is simple arithmetic: blended rate equals input token share times input price, plus output token share times output price.

The Spheron blended-rate table makes the mechanism visible across three workload shapes: input-heavy, balanced, and output-heavy. On an input-heavy split, Grok 4.3 and GPT-5.4-mini are tied dollar for dollar at the same blended rate. Shift the same two models to an output-heavy split, and they diverge substantially. The entire gap traces back to the output multiplier: one model prices output at 2x its input rate, the other at a much higher multiple, and that difference compounds every time the workload skews toward longer generated responses rather than longer prompts.

The practical consequence is concrete: a team has to pull actual input and output token counts from production logs before choosing a provider on price. A substantial swing in effective cost between two providers with similar headline rates appears routinely once the real ratio is computed. There's no universal answer to which provider wins this comparison, because the answer depends entirely on whether the workload leans toward short questions and long answers, or long context and short answers. Running the same blended-rate formula against both will produce two different winners, even using the exact same two providers.

Diagram: Why Identical Headline Prices Produce Different Bills. Visualizes: Visualize how the same two providers with nearly identical input prices diverge sharply once a real workload ratio is applied.

Caching mechanics as a cost driver

The output multiplier changes the bill by a fixed factor tied to how a provider prices generation. Cache hit rate changes it by a factor that depends entirely on how a team architects its prompts, and that factor can be larger than any difference between provider tiers. This is why caching belongs inside the cost benchmarking calculation itself, not treated as a tuning step applied after a provider has already been chosen.

The scale of the discount is easiest to see in a single comparison. DeepSeek V4-Flash charges $0.0028 per million tokens on a cache hit. That rate isn't just cheaper than premium models' base input prices by orders of magnitude, it's far cheaper than DeepSeek's own output rate on the same model. A workload built to maximize cache hits on that model is paying a fraction of a cent for tokens that would otherwise cost multiples of that on a fresh read.

Different providers implement this mechanism differently, and a caching strategy tuned for one won't necessarily transfer to another. Google's context caching typically strips the majority of cost off standard Gemini rates, and its Batch API adds a further substantial discount on top of that. A team that assumes caching behaves the same way across providers will mis-forecast costs the moment it moves workloads between them, since the discount is not a fixed percentage regardless of architecture.

System prompts make the clearest case for investing in caching deliberately. They count as input tokens on every request, so a lengthy system prompt reused across thousands of daily calls turns cache investment into an easy calculation: pay the cache-write premium once, then collect the cache-read discount on every subsequent call that reuses the same block.

Finout's pricing guide flags a specific calculation error that follows from skipping this step: using a blended rate that doesn't separate cache-hit tokens from fresh input tokens. That error understates costs on output-heavy workloads and overstates them on cache-heavy workloads, and in both directions it hides what's actually driving the number on the invoice. Cache hit rate needs to sit alongside input and output token volume as a tracked metric in its own right. Without it, any cost forecast is built on an assumption rather than a measurement.

Context-length tiers and reasoning token billing as the two most overlooked cost traps

Two billing mechanisms are most likely to blow up a budget on a workload that otherwise looks ordinary: pricing tiers tied to context length, and reasoning tokens billed invisibly. Both are dangerous because neither appears in a count of the words in a team's prompts and responses.

Context-length tiers kick in once a prompt crosses a provider-set token threshold, and the rate on the other side of that threshold isn't a small bump. Google Gemini 3.1 Pro doubles its input price once a prompt exceeds a set context length, and xAI Grok 4.5 follows the same pattern. A model with a large advertised context window at a flat headline rate can look equivalent to a model with a smaller window at the same rate, but if the workload actually fills that context, the larger-window model costs more per request once both are filled at the same per-token rate. The hidden cost multiplier here is wasted context: paying to send a large block of retrieved documents when only a small slice of it is relevant to the answer means paying several times over for tokens that never contribute anything to the output. An optimized retrieval step that trims context before it reaches the model closes that gap directly.

Reasoning tokens create a different kind of surprise: they are invisible, billed in full without showing up in either the prompt or the response. Some reasoning models generate internal "thinking" tokens that get billed in full even though the caller only sees the final answer. Finout's guide is explicit that cost estimates for agentic and reasoning workloads need to add this overhead in directly, or the estimate comes in systematically low. Some amount of reasoning overhead is mandatory on that model, not optional. It has to be budgeted for rather than avoided.

Once all four billing dimensions, output multipliers, cache mechanics, context tiers, and reasoning overhead, are accounted for, comparing providers requires running the actual numbers from production traffic against each provider's specific rules rather than reading a price table.

The provider landscape under blended rates

Looking at the providers active as of late September 2026, the spread between blended rates is wide enough that ranking providers by effective cost for a given workload produces a different order than ranking them by input price alone. The budget tier makes this most visible, since several models here post output prices under $2 per million tokens, yet the gap between the cheapest and most expensive options in that tier is still large once caching and workload shape are factored in.

DeepSeek V4-Flash held the cheapest listed spot on the comparison table as of August 2026, at $0.14 per million input tokens and $0.28 per million output tokens under its pre-August-16 flat rate, with peak rates rising after that date and off-peak rates running lower. Its cache hit rate of $0.0028 per million tokens is the figure that made it such an efficient choice for any workload built around reused prompts. By September 10, 2026, V4.1 Flash had superseded it, carrying higher peak rates than the original V4-Flash tier.

GPT-6 Luna shipped September 22, 2026, priced at half the rate of its predecessor, GPT-5.6 Luna. A price cut of that size, cutting the rate in half in a single generation, changes which workloads make financial sense on that model overnight, and it illustrates how quickly a blended-rate comparison can go stale if it isn't recalculated against current pricing.

Mistral Small 4 and Mistral Large 3 are priced at the budget and mid-range tiers respectively. That spread gives teams a way to match workload complexity to price tier within a single provider's model family, rather than switching providers entirely to find a cheaper option for lighter tasks.

GLM-5.3-Flash, GLM-5.3-FlashX, and GLM-4.7-Flash make up this tier, the last of which is free. A free tier changes the calculation entirely for workloads that can tolerate its specific limitations, since it removes the token-cost line item from the budget altogether for whatever volume that tier supports.

Gemini 3.8 Flash carries introductory pricing through December 31, 2026, after which rates rise.

Put together, this landscape confirms the pattern that runs through every section above. The cheapest-looking model on a headline table is not reliably the cheapest model for a specific workload, because the real cost depends on the output multiplier, the cache hit rate, whether the prompt crosses a context-length threshold, and whether the model carries mandatory reasoning overhead. A team that pulls its own input and output token counts from production logs, checks them against each provider's current rate card, and runs the blended-rate formula against its actual cache hit rate will land on a number that the input-price table could never have predicted. That number, not the one in the marketing page, decides whether a product's unit economics hold up at scale.

Sources

  1. LLM Token Cost by Model: 2026 Pricing Data and Optimization Tips
  2. LLM API Providers (2026): 12 APIs Compared by Price per 1M Tokens, Rate Limits, and Context
  3. LLM Token Pricing Comparison 2026 — Cost Per Million Tokens
  4. LLM API Pricing 2026: GPT vs Claude vs Gemini vs DeepSeek

More in AI Spend Management