Gateway & Ground

Hidden Platform Fees in LLM Gateway Pricing Models

Gateway pricing hides platform fees that compound as AI usage scales across enterprises.

Contributing Editor · · 10 min read
Cover illustration for “Hidden Platform Fees in LLM Gateway Pricing Models”
AI Spend Management · September 30, 2026 · 10 min read · 2,236 words

72% of organizations say they plan to increase AI budgets again in 2026. Gartner expects 40% of enterprise applications to have task-specific AI agents built in by the end of 2026, up from under 5% in 2025 LLM Gateway. Each of those agents fires off multiple model calls per task, so the volume running through a gateway doesn't grow in a straight line, it compounds LLM Gateway.

That's the shift that turns gateway pricing from a footnote into a budget line. A fee structure that felt harmless when a team was running a prototype becomes real overhead once every application, and every agent inside every application, routes through the same layer LLM Gateway. Most teams didn't choose their gateway carefully to begin with. They started with whatever came bundled with their cloud provider, a default that was never built for token-based rate limits, multi-provider routing, or the concurrency production AI workloads actually demand.

So here's the situation a lot of teams are quietly sitting in: they're already paying for a gateway, or about to sign up for one, and the pricing page in front of them almost certainly isn't showing the full number. Gateway pricing is built in layers, and only one of those layers tends to get top billing, not because anyone's hiding it maliciously. A Menlo Ventures market update cited in Maxim research shows enterprise LLM API spend climbing from $3.5 billion to $8.4 billion in roughly six months, framing why even small percentage fees matter enormously at this scale.

The two-layer fee structure every gateway uses

Every gateway charges through two separate mechanisms, and they behave nothing alike. The first is a token markup: a percentage tacked onto the provider's published per-token rate. A 10% markup on a $5-per-million-token model turns into $5.50, and that markup applies to every single token sent through the gateway, forever, with no ceiling LiteLLM Pricing 2026: Open-Source & Enterprise Cost Breakdown. The second is a platform or credit fee: a flat percentage taken when credits are purchased, or a flat monthly subscription. It's bounded, unlike the token markup, which applies to every token sent through the gateway, forever, with no ceiling. It happens once per funding event, not once per token.

That difference sounds small on paper. But that difference is not actually small. A token markup scales with consumption, so as usage grows, the tax on every marginal token grows right along with it. A platform fee doesn't behave that way, because it's charged against how much money moves through the account, not against how many tokens get processed. One of these fees punishes growth. The other one doesn't.

Pricing pages rarely spell this distinction out. The platform fee tends to live somewhere else: buried in the payment flow, tucked into an FAQ, mentioned once and never again. That's how a gateway can advertise "zero markup" and still charge a meaningful amount of money. Zero markup on tokens and zero fees overall are two completely different claims, and conflating them is the single easiest way to misjudge what a gateway actually costs.

Diagram: How the Two Gateway Fee Types Behave as Volume Grows. Visualizes: Illustrate the fundamental difference between a token markup and a platform fee as token volume scales up.

What major gateways charge in 2026, line by line

The headline shift in 2026 is this: none of the major gateways mark up tokens anymore. The competition has moved entirely into platform fees, bring-your-own-key (BYOK) terms, and self-hosting options. Here's what that looks like gateway by gateway, on published terms.

LLM Gateway charges no token markup, a 5% platform fee on credits, and drops to 0% if a team brings its own provider keys. It's also free to self-host under an AGPLv3 license LLM Gateway.

OpenRouter charges no token markup either, but takes a 5.5% fee on card credit purchases, with an $0.80 minimum charge per transaction LLM Gateway LLM Gateway LiteLLM Pricing 2026: Open-Source & Enterprise Cost Breakdown. On a $10 top-up, that minimum works out to 8%, which matters a lot for anyone buying credits in small amounts LLM Gateway LLM Gateway LiteLLM Pricing 2026: Open-Source & Enterprise Cost Breakdown. BYOK is free up to $25,000 of list-price inference per month, after which a 5% fee on list-price cost kicks in LLM Gateway LLM Gateway LiteLLM Pricing 2026: Open-Source & Enterprise Cost Breakdown. OpenRouter isn't self-hostable LLM Gateway LLM Gateway LiteLLM Pricing 2026: Open-Source & Enterprise Cost Breakdown.

Vercel AI Gateway skips the token markup too, running on pay-as-you-go credits plus standard payment-processing fees, with BYOK free on its paid tier. It isn't self-hostable. Cloudflare AI Gateway is free on direct billing, and only charges 5% if a team opts into its Unified Billing path; BYOK is available, but self-hosting isn't. Eden AI runs a flat 5.5% on credits with BYOK support, also not self-hostable LLM Gateway.

Portkey offers a free tier and a Production plan priced at $49 a month, which includes 100,000 logs, with overage billed at $9 per additional 100,000 logs beyond that LiteLLM Pricing 2026: Open-Source & Enterprise Cost Breakdown. It supports BYOK and partial self-hosting, and it was acquired by Palo Alto Networks in May 2026, which matters for teams that care about vendor independence LiteLLM Pricing 2026: Open-Source & Enterprise Cost Breakdown.

LiteLLM sits apart from the rest: the open-source software itself is free, and there's no token markup, but a team pays for its own infrastructure. Enterprise support tiers run about $250 a month at the basic level and roughly $30,000 a year at the premium level, and it's fully self-hostable.

None of that is the full bill, either. Vercel, OpenRouter, and Cloudflare's Unified Billing path all pass through card-processing costs, and on a small top-up those processing fees can end up dominating the effective rate a team actually pays. Portkey's $9-per-100k-logs overage sounds trivial until a team running high request volume watches it stack up fast against a monthly baseline that looked cheap at signup. And at least one gateway charges $0.01 per million tokens just for stored request retention, a fee that's easy to miss entirely if the only thing read is the headline number LLM Gateway.

The same model, run through different gateways, can cost noticeably different amounts, even when neither one marks up tokens. A Requesty analysis found certain flex-tier models and deepseek-chat cost roughly half as much through one gateway compared to another, and the gap traced back to which provider pricing tier each gateway had access to, not to any markup at all.

The real cost of "free": what self-hosting adds back in

"Free and open source" is the correct description of LiteLLM's software license LiteLLM Pricing 2026: Open-Source & Enterprise Cost Breakdown. It is not the correct description of what it costs to run that software in production.

TrueFoundry's 2026 analysis breaks down what self-hosting actually adds back in. Initial deployment, standing up Kubernetes clusters, load balancers, CI/CD pipelines, and monitoring, typically eats 2 to 4 weeks of senior DevOps time before anything goes live. After that, ongoing maintenance, security patches, dependency updates, scaling adjustments, general troubleshooting, runs another 10 to 20 hours a month. At a senior DevOps engineer's salary, 20 hours a month of that work works out to roughly $1,730 in labor cost, every single month. When the proxy goes down at 2am, there's no vendor SLA to absorb it. The on-call engineer handles it, and that's a cost too, even if it never appears on a line item. Hidden Platform Fees in LLM Gateway Pricing Models.

There's a technical ceiling worth knowing about as well. LiteLLM runs on Python, and that runtime puts a practical limit on how much sustained concurrency it can handle. Real production deployments generally need PostgreSQL, Redis, and worker-process recycling layered on top, and each of those is its own operational surface to maintain and monitor.

None of this means self-hosting is a bad idea. For a team with strong in-house DevOps that needs full control over its infrastructure and can absorb the total cost of ownership, it's the right call. For a team where the engineering time cost ends up higher than the platform fee it was meant to replace, it isn't, and that math tends to flip a lot faster than most teams expect once volume starts climbing. Infrastructure itself (servers, databases, load balancing) typically runs $200–$500/month for moderate production traffic, before the labor.

Diagram: Self-Hosting's Hidden Monthly Cost vs. a Managed Platform Fee. Visualizes: Show the real cost breakdown of self-hosting an LLM gateway versus paying a managed platform fee.

Fee compounding as token volume scales (the arithmetic that pricing pages omit)

Volume is the variable that changes everything, and it's moving fast. A fee structure that looked harmless at low volume can turn into a serious cost once that volume shows up.

That spread means routing decisions and gateway fees interact multiplicatively, not separately. A team paying a 5% platform fee on credits funding frontier-model spend is paying 5% of a much bigger number than a team routing those same tasks to budget models. The fee percentage might be identical. The dollar amount behind it isn't.

Waste makes this worse. Audits of production AI applications routinely find that 40 to 60 percent of token spend is waste, capability paid for but never actually used, Maxim's research on LLM cost reduction shows. A platform fee charged on credits that fund wasted tokens is, functionally, a fee on the waste itself. And that waste isn't always slow and steady. Particula's research found a single misbehaving prompt can burn through five figures of LLM spend in a week before anyone even notices. Without real-time spend controls sitting at the gateway layer, fees pile up on runaway usage just as readily as they do on usage a team actually intended.

The broader picture makes the gap even more concerning. CloudZero data puts average enterprise monthly AI spend at $85,521 in 2025, yet only 34% of companies had mature cost management processes in place. Most organizations, in other words, are paying gateway fees on top of base spend that's already out of control. The gateway fee alone was never going to be the biggest number on the bill. But it's the number that compounds quietly on top of every other cost driver already at work, and it deserves more scrutiny than a single line on a pricing page usually gets. Enterprise AI token consumption has increased 13x since January 2025, Elvex research shows, meaning fee structures chosen at low volume become expensive at high volume.

Evaluation criteria for a gateway's pricing model

Three levers move total cost more than anything else, ranked in order of impact.

BYOK matters most for teams that already hold provider accounts or have negotiated volume discounts. Where a gateway supports it, bringing your own keys eliminates the platform fee entirely, which makes it the single biggest lever available to high-volume teams. Self-hosting comes next: it removes platform fees too, but shifts that cost into infrastructure and DevOps labor, and it's only worth it when the engineering total cost of ownership comes in under the fee it replaces (the roughly $1,730-a-month labor figure is a reasonable baseline to run that math against). The token never sent costs nothing. Built-in response caching makes repeat requests free, and routing simple requests to cheaper models shrinks the base spend that every fee gets calculated on. Anthropic's Claude, for instance, offers 90% savings on cached prompt tokens compared to its standard input rate.

Before signing anything, a few questions should be put directly to a vendor. Is the token rate shown the provider's actual published rate, or something marked up? By 2026 the answer is almost always passthrough, but it's worth confirming rather than assuming. What exactly is the platform or credit fee, and does it apply to BYOK usage, or only to gateway-managed credits? And if there's a BYOK threshold, what happens above it, does a per-request fee quietly kick back in?

Routing deserves its own line of scrutiny too. Left alone, most teams default to the highest-capability model for every task, even when the task doesn't need it. A gateway that enforces policy-based routing, sending simple requests to cheap models and saving frontier models for the work that actually requires them, shrinks the base spend that fees get applied to.

None of this replaces the need for hard spend limits. Per-key, per-team, and per-project budget caps enforced at the gateway layer stop runaway spend before it piles up, and billing transparency only means something if it comes paired with enforcement, not just a report at the end of the month. The choice between managed and self-hosted isn't a question with one right answer. Managed gateways with transparent flat fees offer predictability, SLA-backed uptime, and no DevOps burden. Self-hosted setups offer full infrastructure control and no software fees, in exchange for engineering time. Which one wins depends entirely on a team's DevOps capacity and the volume at which that math starts to flip. The true cost of any gateway is the provider's token rate, plus the platform fee, plus whatever it costs operationally to run the thing, minus whatever gets clawed back through caching, routing, and cutting waste. Any pricing page that shows only one of those pieces is showing an incomplete picture, and it's worth treating it as exactly that. Teams default to the highest-capability model for every task, but a gateway that enforces policy-based routing (sending simple requests to cheap models and reserving frontier models for tasks that need them) reduces the base spend that platform fees are applied to, and the 4,500x price spread between the cheapest and most expensive models makes routing selection a bigger cost lever than the platform fee itself.

Sources

  1. LiteLLM Pricing 2026: Open-Source & Enterprise Cost Breakdown
  2. 8 Best AI Gateways in 2026 (Compared) | LLM Gateway
  3. AI Gateway Pricing: Fees & Markups Compared (2026) | LLM Gateway
  4. LLM API Pricing — Provider Rates, Free BYOK | LLM Gateway
  5. Cheapest LLM API Prices Compared (2026): Provider by Provider Cost Guide | Requesty

More in AI Spend Management