Gateway & Ground

AI Cost Forecasting and Budgeting for Growing Engineering Teams

Token-based AI billing breaks traditional forecasting—measure cost per task, not per token.

Contributing Editor · · 11 min read
Cover illustration for “AI Cost Forecasting and Budgeting for Growing Engineering Teams”
AI Spend Management · September 30, 2026 · 11 min read · 2,430 words

Forecasting AI spend fails for a structural reason, not a data quality problem: the mental model teams bring to it is wrong. Someone who has forecast cloud infrastructure with confidence for years sits down to project next quarter's AI spend, and the numbers won't hold still, because the mental model itself is wrong. That's not a personal failure. It's what happens when a team applies subscription or cloud billing logic to a cost structure that doesn't work that way at all.

Cloud billing, for all its complexity, is bounded: you provision compute hours or GB-months, and even usage spikes are tethered to capacity someone actually turned on. Subscription billing is even simpler: a fixed price per seat, known months in advance. Both produce cost curves you can draw on a whiteboard and defend in a budget meeting.

Token-based API billing shares none of that discipline. A single request might use 50 tokens or 50,000, depending on prompt length, model tier, and whether an agent chained calls together. There's no provisioned ceiling to bound it. The seat-based instinct, count users and multiply by a rate, breaks down once AI shifts from chatbots to agents, which chain multiple calls per task, so cost stops tracking how many people are logged in.

This isn't a fringe use case teams can wait out. According to BCG's AI Radar, CEOs have already put a significant share of new AI investment behind agentic architecture this year, so the newest dollars flow into the hardest-to-predict part of the system. The FinOps world is living this shift in real time. In 2025, roughly a third of FinOps practitioners had any responsibility for AI spend. By 2026, almost all of them do. A discipline built over a decade to govern predictable cloud costs has been handed a consumption model with no shared playbook.

The three variables that drive token spend (and why only one of them shows up on invoices)

Diagram: The Three Variables That Drive Token Spend. Visualizes: Visualize three distinct variables that drive AI token spend, showing that only one of them appears on an invoice.

Invoices only show the last one clearly. The other two hide inside engineering decisions nobody reviews for cost before shipping.

Model tier sets the price per token, and the spread is large: Claude Haiku 4.5 costs a fraction of Claude Opus 4.7 for the same tokens, and GPT-4o-mini undercuts GPT-4o by a wide margin. Route the same task to a different tier, and the bill changes by orders of magnitude before anyone touches token count. That makes model selection a budget decision first, a quality decision second.

Token volume per request is shaped by prompt architecture: context stuffed in, system prompt length, examples attached, output verbosity. None of that appears in a finance review. A pull request reveals it instead.

Call frequency is the one variable that resembles old-world API billing, at least on paper. In agentic workflows, frequency ties to how the workflow is built: a single user action can trigger many chained model calls.

Uber found this out the hard way. After Claude Code rolled out to its engineering org in December 2025, the company burned through its entire 2026 AI coding tools budget by April. Microsoft pulled Claude Code licenses from its Experiences & Devices division after per-engineer monthly costs climbed similarly. These are organizations with mature engineering cultures, yet the tooling still outran budget within months.

Layer on Jevons paradox and the problem compounds rather than resolves itself over time. Token prices keep falling year over year, but cheaper tokens make more use cases economically viable, so total enterprise AI spend keeps climbing even as the per-token rate drops. Gartner projects inference costs falling sharply through 2030, yet agentic workflows burn so many more tokens per task that enterprise bills rise anyway. Cheaper unit economics do not mean a cheaper total bill. Anyone forecasting AI spend on the assumption that falling prices will flatten the curve is forecasting the wrong direction.

What a practical forecasting unit looks like: cost per task, not cost per token

The unit that actually produces a usable forecast is cost per task, or cost per completed workflow. Raw spend without a denominator tells a team little useful, and tells finance something actively misleading.

Picture a workflow whose monthly token bill doubles. Alarming, on its face. But if that workflow handled triple the volume of work in the same period, it actually got cheaper per unit of value delivered. Finance staring at a bigger invoice with no context reaches the wrong conclusion, and the ensuing panic tends to produce blunt across-the-board cuts instead of targeted fixes.

The comparison that matters is cost of the AI workflow against cost of whatever process it replaced: cost per resolved support ticket, cost per document processed, cost per code review completed. Those are the denominators that turn a token bill into something a business can actually reason about.

FinOps practice offers a pattern to borrow directly: inform, optimize, operate. Inform means real visibility into spend by provider, model, team, and feature. Optimize means cutting waste via routing, caching, and tighter context windows. Operate means setting budgets against task-level unit economics and reviewing them on a schedule.

Building the forecast itself doesn't require anything exotic. Three inputs per workflow: expected task volume, average tokens per task including both prompt and completion, and the price of the model tier doing the work. It's not a formula to memorize so much as a habit of thought: never look at spend without asking what it bought.

DoorDash built accountability around this exact idea rather than trying to predict everything in advance. Every developer gets a high monthly token limit; crossing it requires explaining why and committing to an efficiency plan for the next month. Cost becomes an engineering discipline owned by the team doing the work, not a veto finance issues after the fact, because accountability is one engineer's monthly usage, not an untraceable company-wide number.

Diagram: FinOps for AI: Inform → Optimize → Operate. Visualizes: Visualize the three-phase FinOps cycle as applied to AI spend: Inform (real visibility into spend by provider, model, team, and feature), Optimize (cutting waste via routing, caching…

How workflow architecture shapes token consumption before a single request is sent

Most of the token bill gets decided before a single request leaves the building. Context window design, agent topology, and how a task gets broken into steps: these architectural choices set the ceiling on spend long before any runtime dial gets turned.

Context management is the most direct lever, and the most commonly ignored one. Every token stuffed into a prompt costs money on every single call, forever, until someone rewrites it. A bloated system prompt, a full conversation history dragged along out of habit, a document chunk far bigger than needed: it all adds up invisibly, call after call. Trimming context isn't a nice-to-have cleanup task. It's a design decision with a direct line to the bill.

Agent topology multiplies the effect. A single-turn query costs one call. A multi-step agent chain costs that same call, times however many steps the chain takes to finish the job. Gartner projects the share of enterprise applications wired to task-specific AI agents will grow sharply by end of 2026, up from a small fraction the year before, with each agent firing multiple calls per task rather than one. That's not a marginal shift in how workloads run. It's a structural change in how many calls get made per unit of actual work.

Task decomposition is where architecture and cost meet most directly. Wishroll routed sub-tasks to different model tiers by design, not as a bolted-on afterthought, cutting inference cost sharply while scaling to a large user base. That's the difference between treating routing as a structural decision made at design time versus a runtime patch applied under pressure once the bill arrives.

Different routing philosophies produce different cost profiles for the same workload. Rule-based routing assigns tasks by type (cheap models for classification, stronger ones for generation). Cost-based routing picks the cheapest model clearing a latency bar. Complexity-based routing scores the input and routes accordingly. None of these is universally correct, but each one changes the expected model tier for an entire class of requests. Each one changes the forecast.

Before estimating what anything will cost, map the architecture first. Which tasks route to which tiers, how many steps each workflow takes, what actually lands in context on every call. Skip that mapping, and any token forecast produced afterward is a guess dressed up as a number.

Why model routing is the highest-leverage cost control available after architecture is set

Once architecture is fixed, model routing becomes the single most powerful lever left for controlling cost in real time. And it belongs nowhere near application code.

The price gap between tiers is the reason routing works at all. A flagship model can cost more than an order of magnitude more than a capable smaller model from the same provider, for tasks where the smaller model performs just as well. Routing decisions built around that gap have cut production bills dramatically with no visible quality drop, and at enterprise scale the difference between routing well and not at all has reached hundreds of thousands of dollars a month. That's not a rounding error. That's the difference between a budget that holds and one that blows through its ceiling by April, the way Uber's did.

Routing strategy comes in a few recognizable shapes. Rule-based assigns tasks by category deterministically; cost-based sends requests to the cheapest model clearing a latency threshold; complexity-based scores input first, then decides. Each optimizes for something slightly different, and picking one is itself an architectural choice with cost consequences that ripple forward.

Now consider where that routing logic actually lives, in most organizations. It's buried in application code, service by service: one team writes it well, another copies it badly, a third never gets to it because nobody flagged it mattered. The logic scatters, drifts out of sync, and stays invisible to anyone but its author. Every new service reinvents it or skips it, leaving an org with several uncoordinated routing philosophies running in parallel.

Routing needs a single home that applies the same policy to every request, regardless of which team or service sent it. That home is the gateway layer sitting between application code and the model providers themselves. Putting routing there turns it into infrastructure, because a single home applies the same policy to every request instead of leaving it to per-team judgment calls. It also enables real-time budget enforcement, checking spend limits before a call goes out rather than after the invoice arrives, which scattered app code never could.

What the gateway layer must do to make forecasting and cost control real

A gateway that just proxies requests through to providers doesn't make spend forecasting operational. What makes forecasting operational is a specific set of capabilities sitting at that same layer: real-time spend attribution, hierarchical budget enforcement, token-based rate limiting, semantic caching, and reliable fallback routing.

Spend attribution must be granular from the start, broken out by team, project, virtual key, model, and provider, not reconstructed by hand from a late invoice. TrueFoundry's evaluation framework treats such attribution as a baseline requirement. Skip it, and the inform phase of the FinOps cycle never happens; there's nothing to inform anyone with.

Budget enforcement is what turns visibility into an actual guardrail. The gateway needs the ability to block or redirect a request the moment a threshold gets crossed, at the org level, the team level, or down to a single virtual key. Zuplo's 2026 guide notes a retry-loop bug can burn a monthly budget in minutes without a stop, and hard gateway enforcement provides that stop. That's the operate phase, made concrete instead of aspirational.

Rate limiting has to be counted in tokens, not requests. A request-count limit is nearly useless, since a single request's token consumption can vary by orders of magnitude. Counting requests instead of tokens is like metering electricity by counting how many times someone flips a switch.

Semantic caching detects when differently phrased queries mean the same thing, and returns a cached answer instead of paying for a fresh model call. It cuts cost and latency together, working across every application routed through the gateway.

Fallback and failback routing matter because provider outages and rate-limit errors are among the leading causes of downtime in production AI systems. A gateway that reroutes cleanly during an outage keeps the application running, but it also has to track what that rerouted traffic actually cost. An outage that quietly shifts traffic to a pricier backup model wrecks a forecast if the substitution isn't tracked.

None of this counts if the overhead is too heavy to survive an agentic chain, where a handful of milliseconds gets multiplied across every step. TrueFoundry's framework sets a sub-10ms p95 overhead bar as a baseline requirement. Independent benchmarks put the fastest open-source gateways at negligible overhead even at 5,000 requests per second on modest hardware, while more widely deployed options add a few to tens of milliseconds under heavy concurrency, a gap that matters in agentic chains.

What to look for when evaluating gateway options across the open-source and managed spectrum

Gateway options today span self-hosted open-source projects, integrations built into specific platforms, and fully managed aggregators sitting somewhere in between. There's no single right answer here. The right choice depends on which capability gaps matter most for a given team's forecasting needs and its governance requirements.

Latency measured under real production load affects agentic workloads especially, since overhead compounds across chained calls rather than showing up once.

Governance depth is the next filter: virtual keys, hierarchical budget controls, per-team rate limits, and role-based access control need to exist together, not as separate bolted-on add-ons.

Deployment model decides whether a gateway is even usable for a given workload in the first place. Self-hosted, managed, private cloud, or air-gapped: regulated workloads often can't route sensitive prompts through a third party at all, regardless of that party's routing quality.

Among open-source options evaluated in current comparisons, one Go-based gateway stands out on raw benchmarks, running at negligible overhead even at 5,000 requests per second on a t3.xlarge instance while supporting over 25 providers. It ships with virtual-key governance, semantic caching, automatic fallback, a full client-and-server MCP gateway, and per-key tool filtering, deployable self-hosted or in a private cloud, with drop-in integration via a base URL change. For high-traffic, regulated, mission-critical workloads, that combination of low overhead and deep governance is exactly the profile worth testing first.

Whatever gets chosen, the evaluation should start from the forecasting framework. Ask what the tool does for spend attribution, for budget enforcement, for token-based rate limiting, for routing, for fallback. A gateway that scores well on all five gives a team the infrastructure to run the cost-per-task model, instead of just a nicer dashboard for the same blind spending.

Sources

  1. 6 Best LLM Gateways in 2026
  2. Top 5 LLM Gateways for Production in 2026 (A Deep, Practical Comparison) - DEV Community
  3. Best API Gateways for AI and LLM Workloads (2026): Evaluative - Zuplo
  4. Comparing Open-Source LLM Gateways in 2026 to Run Enterprise AI at Scale - DEV Community

More in AI Spend Management