author photo
Gabriel Zessin
Author
,
September 17, 2026
•
15
min reading time

Your LLM bill arrives as a single number. An aggregated value per account, at the end of the month, with no way to tell which team spent it, which application drove the consumption, or which AI agent went into a loop after Friday's deploy. You know how much it cost, but you don't know whose bill it is.

In the FinOps discipline of cloud providers, this problem was solved long ago with tags, where each compute resource carries a cost center label, and allocation happens naturally through the provider itself. In GenAI, this doesn't work, for a structural reason: there is no resource to tag. There is an ephemeral HTTP call, born and gone in seconds, consuming an amount of tokens no one predicted and leaving no persistent object behind.

‍The cost unit of AI FinOps is not the resource: it's the request. And the only place every request passes through in full — with the chosen model, the token count, the caller's identity, and the response latency — is your AI Gateway.

This article shows, in practical terms, how the Sensedia AI Gateway (as the central AI governance layer) can turn this into a complete financial discipline, by measuring, pricing, attributing, and governing your AI spend.

‍The Four Layers of AI FinOps

Before metrics and visual dashboards, it's worth establishing the mental model. AI FinOps rests on four layers, and the order between them is not an arbitrary choice.

Each layer only exists if the previous one exists. Chargeback without your negotiated price properly configured is an accounting fiction: you allocate across teams a cost that is not what your company actually pays. Budget without real-time cost becomes a mere post-mortem alert: you only find out about the budget overrun after it has happened. FinOps initiatives that start from the top layer, with a beautiful allocation report on a foundation that doesn't exist, tend to stall more easily. Sensedia chose to build it bottom-up.

‍

AI FinOps Layer 1: Measurement

Cost must be born in the data plane

The first architectural decision of the Sensedia AI Gateway is also the most decisive one: cost is calculated in the data plane, on every request. There is no scheduled job reading the provider's invoice, nor a batch reconciliation against a billing CSV.

When an LLM request crosses our Gateway, the data plane resolves the provider + model pair, looks up the current price catalog for that pair (including any custom prices registered by the customer itself), and emits the cost in currency (USD) along with the request's own telemetry. Cost becomes an attribute of the request, available the very instant it completes and becomes telemetry.

This may look like a mere implementation detail, but it isn't. It is precisely what makes layers 3 and 4 possible: you cannot block a budget overrun with data that only arrives at month-end closing. Real-time cost-based rate limiting obviously requires real-time cost. And real-time cost requires the calculation to happen in the request path.

‍What the telemetry carries

Cost is emitted as an accumulated metric and as an attribute of the call's trace, always accompanied by the dimensions that allow slicing it for future reports. For example: 

And token consumption is not just a number. It is broken down into four types: input, output, cache read, and cache write, since each has a different price. This breakdown matters because it is what turns cache from a mere performance metric into a financial metric, as we will see in the showback section.

AI FinOps Layer 2: Pricing

‍List price is not your price‍

Most tools on the market calculate cost using the LLM provider's public price. It works as an estimate, but not as a basis for internal allocation. Large enterprises do not pay list price. There are usually volume discounts, enterprise contracts, or special conditions negotiated per model. Companies that opt for open-weight models on their own infrastructure also need a custom way to configure that price.

If the amount you allocate across your teams is not what your company actually pays, chargeback loses legitimacy at the very first questioning. The Sensedia AI Gateway solves this with a layered price catalog.

‍

Related content: How does an AI Gateway facilitate the use of multiple LLMs in enterprises?

‍

Public base with customer override

The base layer comes from a public source of market prices, updated on a refresh cycle. On top of this base, you can set your own rates, and the override is additive: you override the price of GPT-4o without touching the other hundreds of models in the public catalog. Each price in the list displays its origin based on the existing layers: 

‍

The term is Overridden, not Replaced, and this choice reflects the behavior: the default price is preserved, not erased. When you open the details of an overridden price, your rate and the market rate appear side by side (which is, by itself, a negotiation asset). And when you remove an Overridden price, the system reverts to the default price instead of leaving the model orphaned.

‍

Two decisions that matter

There are two behaviors that often go unnoticed and that matter to anyone operating in a regulated industry:

  1. Price changes are not retroactive. Changing, restoring, or removing a price applies to future consumption. Already-computed historical cost is not recalculated, because it is a historical record, captured under the price table in effect at that moment. If it were a recalculable view, one month's report would change the following month, which would disrupt any accounting close.
  1. A model without an active price does not emit telemetry with zero cost. If a model has no price defined in the catalog, its consumption is still fully accounted for in tokens but excluded from the currency cost. This is because, in our view, uncomputed cost is different from zero cost. A system that silently adds zero to the total produces a report that closes and looks correct but is essentially wrong.

What the first two layers produce

‍Showback

With cost properly measured and priced, and the proper dimensions implemented across the views, showback is a mere consequence. The cost dashboards of the Sensedia AI Gateway break down cost by provider, model, and AI agent, with time evolution and a view of the top consumers. What separates showback from a simple "consumption chart" are the details below.

‍Unit economics

Two indicators change the conversation: cost per request and cost per 1,000 tokens. They answer the question the total-cost chart cannot: did my agent get more expensive, or just more used?

These are questions with opposite answers. Total cost growth that keeps the per-request cost stable is a sign of adoption, and that is good news, likely deserving more budget. An increase in per-request cost with stable volume is a sign of possible degradation: prompts bloating, context growing with each iteration, silent retries, someone who switched models without telling anyone.

Without unit economics, both cases appear on the dashboard as the same rising line. Both indicators come from combining the series emitted by layer 1: the period's cost divided by the number of requests, and the period's cost divided by the total tokens in thousands. Neither requires instrumenting the application, because the AI Gateway already holds both halves.

Cache hit-rate as a financial metric

Back to the four token types from layer 1. Since cache reads are priced separately, and at a fraction of the normal input price, the cache hit-rate stops being a performance indicator and becomes a direct lever for cost reduction.

The math is the ratio between tokens read from cache and total input tokens. The higher this number, the lower the cost per request, without requiring any model change. A team that restructures its prompts to maximize the cacheable prefix is not just doing latency optimization: it is also cutting the bill — and now it has the number to prove it.

‍Cost and performance on the same screen

The cost panels live side by side with metrics such as TTFT (Time to First Token), tokens per second, and latency percentiles. Not just for layout convenience, but because every cost decision in AI is a trade-off with quality or latency. Switching to a cheaper model is a two-variable decision. Seeing both in the same place is what allows you to make it with data instead of opinion.

‍

Below, you can check out some of the metrics available with the Sensedia AI Gateway:

ai-gateway-sensedia-metricas1

ai-gateway-sensedia-metricas2

ai-gateway-sensedia-metricas3

‍

‍

‍

AI FinOps Layer 3: Attribution

‍Identity is the contract

The difference between showback and chargeback is not only the report, but also the attribution dimension. Showback shows that your AI platform spent 40 thousand dollars. Chargeback says that 12 thousand belong to the customer support team, 10 thousand to the sales team, and 18 thousand to marketing, booking it into each one's budget. For that to be achieved, you need a stable attribute, present in every request, that says which area the call came from.

‍JWT claims

The Sensedia AI Gateway uses JWT claims as allocation dimensions. The customer can carry the area, project, or team in the JWT token, and the Gateway splits cost metrics according to that configured claim. This was an architectural decision, and it is worth making clear which alternative it rules out.

The most common path in the market is the use of a virtual key: an API key per team, issued by the Gateway, with allocation derived from which key was used. It is simple to start with, and it ends in API key sprawl: a parallel registry of keys to keep in sync with the company's IdP, which goes stale at the first reorganization, squad merger, or person who changes areas. Allocation granularity becomes limited by your discipline in key management.

By tying attribution to the token your IdP already issues, allocation follows the org chart automatically. The correct area arrives in each request because it is naturally already there, not because someone remembered to update a key spreadsheet.

Since everything in architecture is about trade-offs, this choice obviously requires more initial configuration and carries a shared responsibility, because it is up to the Sensedia customer to ensure the correct claims are in the token and configured on the platform. In exchange, it requires far less maintenance in the long run, and the allocation data inherits the identity governance the company already has.

AI FinOps Layer 4: Governance

‍Why rate limiting does not protect the budget

Every traditional API Gateway knows how to limit requests per minute. It is one of the category's oldest policies, but it does not serve an AI budget, and the reason is arithmetic. One thousand requests can cost 2 or 2,000 dollars, depending on the chosen model, the size of the context sent, and how many tokens the response generated. A limit based only on requests is not a budget limit; it is a traffic limit (which may even behave like a budget while the usage pattern does not change). However, we know the AI usage pattern changes every week.

Therefore, only a cost-based limit is truly a budget limit. The Sensedia AI Gateway allows you to define cost ceilings per route with rate limits measured in dollars, not only in calls. When the defined ceiling is reached, the configured enforcement action takes effect.

Why in the AI Gateway, and not in the billing system?

The question is fair: doesn't finance already control the budget? At closing time, yes. The billing system learns about the overrun after it has happened. The Gateway knows at the very moment of the response, for each request, because it is the one that calculates the cost (layer 1) with the price you actually pay (layer 2). AI cost governance needs to happen in the request path, otherwise it becomes mere accounting.

‍Adoption roadmap

If you are just getting started, the layer order we presented is the starting point:

1. Centralize your AI traffic in a Gateway: without a central AI governance layer, nothing else exists. Direct consumption to the LLM provider is a blind spot you carry forever.

2. Turn on measurement and use the default price for a few weeks: having an approximate number today is worth more than having the exact number a quarter from now.

3. Register your negotiated rates: this is where the number stops being just a metric and becomes a defensible basis for allocation.

4. Publish the showback before charging anyone: give teams a grace period to visualize their own consumption. The chargeback conversation gets much easier when no one is caught off guard.

5. Activate budgets on critical routes: start by monitoring consumption, and then start blocking once teams already know their own consumption patterns.

‍

Related content: Learn how the AI Gateway is already being used in the financial industry

‍

A note on reconciliation

It is worth anticipating an objection from every experienced FinOps professional: the cost reported by the AI Gateway will not exactly match the provider's invoice. It won't, and no tool does that. Scenarios caused by retries, aborted streams, rounding differences, and distinct closing windows can always generate some divergence.

The goal of the cost monitored in the Gateway is not to replace the invoice, but to help explain it. The invoice says how much, and the Gateway says whose, in which model, in which application, and why. A fractional-percentage divergence between the two is noise. An aggregated number with no attribution at all is a much bigger management problem.

The final choice is about friction

The three scenarios below change team behavior in different ways:

  • Showback changes behavior through transparency: nobody wants to appear at the top of the list of top consumers (at least not without a defensible reason).
  • Chargeback changes behavior through budget: the cost of AI starts competing with the team's other financial priorities.
  • Budget changes behavior through blocking: it is a protection layer, not the strategy.

Choose the degree of friction your organization can handle. But note that all three will require exactly the same technical foundation we discussed: real cost, calculated in the request path, with the price you actually pay, attributed to an identity that doesn't go stale.

That is the foundation Sensedia delivers, enabling your organization to build its own strategy on top of our platform. The rest is internal politics, and that is your decision.

Want to learn how to enable the FinOps discipline in your business? Talk to our experts now!

‍

Begin your API journey with Sensedia

Hop on our kombi bus and let us guide you on an exciting journey to unleash the full power of APIs and modern integrations.

Blog

Related content

Check out the content produced by our team.

No items found.

Embrace an architecture that is agile, scalable, and integrated

Accelerate the delivery of your digital initiatives through less complex and more efficient APIs, microservices, and Integrations that drive your business forward.