Gartner's latest forecast carries a finding that surprises most finance teams: AI is getting cheaper per token and more expensive per task at the same time. In an August 2026 release, Gartner predicted that inference cost per agentic workflow will rise more than fivefold through 2028, even as unit prices keep falling.
That is the budget problem many Hong Kong enterprises are now meeting in their monthly invoices. The pilot looked affordable. Production does not. This guide defines AI FinOps, explains why AI spend behaves differently from cloud spend, and gives you a four-layer framework you can take into your next budget review.
What is AI FinOps?
AI FinOps is the discipline of making AI spending visible, attributable and accountable. It applies cloud financial management practices to tokens, model calls, GPU capacity and AI subscriptions, so that every dollar of AI cost can be traced to a team, a workflow and a business outcome, and then optimised against that outcome.
The term borrows from cloud FinOps, the practice that emerged a decade ago when organisations realised nobody could explain their AWS or Azure bills. AI has recreated that problem in a faster and more fragmented form.
AI cost now arrives through several doors at once:
--- Model API usage, billed per million input and output tokens, with discounts for cached input.
--- Seat subscriptions for copilots and assistants, billed per user per month.
--- Infrastructure, including GPU capacity, vector databases, embeddings, orchestration and logging.
--- Embedded AI inside SaaS tools, often priced as credits or add-ons that sit in departmental budgets.
According to the FinOps Foundation's State of FinOps 2026 report, 98% of practitioners now actively manage AI costs, up from 63% in 2025 and 31% in 2024. AI cost management is also the skill those teams most want to add in the next 12 months.
Why are AI bills rising when token prices keep falling?
Token prices fall, but token consumption rises faster. Gartner calls this the Inference Paradox: cheaper tokens fund more complex agentic workflows that read, reason, call tools and check themselves. Gartner expects inference cost per agentic workflow to rise more than fivefold through 2028, so unit price cuts alone will not protect the budget.
Unit prices really are falling. Frontier model launches in September 2026 continued the trend, with providers cutting per-token and cache-read prices on new releases. Gartner separately projects that inference on a trillion-parameter model will cost providers more than 90% less by 2030 than in 2025.
Consumption is the other half of the equation. A chatbot reads a question and answers it. An agent plans, calls tools, reads the results, checks its own work and often tries again. Gartner's August 2026 analysis states that routing a task to an agentic reasoning model raises inference cost by at least five times compared with a basic chatbot interaction, and often much more as complexity grows.
Gartner's Will Sommer put the conclusion bluntly: leaders cannot rely on better token economics to rationalise AI costs. For a COO, the practical meaning is simple. A lower price list does not guarantee a lower bill.
Who should own AI cost in an enterprise?
One named executive should own AI cost, with engineering owning the technical levers and finance owning visibility and unit economics. Harness found that 52% of organisations have no clear AI cost owner. Without one, spikes go unexplained, policies go unenforced and nobody can tell the board whether AI spend is paying off.
The Harness 2026 State of AI in FinOps survey of 700 engineering and FinOps leaders across five countries shows what happens without ownership:
--- 72% experienced an unexpected AI cost spike or bill in the past year, and 33% were caught out more than once.
--- Only 20% could identify the reason within hours if spend doubled overnight.
--- Respondents estimated that 26% of all AI spend is wasted.
--- Only 26% have a robust method for measuring the business value of AI spend.
The split that works in practice pairs two roles. Engineering or the platform team owns the model gateway, routing rules and budgets. Finance owns the unified view and the unit economics. A single executive, often the COO or Head of Digital Transformation, answers to the board for both.
What should an AI cost dashboard measure?
An AI cost dashboard should measure cost per business outcome, not total spend. The core metrics are cost per resolved ticket, per processed document or per completed workflow, alongside tokens per task, model mix, cache hit rate and the manual baseline. Total spend alone hides whether AI is getting cheaper or dearer per unit of work.
A token bill that doubles while volume triples is good news. A flat bill with falling volume is bad news. Only unit cost tells you which one you are looking at.
The metrics worth putting in front of management are:
--- Cost per outcome: per resolved enquiry, per processed invoice, per drafted report.
--- Tokens per task: a rising number often signals prompt bloat or runaway agent loops.
--- Model mix: the share of work handled by frontier, mid-tier and small models.
--- Cache hit rate: repeated context that is billed at a fraction of the normal input price.
--- Manual baseline: the fully loaded cost of the process before AI took it over.
What is the four-layer AI FinOps framework?
The four layers are Inform, Attribute, Govern and Optimise. Inform creates one view of all AI spend. Attribute tags each cost to a team and workflow. Govern sets budgets, thresholds and approval rules before spend happens. Optimise routes work to the cheapest model that meets the quality bar. Each layer depends on the one before it.
Layer 1: Inform. Consolidate API invoices, seat licences, cloud GPU charges and SaaS AI add-ons into one view. Most organisations discover AI spend in budgets they did not expect, from marketing tools to HR platforms.
Layer 2: Attribute. Route model calls through a gateway that tags each request with a team, an application and a workflow. Without tags, every conversation about optimisation becomes an argument about whose cost it is.
Layer 3: Govern. Set monthly ceilings per workflow, token thresholds per task, alerts at 70% and 90% of budget, and a rule that agents escalate to a human before they exceed a spend limit. Governance should stop a problem before the invoice, not explain it afterwards.
Layer 4: Optimise. Build a three-tier model policy. Frontier models handle complex reasoning, mid-tier models handle standard work, and small or open models handle high-volume routine tasks. Add prompt caching, shorter context and output limits. Gartner calls this inference tiering and identifies it as critical to protecting margins.
How does AI FinOps work in a Hong Kong enterprise?
In practice, AI FinOps starts with one high-volume workflow, such as claims triage or customer enquiries, and puts a cost-per-outcome number on it. A Hong Kong logistics group or insurer can then compare that number against the manual baseline, set a monthly ceiling, and route routine requests to smaller models while reserving frontier models for exceptions.
Consider a Hong Kong insurer running AI-assisted claims triage. In month one it knows only the total invoice. After tagging, it learns that 80% of spend comes from one agent that re-reads the full policy document on every step. Caching that document and routing simple claims to a smaller model cuts cost per claim sharply, and the saving can be shown against the manual handling baseline.
Or consider a professional services firm where each partner's team bought its own assistant subscriptions. The Inform layer reveals overlapping seats across three vendors. Consolidation, plus a clear owner, turns a scattered cost into a single line the CFO can plan around.
Data governance matters here too. Cost controls sit alongside retention and privacy terms in every vendor contract. If you are negotiating with model providers, our guide to zero data retention for enterprise AI buyers covers the parallel set of questions.
What mistakes do enterprises make with AI cost control?
The most common mistakes are writing policy before building visibility, measuring total spend instead of unit cost, defaulting every task to the most expensive model, and treating cost as a finance problem only. Harness found 73% of organisations have AI cost policies but only 13% have basic visibility, which means most policies cannot actually be enforced.
--- Policy before visibility. A rule you cannot measure is a rule you cannot enforce.
--- Frontier by default. Sending every request to the most capable model is the fastest way to inflate a bill.
--- Rewarding usage. Harness found 57% of engineers say their organisation encourages maximising AI usage regardless of value.
--- Ignoring the hidden stack. Vector databases, embeddings, logging and data transfer rarely appear on the model invoice.
--- Cutting instead of steering. Blanket freezes stall adoption. Unit economics let you invest more where AI pays and less where it does not.
How should you present AI costs to the CFO or board?
Present AI cost as unit economics with a trend line. Show cost per outcome, the manual baseline it replaces, the monthly ceiling and the guardrails that stop runaway spend. Boards respond to a clear owner, a predictable envelope and evidence that each additional dollar buys measurable output, not to token counts or model names.
A one-page AI cost brief for the board should answer five questions: who owns the spend, what the monthly envelope is, what each unit of output costs today, how that compares with the manual baseline, and which guardrails prevent a surprise. If you can answer those five, you are ahead of most peers in the region.
Frame the request the way finance frames any operating investment. Show the current run rate, the forecast under two volume scenarios, and the trigger points at which you will scale up or pull back. A CFO who can see the exit conditions is far more willing to approve the entry.
It also helps to separate experimentation from production in the budget. A small, capped sandbox allowance lets teams test new models and agents without approval friction, while production workflows carry named owners, unit-cost targets and quarterly reviews. That split protects innovation and accountability at the same time, and it gives the board a clear line between learning spend and operating spend.
Conclusion: Make AI Spend Predictable Before You Make It Bigger
AI FinOps is not about spending less on AI. It is about knowing what each dollar buys, so that you can spend more with confidence where the return is proven. The enterprises that scale AI successfully in 2027 will be the ones that treated cost visibility as infrastructure in 2026, not as an afterthought.
We understand AI. We understand you. With UD by your side, AI never feels cold.
Reviewed by the UD enterprise AI team. Sources verified on 28 September 2026.
Not Sure Where Your AI Spend Is Going?
Now that you have the framework, the next step is finding the workflow where it will pay back first. We'll walk you through every step, from an AI readiness assessment to cost baselining, model selection, deployment and performance tracking, backed by 28 years of serving Hong Kong enterprises.