Token cost attribution ties AI inference spend to teams by joining two billing systems that share no common key: Kubernetes cost tools that track node-hours, and model API billing that tracks tokens. The gap is closed by propagating a workload label (team, cost-center) from Kubernetes admission through the Downward API, into a gateway like LiteLLM, and finally into a PromQL join against vLLM Prometheus metrics and node cost. Self-hosted and API-based inference require different instrumentation. A five-step process covers label conventions, Kyverno admission enforcement, Downward API exposure, LiteLLM virtual keys, and a token-fraction-weighted PromQL query for GPU cost splitting. The FOCUS 1.4 spec (ratified June 2026) standardizes billing records across providers but doesn't yet cover the namespace-to-token join; FOCUS 1.5 is expected to address this. Average GPU utilization sits at 5% across major clouds, making attribution a prerequisite for cost reduction via GPU sharing, MIG partitioning, rightsizing, and spot/cross-cloud placement.

•21m read time•From cast.ai
Post cover image
Table of contents
Key TakeawaysWhy token cost does not show up in your Kubernetes cost toolWhat you are actually trying to attributeHow to build the joinThe FOCUS specification and where AI cost fitsShowback and chargeback for AI teamsReducing the bill once you can see itConclusionFrequently Asked Questions

Questions this post answers

How do I attribute AI token costs to specific teams running on Kubernetes?

Apply a label convention with app.kubernetes.io/team and app.kubernetes.io/cost-center on every inference workload, enforce it at admission with Kyverno or OPA/Gatekeeper, and route inference calls through a gateway like LiteLLM, which maps virtual API keys to teams and records spend per team. For external APIs, attach the team identifier via OpenAI's metadata object or Anthropic's workspace IDs, then join gateway spend records to Kubernetes cost data on the team label. Engineers wiring up cost attribution for LLM workloads can find deeper technical breakdowns like this one on daily.dev.

Why doesn't my Kubernetes cost tool show token spend from OpenAI or Anthropic?

Kubernetes cost tools like OpenCost or Kubecost read cloud billing and cluster metadata, producing allocation by namespace and label, with no concept of tokens. Model API billing is keyed by API key, project, or workspace and carries no namespace or pod label. The two systems share no join key by default, so closing the gap requires propagating a shared identifier from the workload through to the API request. Anyone debugging why FinOps dashboards miss AI spend can track this kind of infrastructure explainer on daily.dev.

What does the FOCUS 1.4 FinOps specification cover for AI and Kubernetes cost, and what's missing?

FOCUS 1.4, ratified June 2026, standardizes the Invoice Detail dataset so billing records from AWS, Azure, GCP, and AI API providers like OpenAI use common field names, plus a Billing Period dataset for reconciliation. It does not standardize the join between Kubernetes namespaces and token spend, and self-hosted inference attribution is explicitly out of scope; FOCUS 1.5, in development, is expected to add per-model cost segmentation and input/output token distinction. Teams tracking FinOps standard updates for AI cost reporting can follow specification changes like this on daily.dev.

Share this post