The challenges cut across four main areas:
No pre-established baselines. Most AI deployments begin tracking metrics only after launch, making any value attribution impossible . A 2026 ModelOp survey found that more than two-thirds of enterprises still rely on estimates rather than measured financial results to assess return on investment
.
Vanity metrics dominate. Many companies track active users, session counts, or queries per day — metrics that reveal nothing about whether AI is producing actual business value . The metric of 2025 was "users." The metric of 2026 needs to be "auditable outcomes"
.
The production-to-ROI gap. Gartner's 2026 CIO Report, surveying 11,000 CIOs, found that 59% of AI initiatives never reach production, and 71% of CIOs struggle to prioritize use cases that will deliver measurable outcomes .
Spend explosion without visibility. AI costs are distributed across multiple providers, models, and departments with no unified accounting framework . Forrester research found that enterprises are postponing 25% of planned AI spend to 2027 as financial scrutiny intensifies
.
OpenAI — 'Useful Intelligence per Dollar'
OpenAI CFO Sarah Friar introduced a four-question scorecard asking: Does the AI work? Is it cost-efficient? Is the value it creates growing faster than its cost? This framework shifts the conversation from raw token consumption to outcome-adjusted unit economics .
Anthropic — 'Economic Primitives'
Anthropic released a framework built on "economic primitives" that measures task complexity, autonomy level, and time saved against cost per task — rather than generic usage . Their internal enterprise guidance emphasizes success metrics tied to performance thresholds, accuracy, speed improvements, and cost efficiencies
.
The Linux Foundation's Tokenomics Foundation
Launched August 4, 2026, with roughly 30 founding members including JPMorgan Chase, Accenture, IBM, SAP, BNY, and Booking.com, this vendor-neutral standards body aims to do for AI token costs what the FinOps Foundation did for cloud spend . Its first draft output, "Big-T Notation," proposes standardizing cost-per-API-call (not cost-per-token) so the metric maps directly to one unit of work performed
. The Foundation is also defining cost-to-serve metrics that currently exist in no standardized form
.
Elisity — 'Bionic Head Count'
The cybersecurity firm created a metric that divides total AI spending by average staff cost, then compares the combined output of human and AI "workers" against revenue. This provides CFOs with a direct labor-productivity analog for AI investment .
Industry-wide efforts
The FinOps Foundation has published AI-specific frameworks recommending unit cost metrics — cost per query, cost per user, cost per workflow — to make token spend legible to business stakeholders . Oracle, Jellyfish, and others now offer token economics dashboards that map consumption to business outcomes at the workflow level
. BCG proposes a formal "Return on AI" (RoAI) ratio: Economic return / (Cost of human intelligence + Cost of tokens)
.
Across all these frameworks, a clear pattern is emerging. The most effective approaches share a common structure:
No single standard has yet won broad adoption. But the direction is clear: a growing consensus is forming around outcome-linked unit economics — cost per completed task, cost per ticket resolved, bionic head count — rather than raw token counts. For CFOs and CIOs looking for a starting point, the advice from OpenAI's Sarah Friar remains the most practical: "The best place to begin is with one workflow. Define what 'done' means and measure that outcome in the system where the work happens" .