OpenAI is reportedly testing a limited enterprise model in which customers pay when an AI agent completes an agreed task, rather than for every token used. The model addresses a core agent economics problem: a developer reportedly spent $1.3 million on OpenAI tokens in 30 days while running 100 agents, even though t...
Research answer

Create a landscape editorial hero image for this Studio Global article: What does OpenAI’s reported pilot of outcome-based pricing for select large enterprise customers involve—including how it differs from token. Article summary: OpenAI is reportedly piloting outcome-based pricing with a small set of large enterprises: instead of charging for the volume of model input and output, it would charge when an AI agent completes a pre-agreed business ta. Topic tags: general, general web, news, user generated, documentation. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, water
OpenAI is reportedly allowing a small number of major enterprise customers to pay only when an AI agent completes an agreed business task. The arrangement is described as a limited pilot, not a publicly announced pricing plan, and OpenAI has not disclosed the participating tasks, rates, or contract terms. 5
The experiment matters because autonomous agents can consume far more compute than a conventional chatbot. A token-based bill grows with prompts, responses, retries, and tool calls. An outcome-based bill would instead tie the charge to a result that the customer and vendor define in advance.
OpenAI’s published API and enterprise rate cards price input, cached-input, and output tokens per million tokens. 1
2 Under that model, the customer pays for model usage regardless of whether the agent ultimately completes the task.
Outcome-based pricing changes the billing event. Instead of asking, “How many tokens did the system process?” the contract asks, “Did the agent achieve the agreed result?” A qualifying result might be a successfully resolved case, a completed workflow, or a verified update to a business system—but the specific outcomes in OpenAI’s reported pilot have not been made public. 5
That distinction shifts some execution risk from the buyer to the vendor. Multiple attempts, long reasoning chains, or failed runs could increase OpenAI’s costs without necessarily producing a billable outcome.
Traditional usage-based pricing is relatively straightforward for predictable applications. It becomes harder to forecast when agents can choose their own paths, call tools repeatedly, delegate work, or operate in parallel.
The most striking reported example involved a developer running 100 agents and accumulating $1.3 million in OpenAI token charges over 30 days. 5 The figure is an extreme case, but it illustrates the commercial concern: for autonomous systems, expenditure can scale with attempts and activity rather than with the value of the completed work.
An outcome contract could make budgets easier to model. It could also reduce the risk that a customer pays the full cost of an unsuccessful attempt, assuming the contract treats failure as non-billable. Whether that protection applies, and how exceptions would work, remains unknown in OpenAI’s pilot.
Outcome pricing is easiest to administer when a task has a clear boundary and an observable completion event. Customer support is the leading example because a resolution can often be checked against a ticket status, a customer interaction, or the absence of further human intervention.
Other candidate workflows could include narrowly defined coding, claims, lead-qualification, or back-office actions. These are possibilities rather than disclosed OpenAI use cases. The common requirement is that the parties can specify what counts as success before the agent begins.
A practical contract would likely need to define:
Without those rules, “pay only when the AI works” is an appealing slogan but an ambiguous commercial promise.
OpenAI’s reported experiment follows a broader shift in how software companies are charging for AI agents. The models are related, but they are not identical:
These examples show why “outcome-based” and “consumption-based” should not be treated as interchangeable. Per-action or per-conversation pricing can make usage easier to understand, but it may still charge when the broader business objective is not achieved.
The available Futurum findings point to a fragmented market rather than the end of seat-based pricing.
One Futurum survey of 830 global IT decision-makers found that 43% preferred consumption pricing for generative-AI functionality, while 27% preferred outcome-based pricing. 14 A separate 2H 2026 survey found that buyers preferred per-seat pricing for separately billed AI functionality at 42.3%, compared with 36.6% for consumption pricing and 21.1% for outcomes. For core software, the same survey reported stronger support for consumption-based pricing at 28.9% and outcome-based pricing at 22.2%, while per-seat pricing ranked at 12.6%.
The takeaway is not that enterprises have chosen one universal model. Their preference appears to depend on what the AI is doing. A predictable assistive feature may fit a seat or add-on price. An autonomous service that performs measurable work may be easier to justify through consumption or outcome-linked billing.
The model has a straightforward appeal for both sides. Buyers can connect spending to delivered value, while vendors have a stronger incentive to improve reliability and reduce wasted execution. It may also help finance teams approve agent deployments that would be difficult to budget under open-ended token usage.
But outcome pricing exposes vendors to costs they can no longer automatically pass through to customers. It also creates incentives that must be managed carefully. If an agent is optimized for a narrow metric, it could technically satisfy the contract while damaging the broader customer experience.
Attribution is another obstacle. Consider a workflow that uses several models, external tools, employees, and enterprise systems. If the agent completes only part of the process, or a human makes the final decision, it may be difficult to determine whether OpenAI delivered the billable outcome.
For that reason, outcome pricing is most likely to expand from narrow, verifiable workflows rather than from open-ended knowledge work. Hybrid arrangements—combining seats, usage, credits, and selected outcome fees—may remain more practical for complex deployments. Futurum has described hybrid models as an important middle ground as vendors test how to connect AI pricing with customer value. 14
If the reported arrangement expands, it would mark a meaningful change in enterprise AI economics: vendors would increasingly sell completed work, not just access to models or consumption of tokens.
For now, the evidence supports a more limited conclusion. OpenAI is reportedly testing outcome-based pricing with selected large customers, but the company has not published a general plan or disclosed enough contractual detail to judge its rates, scope, or commercial success. 5 The experiment is best understood as an early test of whether AI-agent vendors can make “pay for results” precise enough for enterprise procurement.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
OpenAI is reportedly testing a limited enterprise model in which customers pay when an AI agent completes an agreed task, rather than for every token used.
OpenAI is reportedly testing a limited enterprise model in which customers pay when an AI agent completes an agreed task, rather than for every token used. The model addresses a core agent economics problem: a developer reportedly spent $1.3 million on OpenAI tokens in 30 days while running 100 agents, even though token billing measures activity rather than business resu...
Outcome pricing works best when success is narrow and auditable, such as a verified resolution; complex workflows still create difficult questions about quality, attribution, exceptions, and disputes.