The catch is that not every token is priced the same. If your application uses prompt caching, you need to separate standard input, output, cache writes and cache reads instead of multiplying one “total tokens” number by a blended rate.
In the table below, MTok means 1,000,000 tokens. Anthropic’s pricing page separates Base Input Tokens, Cache Writes, Cache Hits and Output Tokens, so your own cost model should do the same.
| Billing item | Price | How to read it |
|---|---|---|
| Base input tokens | $5 / MTok | Standard input sent to the model, excluding tokens billed as cache writes or cache reads. |
| Output tokens | $25 / MTok | Tokens generated by Claude in the response. |
| Prompt cache write, 5-minute TTL | $6.25 / MTok | Charged when reusable prompt content is written to the cache with a 5-minute time to live. |
| Prompt cache write, 1-hour TTL | $10 / MTok | Charged when reusable prompt content is written to the cache with a 1-hour time to live. |
| Cache read / hit | $0.50 / MTok | Charged when previously cached content is read from the prompt cache. |
For a basic request with no prompt caching, the formula is:
cost = input_tokens / 1,000,000 × 5 + output_tokens / 1,000,000 × 25
For example, a request with 200,000 input tokens and 20,000 output tokens would cost $1.00 for input plus $0.50 for output, or $1.50 before any platform-specific fees.
Once prompt caching is involved, calculate each billing category separately:
cost = base_input_tokens / 1,000,000 × 5 + output_tokens / 1,000,000 × 25 + cache_write_5m_tokens / 1,000,000 × 6.25 + cache_write_1h_tokens / 1,000,000 × 10 + cache_read_input_tokens / 1,000,000 × 0.50
If you only use one cache TTL, keep only the matching cache-write term. Anthropic’s streaming examples show usage fields such as input_tokens, output_tokens, cache_creation_input_tokens and cache_read_input_tokens, while the pricing page bills cache writes and cache hits separately.
Do not estimate Claude API cost from word count, character count or file size. Anthropic provides the /v1/messages/count_tokens endpoint so developers can count tokens before sending a message to Claude. The endpoint accepts a structured input similar to a Messages API request, including system prompts, tools, images and PDFs, and returns the total input tokens. Anthropic says all active models support token counting.
The safest workflow is to send the exact payload you plan to use — system prompt, messages, tools and attached content included — to count_tokens first. That gives you a pre-request estimate for input-token cost and makes it easier to enforce product-level budget limits, warnings or request caps.
After the request completes, log the API response’s usage data instead of trying to infer output tokens from the returned text. Anthropic’s Messages API examples include usage fields such as input_tokens and output_tokens; streaming examples also show cache-related fields including cache_creation_input_tokens and cache_read_input_tokens.
There is one important streaming gotcha: Anthropic’s streaming documentation says token counts in message_delta.usage are cumulative, not the incremental tokens for that single event. If you add every delta as though it were a new slice of usage, you will overcount.
Per-request logs are useful for real-time controls, but they are not the whole finance workflow. For month-end reporting, workspace chargebacks or historical cost analysis, use Anthropic’s Usage & Cost Admin API. Anthropic describes it as programmatic, granular access to historical API usage and cost data for an organization, with breakdowns by dimensions such as model, workspace and service tier.
A practical setup is to log response usage inside your application for immediate controls, then use the Usage & Cost Admin API as the source for formal reconciliation and reporting.
Opus 4.7 introduces a new tokenizer. Anthropic says the tokenizer may use roughly 1x to 1.35x as many tokens when processing text compared with previous models — up to about 35% more, depending on content — and that /v1/messages/count_tokens can return a different token count for Claude Opus 4.7 than for Claude Opus 4.6.
That means the same $5 input / $25 output sticker price does not guarantee the same real-world bill after an upgrade. If you are moving from Opus 4.6 or an earlier Claude model, rerun count_tokens on high-traffic prompts, long-context prompts, tool-definition-heavy payloads and your most expensive workflows before updating alerts, rate limits or customer-facing budgets.
claude-opus-4-7 when using the Claude API./v1/messages/count_tokens on representative production-like payloads.message_delta.usage is cumulative and should not be summed event by event.Bottom line: Claude Opus 4.7’s base API price is easy to remember — $5/MTok for input and $25/MTok for output — but accurate budgeting depends on token counting before the call, usage logging after the call, prompt-cache accounting and a fresh look at token budgets under the new tokenizer.