| Billing item |
|---|
| Claude Opus 4.7 public price |
|---|
| Base input tokens | $5 / 1M tokens |
| Output tokens | $25 / 1M tokens |
| 5-minute cache write | $6.25 / 1M tokens |
| 1-hour cache write | $10 / 1M tokens |
| Cache hit / refresh | $0.50 / 1M tokens |
Without prompt caching, the basic cost formula is:
Cost = input_tokens / 1,000,000 × 5
+ output_tokens / 1,000,000 × 25
With prompt caching, split reusable context from new context. The first write of reusable context is charged at the cache-write rate: $6.25/MTok for a 5-minute cache or $10/MTok for a 1-hour cache. Later cache hits or refreshes are charged at $0.50/MTok. New, uncached user input is still billed at the normal input rate, and model responses are still billed at the output rate.
If you analyze a document once and do not ask follow-up questions, the budget is straightforward: the document, system prompt, and user question are input tokens; the model’s answer is output tokens.
Using the public Claude API rates:
| Scenario | Input | Output | Estimated cost |
|---|---|---|---|
| Shorter long-document summary | 100,000 | 5,000 | About $0.625 |
| Mid-to-large document analysis | 300,000 | 8,000 | About $1.70 |
| Very large document analysis | 1,000,000 | 10,000 | About $5.25 |
For example, 300,000 input tokens plus 8,000 output tokens works out as:
300,000 / 1,000,000 × 5 = 1.50
8,000 / 1,000,000 × 25 = 0.20
Total = 1.70 USD
If you are migrating from an older model, do not blindly reuse old token estimates. Anthropic’s pricing documentation says Opus 4.7 uses a new tokenizer, and the token count for fixed text can increase by up to 35%.
So an earlier estimate of 300,000 input tokens may be safer to budget as 405,000 input tokens. With the same 8,000-token output:
405,000 / 1,000,000 × 5 = 2.025
8,000 / 1,000,000 × 25 = 0.20
Total ≈ 2.23 USD
The most common cost mistake in long-document products is billing the same large file as fresh input on every turn. If users will ask multiple questions about the same document, include prompt caching in the budget model from the start.
Assume:
| Approach | Cost components | Estimated cost |
|---|---|---|
| First turn: create 5-minute cache | 300k × $6.25/MTok + 2k × $5/MTok + 2k × $25/MTok | About $1.935 |
| Later turn: cache hit | 300k × $0.50/MTok + 2k × $5/MTok + 2k × $25/MTok | About $0.21 |
| No cache: resend full document each turn | 302k × $5/MTok + 2k × $25/MTok | About $1.56 |
In this example, the first cached request is more expensive than sending the full document once. But by the second turn, caching is already cheaper overall:
No cache, two turns: about 1.56 × 2 = 3.12 USD
5-minute cache, two turns: about 1.935 + 0.21 = 2.145 USD
The budget question is therefore not just document size. It is cache hit rate: will the same context really be reused, will the follow-up arrive before the cache expires, and how much new uncached material is added on each turn?
Long chats follow the same cost logic as long documents. If your app sends a large conversation history back to the model on every turn, input costs accumulate quickly. Stable, reusable context should be evaluated for prompt caching.
Assume:
| Approach | Estimated cost |
|---|---|
| No cache: 200k history + 1k new message + 2k output each turn | About $1.055 / turn |
| Write 200k history to 5-minute cache: first turn | About $1.305 |
| 5-minute cache hit: later turns | About $0.155 / turn |
| Write 200k history to 1-hour cache: first turn | About $2.055 |
| 1-hour cache hit: later turns | About $0.155 / turn |
Choosing between a 5-minute and 1-hour cache is a product-behavior decision, not just a price-table decision:
Batch workloads are common for offline analysis, data labelling, bulk summarisation, and large-scale classification. But unless you have confirmed the exact batch price for your account, contract, or platform endpoint, avoid putting an unverified discount into the official budget.
A conservative baseline is to estimate the job using the public synchronous Claude API price, then revise downward only after the applicable batch price is confirmed.
The same basic formula applies:
Total cost = total input tokens / 1,000,000 × 5
+ total output tokens / 1,000,000 × 25
Example: 10,000 tasks, each with 2,000 input tokens and 500 output tokens.
Total input = 10,000 × 2,000 = 20,000,000 tokens
Total output = 10,000 × 500 = 5,000,000 tokens
Input cost = 20 × 5 = 100 USD
Output cost = 5 × 25 = 125 USD
Total = 225 USD
That $225 estimate does not include any batch discount. If you later confirm a lower batch price, replace the unit prices in the formula with the actual rates.
Also check where you are buying access. If you are not calling Anthropic’s Claude API directly, the bill may differ. Third-party data from CloudPrice lists Opus 4.7 at $5 input and $25 output per MTok for Anthropic/global-style entries, while some AWS Bedrock regional codes are listed at $5.50 input and $27.50 output per MTok. Treat that as a prompt to verify your actual platform billing page, contract, and official documentation.
A token spreadsheet is usually too optimistic before launch. At minimum, build in room for three things:
A practical, non-official budgeting buffer could look like this:
| Stage | Suggested budget multiplier |
|---|---|
| Proof of concept or pilot | Theoretical cost × 1.2 to 1.5 |
| Production with stable traffic | Theoretical cost × 1.35 to 1.6 |
| Migration from an older model to Opus 4.7 with heavy long-context use | Theoretical cost × 1.5 to 1.8 |
These multipliers are not Anthropic quotes. They are conservative planning ranges. Once the system is live, replace assumptions with real token logs, cache hit rates, and invoice data.
Without prompt caching:
Monthly cost ≈ daily requests × 30
× (average input tokens / 1,000,000 × 5
+ average output tokens / 1,000,000 × 25)
With prompt caching, separate the cost buckets:
Monthly cost ≈ normal input cost
+ cache write cost
+ cache hit / refresh cost
+ output cost
Before implementation, define at least these variables:
| Variable | Example |
|---|---|
| Average input tokens per request | 300,000 |
| Average output tokens per request | 8,000 |
| Daily request count | 1,000 |
| Cache write tokens | 300,000 per document |
| Cache hit tokens | 300,000 per hit |
| Cache hit rate | 60% |
| Tokenizer migration buffer | Up to × 1.35 |
| Operating buffer | For example, × 1.35 to 1.6 |
For one-off long-document analysis, estimate with $5/MTok input plus $25/MTok output.
For repeated Q&A over the same large document, or long chats that carry a large history, model prompt caching before you finalise the architecture. In the 300,000-token document example above, a cached second turn costs about $0.21, compared with about $1.56 when the full document is resent each time.
For batch workloads, start with the synchronous public Claude API price unless your actual batch, cloud-platform, or contract pricing has been confirmed. And if you are moving older workloads to Opus 4.7, apply the possible 35% tokenizer increase first, then add an operating buffer. That will usually be closer to the invoice than a budget based on the headline token price alone.