Claude Opus 4.7 is a clear example. Anthropic says the model’s new tokenizer may use roughly 1x to 1.35x as many tokens when processing text compared with previous models—up to about 35% more, varying by content. Anthropic also says /v1/messages/count_tokens will return a different token count for Claude Opus 4.7 than it did for Claude Opus 4.6.
If the same prompt is split into more input tokens, and the relevant input-token price is unchanged, the input portion of the request will cost more. But “up to 35% more tokens” is not the same as “every prompt costs 35% more.” Anthropic gives a range of roughly 1x–1.35x, and explicitly says the result varies by content.
It is also not the same as saying your whole bill rises by 35%. Anthropic’s pricing separates Base Input Tokens, Cache Writes, Cache Hits and Output Tokens; OpenAI and Gemini also publish their own API pricing pages. So the final cost depends on the full request: input tokens, output tokens, cache behavior, the selected model and the pricing fields that apply.
Tokens are not words. OpenAI’s tiktoken guidance shows that token counts depend on the specific encoding used for a model, while Gemini’s documentation says input and output are tokenized, including text and image inputs.
That means word counts, character counts and rough “tokens per word” estimates are useful only for back-of-the-envelope budgeting. For billing-grade comparisons, you need the token count returned for the exact model you plan to use. Anthropic’s note that Opus 4.7 and Opus 4.6 return different counts for the same input is precisely why tokenizer changes matter.
| Common claim | More accurate reading |
|---|---|
| “Opus 4.7 makes every prompt 35% more expensive.” | Too broad. Anthropic says roughly 1x–1.35x tokens, up to about 35% more, and that the increase varies by content. |
| “The same text can be counted as more tokens.” | Correct. Anthropic says Opus 4.7’s new tokenizer may use more tokens and that token counts differ from Opus 4.6. |
| “Tokenizer changes only affect context limits, not cost.” | Incomplete. API pricing commonly uses input, output and cache-related token fields, so a token-count change can affect cost calculations. |
| “The safest approach is to test with the official counter.” | Correct. OpenAI documents input token counting and tiktoken; Gemini documents count_tokens; Anthropic points to /v1/messages/count_tokens for this Opus 4.7 comparison. |
For the input side only, the basic calculation is:
Extra input cost ≈ (new tokenizer input tokens − old tokenizer input tokens) × input-token unit price
That formula is deliberately narrow. It does not include output tokens, cache writes, cache hits or other pricing fields. Anthropic lists those categories separately, and OpenAI and Gemini maintain their own pricing references for their APIs.
A production request may include system instructions, retrieved context, tool results, files, images or other structured input. Gemini says all input and output are tokenized, including text and image inputs; OpenAI’s token-counting guide also shows input counting for mixed text-and-image requests.
Use the counter that matches the model you are evaluating. OpenAI provides responses.input_tokens.count and tiktoken guidance, Gemini provides count_tokens, and Anthropic says /v1/messages/count_tokens will show different results for Opus 4.7 versus Opus 4.6.
Do not test only one short prompt. Because Anthropic says the Opus 4.7 token increase varies by content, compare the payloads that actually matter to your product: high-volume prompts, long-context requests, expensive workflows and common user patterns.
First measure the old and new input token counts. Then apply the relevant model’s official pricing, and only after that add output, cache and other fields back into your cost model. Anthropic, OpenAI and Gemini each provide pricing documentation for this step.
If the token delta is small, you may only need to update budgets and monitoring. If high-volume payloads become materially more expensive, consider trimming repeated prompt text, shortening context, improving cache strategy or revisiting per-request cost assumptions. The point is not to panic at the “35%” headline; it is to quantify the change with official counters and official pricing.
Claude Opus 4.7’s tokenizer can make the same text consume more tokens. Anthropic’s own documentation says the new tokenizer may use roughly 1x to 1.35x as many tokens as previous models when processing text, up to about 35% more, with the increase varying by content.
But the practical question is narrower: how many extra input tokens do your real payloads produce, what happens to output length, how do cache fields apply and which pricing table governs the request? Run the official token counter first, then apply the official pricing. That is the reliable way to tell whether your prompts will actually cost more.