But a whole-repo prompt has to satisfy three constraints:
/v1/messages/count_tokens returns different results for Opus 4.7 than for Opus 4.6.| Question | What the sources say | Practical meaning |
|---|---|---|
| How large is the context window? | Opus 4.7 supports a 1M-token context window. | Very large inputs are possible, but there is still a hard limit. |
| How much can it output? | Opus 4.7 supports up to 128k output tokens. | Long reports, big patches and bulk analysis require output planning. |
| Did tokenization change? | The new tokenizer may use roughly 1x to 1.35x as many tokens as previous models, and token-counting results differ from Opus 4.6. | Do not rely on old token estimates or rough character counts. |
| Is it aimed at repo-scale work? | Anthropic positions Opus 4.7 for complex agentic workflows, long-running work and reliable operation in larger codebases. | It is more appropriate for large code tasks, but not an unconditional guarantee. |
| Is it stable for long tasks? | Anthropic says Opus 4.7 handles complex, long-running tasks with rigor and consistency. | The official framing is positive, but production teams should still validate it on their own repos. |
A repository is rarely a clean, compact document. A useful codebase review may need README files, configuration, tests, dependency manifests, CI errors, stack traces, search results and recent diffs. Those can matter more than simply dumping every file into the prompt.
Some files also add volume without much value for the task: generated code, build artifacts, minified bundles, lockfile noise, vendor directories, cache folders, large logs and duplicated documentation. Including them can crowd out the source files, tests and diagnostics that actually help the model reason about the problem.
That matters even more with Opus 4.7 because the tokenizer can use up to about 35% more tokens than previous models for some content. A repo that looked comfortably under the limit using an older estimate may be much tighter when counted with the Opus 4.7 tokenizer.
It is reasonable to be optimistic, but not absolute.
Anthropic’s product page frames Opus 4.7 as suitable for complex agentic workflows, long-running work and larger codebases. Its announcement also says the model can handle complex, long-running tasks with rigor and consistency.
Those claims support a cautious conclusion: Opus 4.7 is officially positioned as a stronger fit for long-context, multi-step and large-codebase work. They do not prove that every oversized repository, every agent loop or every one-shot prompt will complete reliably.
For production uses such as security review, CI/CD auto-fixes, large refactors or long-running autonomous coding workflows, the safer approach is to test against your own repository, your own test suite and the failure cases you actually care about.
List the main directories, languages, entry points, tests, configuration files and recent changes. Decide what the model needs for the job before loading files into context.
For many tasks, you should exclude build outputs, generated files, vendored dependencies, caches, giant logs and repeated documentation unless they are directly relevant.
Do not estimate from Opus 4.6, another model or a simple characters-to-tokens rule. Anthropic says the Opus 4.7 tokenizer may use roughly 1x to 1.35x as many tokens as previous models, and its token-counting endpoint returns different numbers from Opus 4.6.
Even if the input technically fits, the task may still suffer if there is no practical room for reasoning, summaries, patches or follow-up. If you expect a detailed report or large code changes, keep the prompt lean and reserve output space.
For a big repo, a staged workflow is often more reliable than one massive prompt. A sensible pattern is:
That style fits Anthropic’s positioning of Opus 4.7 for complex agentic workflows and larger codebases without assuming that every file must be present from the start.
For repo analysis, require the answer to include what was read, what was not read, key assumptions, unresolved risks and suggested tests. That does not guarantee correctness, but it helps prevent a common mistake: treating partial context as if it were full codebase understanding.
Claude Opus 4.7 really does support a 1M-token context window and up to 128k output tokens. Anthropic also presents it as a model built for long-running, agentic and larger-codebase work.
So, can it read an entire repo at once? Sometimes. If the repository and all task context fit comfortably, one-pass analysis can make sense. If the repo is large, noisy or requires a substantial report or patch, the better answer is still to filter first, work in stages and validate the result with real tests.