This benchmark tests planning, tool use, and command-line agentic capability — a "brutal test" of a model's ability to operate autonomously in a terminal environment Y. Reported scores:
| Model | Terminal-Bench 2.1 Score |
|---|---|
| GPT-5.6 Sol Ultra | 91.9% BI |
| GPT-5.6 Sol | 88.8% B |
| Claude Mythos 5 | 88.0% I |
| GPT-5.5 | 83.4% BI |
Sol Ultra's 91.9% represents a significant jump — roughly 4 percentage points ahead of Anthropic's best model and more than 8 points ahead of GPT-5.5 BI.
OpenAI rates all three GPT-5.6 models as "High" in both cybersecurity and biological risk, but none cross the "Critical" threshold A.
| Model | Input ($/1M tokens) | Output ($/1M tokens) | Role |
|---|---|---|---|
| GPT-5.6 Sol | $5.00 | $30.00 | Flagship — complex reasoning, coding, cybersecurity, long-horizon agentic work |
| GPT-5.6 Terra | $2.50 | $15.00 | Balanced — matches GPT-5.5 performance at half the cost |
| GPT-5.6 Luna | $1.00 | $6.00 | Volume — fastest and most cost-efficient for classification, extraction, routing |
OpenAI also introduced prompt caching for GPT-5.6. Cache writes are billed at 1.25x the model's uncached input rate, while cache reads receive a 90% discount on cached-input tokens. GPT-5.6 supports explicit cache breakpoints and a 30-minute minimum cache life H.
Primary reason: Direct U.S. government intervention. The Trump administration explicitly requested that OpenAI restrict access before release, citing national security concerns TAK. White House officials personally cleared the roughly 20 approved partner organizations one by one AI.
The administration considered GPT-5.6 to have "Mythos-like" capability — referring to Anthropic's Claude Mythos, whose predecessor (Claude Fable 5) was forcibly suspended worldwide by the Commerce Department on June 12, 2026, under export control rules A. The GPT-5.6 restriction stems from the same policy framework.
OpenAI stated publicly that this kind of government gatekeeping "keeps the best tools from users, developers, enterprises, cyber defenders, and global partners who need them" A.
OpenAI says it "plans to expand availability as soon as possible" but has not announced a general-availability date H. Sam Altman has stated he expects broader access "within weeks" A. The company's system card notes: "We believe in broad access, and we plan to make GPT-5.6 Sol, Terra, and Luna generally available in the coming weeks" D.
For now, the most powerful AI model ever released by OpenAI is only available to those Washington approves.