Cursor Composer 2.5: Benchmarks, Pricing, and How It Compares to Claude Opus 4.7 and GPT‑5.5
Cursor’s Composer 2.5, released May 18, 2026, delivers near‑frontier coding performance—79.8% on SWE‑Bench Multilingual and 69.3% on Terminal‑Bench 2.0—while costing about $0.50 per million input tokens and $2.50 per... The model targets long‑running software engineering workflows inside the Cursor IDE, improving su...
Published byEdited with GPT-5.5Images generated with GPT Image 2
Cursor’s Composer 2.5, released May 18, 2026, delivers near‑frontier coding performance—79.8% on SWE‑Bench Multilingual and 69.3% on Terminal‑Bench 2.0—while costing about $0.50 per million input tokens and $2.50 per...
The model targets long‑running software engineering workflows inside the Cursor IDE, improving sustained reasoning, multi‑file editing, and reliability on complex instructions.
Built on Moonshot AI’s Kimi K2.5 checkpoint and heavily trained with synthetic coding tasks and reinforcement learning, Composer 2.5 reflects Cursor’s strategy to reduce reliance on external AI providers.
Cursor Composer 2.5: Benchmarks, Pricing, and How It Stacks Up to Claude Opus 4.7 and GPT‑5.5Cursor’s Composer 2.5 aims to deliver frontier‑level coding performance while dramatically lowering the cost of running AI coding agents.
AI Prompt
Create a landscape editorial hero image for this Studio Global article: Cursor Composer 2.5: Benchmarks, Pricing, and How It Stacks Up to Claude Opus 4.7 and GPT‑5.5. Article summary: Cursor’s Composer 2.5 is an in‑house coding model released May 18, 2026 that scores about 79.8% on SWE‑Bench Multilingual and 69.3% on Terminal‑Bench 2.0—roughly matching Claude Opus 4.7 on some benchmarks while costi.... Topic tags: cursor, ai coding, developer tools, ai models, benchmarks. Reference image context from search candidates: Reference image 1: visual subject "Composer 2.5 matches Opus 4.7 and GPT-5.5 on CursorBench 3.1 but costs less than a dollar per task - compared to up to eleven dollars for the competition. | Image: Cursor" source context "Cursor's Composer 2.5 matches Opus 4.7 and GPT-5.5 benchmarks ..." Reference image 2: visual subject "Composer 2.5 vs Opus | The Results Are Brutal Merv
openai.com
Cursor’s Composer 2.5 is the latest coding‑focused model from Anysphere, the company behind the Cursor IDE. Released on May 18, 2026, the model is designed specifically for AI‑assisted software engineering workflows such as navigating large repositories, editing multiple files, executing terminal commands, and iterating on tests.
The release is notable for two reasons: its competitive benchmark performance against frontier models like Anthropic’s Claude Opus 4.7 and GPT‑5.5, and its significantly lower token pricing, which changes the economics of running long‑lived coding agents inside development environments.
Studio Global AI
Continue your research
This page includes a source-backed answer you can continue inside Studio Global.
What is the short answer to "Cursor Composer 2.5: Benchmarks, Pricing, and How It Compares to Claude Opus 4.7 and GPT‑5.5"?
Cursor’s Composer 2.5, released May 18, 2026, delivers near‑frontier coding performance—79.8% on SWE‑Bench Multilingual and 69.3% on Terminal‑Bench 2.0—while costing about $0.50 per million input tokens and $2.50 per...
What are the key points to validate first?
Cursor’s Composer 2.5, released May 18, 2026, delivers near‑frontier coding performance—79.8% on SWE‑Bench Multilingual and 69.3% on Terminal‑Bench 2.0—while costing about $0.50 per million input tokens and $2.50 per... The model targets long‑running software engineering workflows inside the Cursor IDE, improving sustained reasoning, multi‑file editing, and reliability on complex instructions.
What should I do next in practice?
Built on Moonshot AI’s Kimi K2.5 checkpoint and heavily trained with synthetic coding tasks and reinforcement learning, Composer 2.5 reflects Cursor’s strategy to reduce reliance on external AI providers.
Composer models are optimized for agentic software engineering, meaning the AI is expected to perform multi‑step development workflows rather than just produce short code snippets. Tasks can include planning changes across a codebase, editing multiple files, compiling code, and debugging through iterative test runs.
Compared with earlier versions, Composer 2.5 improves reliability on long‑running tasks, follows complex instructions more consistently, and behaves more cooperatively during collaborative development sessions inside the IDE.
This focus reflects a broader shift in AI coding tools—from simple autocomplete or snippet generation to persistent agents that can execute full development workflows.
Benchmark Performance vs. Opus 4.7 and GPT‑5.5
Cursor reports several benchmark results that place Composer 2.5 within the same general performance tier as leading frontier models.
On SWE‑Bench Multilingual, which measures an AI’s ability to fix real GitHub issues across languages, Composer 2.5 performs roughly at frontier level and slightly ahead of GPT‑5.5 in the reported comparison.
On Terminal‑Bench 2.0, designed to evaluate agent performance in terminal environments, Composer 2.5 performs almost identically to Claude Opus 4.7 but trails GPT‑5.5 by a substantial margin.
Compared with the previous generation, Composer 2.5 shows large improvements—for example rising from 73.7% to 79.8% on SWE‑Bench Multilingual.
Overall, the benchmarks suggest Composer 2.5 is competitive with top models on some software engineering tasks, though it does not consistently outperform them across all agent evaluations.
Why the Pricing Is So Disruptive
The most striking aspect of the release is pricing.
Composer 2.5 is priced at approximately:
$0.50 per million input tokens
$2.50 per million output tokens
A faster variant is available at $3.00 per million input tokens and $15.00 per million output tokens, still competitive with fast tiers from other frontier models.
For comparison, some reports estimate that Claude Opus models can cost around $5 per million input tokens and $25 per million output tokens, meaning the standard Composer tier can be dramatically cheaper—particularly for output tokens.
This matters because agentic coding workflows consume very large numbers of tokens. A single task might involve repository search, planning steps, editing code, compiling, and executing tests, each triggering additional model calls.
Lower token prices allow Cursor to run many more reasoning steps per task without dramatically increasing costs.
What the Model Is Built On
Composer 2.5 builds on Moonshot AI’s Kimi K2.5 open‑weight checkpoint, which Cursor then extends through additional training tailored to software engineering tasks.
Reports about the training approach indicate that the model used:
About 25× more synthetic coding tasks than the previous generation
A training process where roughly 85% of compute was devoted to additional training and reinforcement learning, rather than relying primarily on the base checkpoint.
Synthetic tasks allow the model to repeatedly practice structured development workflows—planning edits, modifying code, running tests, and iterating—helping improve reliability on real engineering problems.
Why This Release Matters for Cursor’s Strategy
Composer 2.5 also reflects a broader strategic shift inside Cursor.
Early versions of the IDE relied heavily on external AI providers such as OpenAI, Anthropic, and Google to power coding features. Developing competitive in‑house models changes that dynamic.
Owning more of the model stack provides several advantages:
Lower inference costs for long‑running coding agents
Reduced dependence on external model providers
Greater control over model behavior within the IDE
This is especially important as competitors like Anthropic’s Claude Code benefit from tight integration between the underlying model and the coding agent itself.
By developing its own Composer models, Cursor is attempting to compete more directly in that integrated model‑plus‑tool category rather than simply routing requests to third‑party AI systems.
The Bottom Line
Composer 2.5 does not clearly dominate the frontier across all benchmarks. GPT‑5.5 still leads in some agent evaluations, and Claude Opus 4.7 remains highly competitive.
What makes the model notable is the combination of near‑frontier coding performance and dramatically lower cost. If Cursor continues improving its in‑house models while maintaining this pricing advantage, it could significantly shift the economics of AI‑assisted software development—especially for long‑running coding agents operating directly inside the IDE.