For general readers, the key phrase is well-specified knowledge work. In plain English, GDPval is about whether a model can produce defined work outputs across a range of professional tasks. It is useful for judging GPT-5.5 as a work-oriented model, but it should not be treated as a single all-purpose scorecard.
| Benchmark or comparison | Reported result | What it measures | How to read it |
|---|---|---|---|
| GDPval | 84.9% | Well-specified knowledge work across 44 occupations | The strongest short benchmark because OpenAI states it directly in its GPT-5.5 announcement. |
| Expert-SWE | 73.1% | Coding tasks; reported as OpenAI’s internal evaluation for tasks with an estimated 20-hour completion time | More relevant to software development than GDPval, but not directly comparable with it. |
| BixBench | 80.5% | Real-world bioinformatics tasks | Relevant for bioinformatics, but in the available source set it is less directly evidenced than OpenAI’s GDPval figure. |
| Artificial Analysis Intelligence Index | No. 1, by 3 points | A third-party model index from Artificial Analysis | Useful for broad model comparison, but it is not a single official OpenAI benchmark. |
Numbers such as 84.9%, 73.1% and 80.5% look easy to rank. That can be misleading.
So the better question is not “Which percentage is highest?” It is “Which benchmark matches the job?” For general workplace knowledge tasks, GDPval is the better reference point. For software engineering, Expert-SWE is closer to the target. For bioinformatics, BixBench is more relevant to the domain.
Artificial Analysis says GPT-5.5 tops its Intelligence Index by three points. It also says OpenAI leads five of its headline evaluations and places second to Gemini 3.1 Pro Preview on three others.
That distinction matters. A No. 1 ranking in an external index does not mean a model wins every individual test. It means GPT-5.5 comes out ahead overall under that third-party index’s methodology.
Some reports cite additional GPT-5.5 figures, including 91.7% in connection with legal AI capabilities and 82.7% in the context of agentic coding. Those numbers may be useful for their specific areas, but they are weaker as a general answer unless the benchmark design, comparison group and measurement target are just as clear.
That is why 84.9% on GDPval remains the cleanest general-purpose benchmark to quote: it is directly stated by OpenAI and tied to a defined evaluation area.
Use the benchmark that matches the context:
The best short benchmark for GPT-5.5 is 84.9% on GDPval. It is official, easy to cite and clearly tied to well-specified knowledge work across 44 occupations. Other scores may matter more for coding, science or specialised professional tasks, but they should always be named with the benchmark they come from.