OpenAI’s GPT 6 Astra was the only model in DrivingBench’s four model test to finish a real car cone course: 134.7 meters in 5 minutes 22 seconds on its second attempt. The successful attempt used 6.6 million tokens and cost $7.74 in model inference, or about $92.47 per mile at that rate.[34] The result shows a gener...
Published byEdited with GPT-6 SolImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: According to the DrivingBench report, how did OpenAI’s GPT-6View Astra become the first commercial large language model to steer a real car. Article summary: DrivingBench’s result was a narrow but notable proof of concept: OpenAI’s GPT-6 Astra—not “GPT-6View Astra” in the available sources—was the only tested model to complete a real-car cone course, succeeding on its second . Topic tags: general, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fak
A general-purpose AI model completed a course in a real car, but the conditions matter as much as the finish. In DrivingBench, OpenAI’s GPT-6 Astra—not “GPT-6View Astra”—was the only one of four tested models to navigate the full parking-lot cone course. It succeeded after a failed first attempt, at walking pace and with a human in the driver’s seat ready to brake.14
15
5
Researchers Tobias Gessler, Simon Mahns and Aditya Ramabadran built DrivingBench to test frontier models on a physical driving task rather than in a simulator. They used a Toyota Corolla connected through Comma 4 hardware and an openpilot-based control harness. Each model could make up to three attempts in a continuous conversation, giving it a chance to use feedback from earlier runs.11
36
The model viewed the course through two cameras and controlled the car one tool command at a time. It could call observe for camera views and telemetry, set_motion for steering and speed, or stop_now to stop. The course was in an empty parking lot, and a human sat in the driver’s seat as the braking backup. DrivingBench published recordings and command traces for the attempts.5
11
Astra reached 49% of the course on its first attempt. After the researchers asked it to reflect on its mistakes, it completed the 134.7-meter course on attempt two in 5 minutes 22 seconds. That works out to an average of about 0.94 mph, including stops; it does not mean the car’s speed never exceeded one mph.15
34
None of the other tested models finished. Claude Fable 5.1 reached 45% at best, while Grok 4.6 and GPT-5.6 Sol made much less progress. Reports of the test describe the first corner as a common point of failure.4
40 Astra’s finish is therefore a first within this reported comparison, not evidence that every commercial model or every driving scenario has been tested.
11
14
DrivingBench’s trace for Astra’s successful attempt records 6.6 million tokens and $7.74 in model-inference cost for the 134.7-meter trip. Extrapolating that bill over a mile gives approximately $92.47 per mile. That is a calculation from one short run, not a measured cost for sustained driving, and it does not include the car, hardware or human supervision.34
11
Cost figures elsewhere are not always labeled the same way. A secondary comparison puts Astra’s two runs combined at roughly $9.75, while DrivingBench’s attempt-level trace assigns $7.74 to the successful second run. The figures should not be treated as competing prices for the same attempt.8
34
DrivingBench demonstrates that a general-purpose model can use visual observations and tool calls to complete one real-world driving task after feedback. That is more concrete than success in a simulation, but it is a much narrower claim than road readiness.36
15
The test used an empty lot, a short fixed route, walking-pace driving and a human braking backup. It did not demonstrate reliable behavior in traffic, emergency handling without an operator, or economical long-distance operation.5
11
34 The published result is best understood as a proof of capability under carefully controlled conditions—not a replacement for a tested autonomous-driving system.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
OpenAI’s GPT 6 Astra was the only model in DrivingBench’s four model test to finish a real car cone course: 134.7 meters in 5 minutes 22 seconds on its second attempt.
OpenAI’s GPT 6 Astra was the only model in DrivingBench’s four model test to finish a real car cone course: 134.7 meters in 5 minutes 22 seconds on its second attempt. The successful attempt used 6.6 million tokens and cost $7.74 in model inference, or about $92.47 per mile at that rate.[34]
The result shows a general purpose model can control a car through camera observations and tools in one constrained test; it does not establish safe, affordable autonomous driving.[5][11]