TL;DR
Get garage and car supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Researchers tested general-purpose AI agents using Comma vision hardware and OpenPilot to steer a Toyota Corolla around a short cone course. GPT-6 Astra was the only agent to finish, on its second attempt; eight of 11 runs failed before completing more than 11% of the route, and Grok’s first run ended after two commands.
Researchers testing general-purpose AI models behind the wheel found that GPT-6 Astra was the only agent to complete a short cone-marked course, while eight of 11 runs failed before covering more than 11% of the route. The DrivingBench experiment used a Toyota Corolla equipped with Comma vision hardware and OpenPilot, and its results do not show that ChatGPT or other chatbots are ready to drive on public roads.
DrivingBench was created by Aditya Ramabadran, Simon Mahns and Tobias Gessler to test how general-purpose models handle a basic driving task. Four agents—models from Claude, GPT and Grok—were given the same instructions and control of the Corolla on a short point-to-point course marked by small cones. The models had up to three attempts and were asked to report what they saw and did after each step.
The system issued commands through a wireless connection, with pauses as requests were sent to remote data centers. The runs took place within continuous chat sessions, and DrivingBench published video and conversation records for the trials. Across the 11 runs, eight failed to get beyond 11% of the course, generally ending at or near the first turn.
DrivingBench’s account describes different failures among the agents. Grok misread a gap between boundary cones as a gate and drove off the course in its first attempt, which ended after two commands. A GPT model incorrectly inferred a pattern in cone colors despite the written instructions warning that the cones did not follow that pattern. Other runs did not steer enough to make the bends. GPT-6 Astra finished on attempt two, but followed an uneven path and nearly left the course near the finish. DrivingBench put the cost of that run at $7.74.
A Cone Course Is Not Road Readiness
The results offer a small, practical test of a gap between language-model capability and reliable physical control. An agent can describe what it sees and issue steering commands, yet still misread the course, invent a visual rule or fail to turn far enough. Those errors matter when a system’s actions affect a moving vehicle rather than just a text response.
At the same time, this was a limited experiment, not a public-road safety evaluation. The course was short and set up in a parking lot, and the report does not establish how the models would perform across broader driving conditions. Astra’s finish shows that one run could succeed under the test setup; it does not establish consistent performance. The findings are a reason for caution about treating general-purpose chatbots as drivers, not evidence that they are being used to operate everyday cars.
As an affiliate, we earn on qualifying purchases.
How DrivingBench Set Up the Test
The test differs from purpose-built automated-driving systems. DrivingBench asked general-purpose AI agents to act on camera input and control a vehicle through OpenPilot and Comma hardware. The report frames the exercise as a way to see whether models built for broad tasks could manage a basic driving route, rather than as a comparison of commercial self-driving products.
Each agent received the same command and directions, and could attempt the course up to three times. The researchers’ published session videos and chat records let viewers follow the agents’ reasoning and actions. That record also exposes a constraint of the setup: commands and observations were relayed over a wireless hotspot, with the car stopping regularly while responses were processed. The article does not provide enough detail to assess how much those pauses or other hardware and setup choices shaped the outcomes.
“The test was designed as “a simple test for Claude, GPT, and Grok models to navigate a short point-to-point course around a parking lot with a few curves.””
— The DrivingBench researchers, as described by The Drive
As an affiliate, we earn on qualifying purchases.
Limits of the Reported Results
The reported results cover 11 runs across four agents, but the source does not give a full breakdown of every model version, each run’s exact conditions, or a detailed scoring method beyond course completion and route progress. It also does not say whether the agents were tested in other environments or on longer routes. The small test cannot establish how often an agent might succeed under different conditions.
The report does not provide an independent technical assessment of the hardware, OpenPilot configuration, connection delays or safety controls. It is also unclear how comparable the models’ computing costs were beyond the stated $7.74 for Astra’s successful attempt. No evidence in the source shows that these general-purpose agents are approved, deployed or suitable for driving on public roads.
As an affiliate, we earn on qualifying purchases.
Further Testing Would Be Needed
The published videos and chat records allow readers to inspect the attempts, but the source does not announce a follow-up test or publication date. Further evidence would be needed to judge repeatability: more runs, clearly documented model and hardware configurations, and tests across varied routes and conditions could show whether Astra’s finish was reproducible or an isolated result.
For now, the confirmed takeaway is limited to this course: most reported runs failed early, and one agent completed it once on its second attempt. DrivingBench does not establish road-going capability, so the experiment should not be read as a reason to hand driving control to a chatbot.
Toyota Corolla OpenPilot compatible
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
Did ChatGPT drive a car in the experiment?
The DrivingBench test gave GPT-family agents control of a Toyota Corolla using Comma vision hardware and OpenPilot on a parking-lot course. GPT-6 Astra completed the route on its second attempt. This was a controlled, limited test, not evidence that ChatGPT can safely drive on public roads.
Did Grok complete the course?
According to The Drive’s account of DrivingBench, Grok’s first attempt ended after two commands. It treated a gap between boundary cones as a gate and drove off the course. The supplied source does not give a successful Grok run.
How many runs succeeded?
GPT-6 Astra was the only agent reported to complete the course, doing so on its second attempt. Across 11 runs by four agents, eight did not complete more than 11% of the route.
Does the test show AI is ready for self-driving cars?
No. The experiment tested a small cone course in a parking lot and does not establish dependable performance on public roads or in varied driving conditions. Its results instead show several ways general-purpose agents struggled with visual interpretation and steering in this setup.
Source: rss
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
