Three AI engineers have pushed general-purpose language models beyond screens and into the controls of a real car, according to a report by WIRED’s Will Knight published on October 7. The experiment is notable less as a new self-driving system than as evidence that models built to handle text, images, code and other digital tasks may be acquiring a limited ability to reason about motion and space in the physical world.

Aditya Ramabadran, Simon Mahns and Tobias Gessler, engineers at the startup Axiom, connected a chat interface to a server receiving video from cameras mounted on a 2024 Toyota Corolla. The system could also operate the car’s power steering. With a safety driver keeping a foot ready above the brake, OpenAI’s GPT-6 Astra slowly guided the vehicle toward an In-N-Out takeout window and completed the lunchtime trip, WIRED reported.

Abstract camera frames and steering vectors guide a car through a controlled parking course.
The researchers connected camera video and steering controls to models that were not built specifically for autonomous driving.

The setup differs from conventional autonomous-driving programs, which use software developed and trained specifically for driving. Here, the engineers asked a general-purpose model to issue steering commands from visual input without first coaching it for their course. That distinction makes the result unusual, but it does not make the experiment equivalent to a driverless service: a human safety operator remained in place, and the broader testing took place on a simple parking-lot course.

The project began after the three engineers noticed that leading models could construct complex three-dimensional simulations and wondered whether that ability would transfer to a vehicle. The models initially refused to control the car, returning warnings that they could interpret road images but should not issue motion commands. The researchers found that careful prompting could move them past those refusals. They told WIRED that the models’ driving ability may be an emergent byproduct of training on multimodal and 3D data, rather than evidence that the systems were explicitly trained to drive.

Performance on the team’s DrivingBench test was sharply uneven. The benchmark asks a model to steer around a basic course laid out in a parking lot. GPT-6 Astra was the only model to finish, and it did so very slowly. Anthropic’s Claude Fable 5.1 covered 45 percent of the course, while Grok reached 11 percent, according to WIRED. Those results underline how far the systems remain from passing anything resembling a real driving test.

Three abstract motion trails reach different points on a closed driving loop.
DrivingBench exposed a wide performance gap: one model finished the course slowly, while the other two stopped far short.

The researchers also said the models appeared to adjust to the unfamiliar controls as the trials progressed. Ramabadran described them as learning from mistakes within the active session and improving their navigation. That observation comes from the experimenters rather than an independent safety evaluation, so it is better read as a clue about possible in-context adaptation than as proof that the models can reliably learn physical control.

The work sits within a larger effort to measure physical reasoning. WIRED reported that Elorian AI and Scale AI recently developed Humanity’s Sixth Sense, a benchmark designed to test whether models understand physical scenes rather than merely identify what is visible. Elorian chief executive Andrew Dai argued that stronger visual reasoning could support applications ranging from robots in homes to systems that interpret activity in restaurants, while also calling robotics an essential test of the capability.

For now, the Corolla experiment offers a vivid but narrow result: a general-purpose model can sometimes connect camera input to useful steering actions under close human supervision. It also exposes the risk of treating an impressive one-off demonstration as mature autonomy. The benchmark failures, slow completion and need for a safety driver all point in the same direction. Physical reasoning may be improving, but the report provides no basis for putting today’s general-purpose models in charge of ordinary road travel.