Agents

OpenAI's Astra Drives a Toyota Corolla to In-N-Out

Three engineers used OpenAI's GPT-6 Astra to drive a Toyota Corolla to a fast-food restaurant, proving that general-purpose AI models are gaining real-world physical reasoning.

WIRED AI16 hrs agoAgents
Image: WIRED AI

Aditya Ramabadran, Simon Mahns, and Tobias Gessler, engineers at the startup Axiom, conducted the experiment using a 2024 Toyota Corolla. They connected OpenAI's GPT-6 Astra to a server linked to windshield-mounted cameras and the vehicle's power steering system. With a safety driver's foot hovering over the brake, the general-purpose model successfully steered the car through an In-N-Out drive-thru. The feat highlights how multimodal models are developing emergent spatial reasoning capabilities without explicit training for physical navigation.

To evaluate these capabilities systematically, the trio created DrivingBench, a test measuring a model's ability to navigate a simple parking lot course. OpenAI's Astra was the only model to complete the course, albeit at a very slow pace. Anthropic's Claude Fable 5.1 completed 45 percent of the track, while SpaceXAI's Grok managed just 11 percent. Initially, these models refused to issue motion commands, but the engineers bypassed these restrictions using careful prompting.

This development aligns with broader industry efforts to teach AI physical reasoning. Startups like Elorian AI, led by former Google DeepMind researcher Andrew Dai, are focusing on this frontier. Elorian and Scale AI recently introduced Humanity's Sixth Sense, a benchmark designed to evaluate how intuitively models understand physical environments. Scale AI research scientist Xingang Guo noted that the benchmark aims to measure what it takes for a model to "understand a scene intuitively," much like a human does.

For AI practitioners and roboticists, these experiments show that scaling up multimodal training data—including images, video, and 3D models—yields unexpected real-world utility. Ramabadran observed that the models demonstrated in-context learning, adjusting to the vehicle's controls and correcting errors in real time. Rather than relying on highly specialized, single-purpose autonomous driving algorithms, developers may soon leverage general-purpose foundation models that can intuitively adapt to physical systems on the fly.

This is our own summary of reporting by WIRED AI

More in Agents