A group of tech workers operating under the name DrivingBench successfully connected several large language models (LLMs) to a Toyota Corolla. The LLMs, including GPT-6 Astra, Claude Fable 5.1, Grok 4.6, and GPT-5.6 Sol, were given control over the car's steering, accelerator, and brakes. The experiment took place in a public parking lot where a small cone course was set up for navigation.
The primary objective of DrivingBench was to assess whether general-purpose, untrained LLMs could operate a real car. Unlike purpose-built autonomous driving systems that rely on extensive training data, these LLMs ran on a laptop. While Grok, Sol, and Fable only managed to drive short distances, GPT-6 Astra eventually learned to complete the entire cone course after troubleshooting.
Aditya Ramabadran, Tobias Gessler, and Simon Mahns, the individuals behind DrivingBench, conceived the idea after observing LLMs demonstrate advanced robotic and spatial reasoning capabilities. They emphasized that the experiment was not intended to prove the immediate practicality of using chatbots for everyday driving. Instead, their goal was to demonstrate the potential of these models in real-world tasks and establish a new benchmark for LLM performance in physical environments.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
A group of tech workers named DrivingBench connected general-purpose large language models (LLMs) to a Toyota Corolla, enabling one model, GPT-6 Astra, to navigate a cone course in a parking lot. This experiment aimed to test the real-world driving capabilities of untrained, off-the-shelf AI models and establish a new benchmark for LLM performance in physical environments.