THE FUTURELESS

RADAR ·

Alibaba's Qwen-Drive 1.0 combines driving and cockpit tasks in one model

Alibaba's research division has released Qwen-Drive 1.0, a model that combines environmental perception, traffic question answering and route planning in one system. Built on Qwen3.5-4B, released in February, it adds two components: one produces a bird's-eye map from camera images, the other plans the car's path for the next few seconds. When the team trained only the added component, spatial accuracy stayed low; results improved markedly only once the vision-language model itself was trained on spatial tasks. In simulation, the version refined with reinforcement learning cut the rate at which the car veered off the road from 24 to 12 percent, while driving more cautiously and covering less ground. The paper notes the model's stated reasons do not always match its maneuvers, and perception weakens on footage from other camera setups.

Source: The Decoder

← Back to the radar