Alibaba Open-Source Qwen-Drive-1.0: The World’s First Vision Language Base Model for Autonomous Driving
2026-09-06 17:15Models🔥 49.3 heat score
2sources
1days unfolding
49.3heat score
5mentions
SummaryAI generated
On September 6, 2026, Alibaba Tongyi Lab officially opened-source the Qwen-Drive-1.0-4B model. Built upon Qwen3.5-4B, this is the world’s first visual language base model for autonomous driving. Its core innovation lies in unifying 3D perception and visual question-answering capabilities during pre-training, extending them to motion planning while fully maintaining the original architecture. The model uses a phased training approach, combining driving supervision data with general visual language data, offering both reinforcement learning (RL) and supervised fine-tuning (SFT) for planning. Official evaluations cover 3D perception, driving scenario understanding, general visual language capabilities, as well as open-loop, pseudo-closed-loop, and closed-loop motion planning. All planning samples are derived from public data, and the trajectory format is standardized. After optimization through reinforcement learning, the model has been significantly improved in terms of human preference alignment and closed-loop safety, albeit at the cost of minor open-loop displacement errors.
On September 6, Alibaba Tongyi Laboratory officially opened-source the Qwen-Drive-1.0-4B autonomous driving visual language model. This model is built based on Qwen3.5-4B and is the world’s first foundational visual language model for autonomous driving. Its innovation lies in unifying 3D perception and visual question-answering pre-training, extending it to motion planning while maintaining the original VLM architecture. The evaluations cover 3D perception, scene understanding, general visual language capabilities, and open-loop/pseudo-closed-loop/closed-loop planning. All planning samples come from public data, with a unified trajectory format, and significant improvement in human preference alignment and closed-loop safety after reinforcement learning.
This week, Alibaba Qwen opened-source Qwen-Drive-1.0-4B, which is the first visual language base model designed for autonomous driving. This model is built upon Qwen3.5-4B, integrating 3D perception and visual question answering during pre-training, and extending to motion planning while fully maintaining the original architecture. The model uses a phased training approach, combining driving supervision with general visual language data, and offers both SFT and RL planning experts. Official evaluations cover 3D perception, driving scenario understanding, general visual language capabilities, as well as open-loop, pseudo-loop, and closed-loop motion planning. Qwen-Drive-1.0 uses only publicly available planning samples to unify the trajectory formats of various datasets. The SFT model is competitive across all benchmarks; after reinforcement learning, the model achieves improved performance in human preference alignment and closed-loop safety at the cost of slight open-loop displacement errors.