Antbird released its first native multimodal model, Ling-3.0-flash-VL, introducing a visual feedback closed-loop mechanism
2026-09-09 09:23Models🔥 44.2 heat score
1sources
1days unfolding
44.2heat score
4mentions
SummaryAI generated
On September 9, 2026, Ant Group released Ling-3.0-flash-VL, the first native multimodal large model in its Bailing series. The total number of parameters of this model is 124B, with a single inference activation parameter of 5.5B. It is developed based on Ling-3.0-flash and natively supports input from images, text, and videos, with a context window of 256K Tokens. The model incorporates a visual feedback闭环 mechanism, designed to support tasks such as medical report interpretation, code generation, and GUI automation. In the Artificial Analysis Intelligence Index v4.1.1 evaluation, it scored 4 points higher than the pure text version; in the Image-to-WebDev Arena test, its score was higher than that of GPT-5.4. The model uses a visual encoder with arbitrary resolution and VideoRoPE, with a language backbone based on a 42-layer hybrid architecture. Currently, the model is available on Ling Studio for free…