Alibaba's Qwen releases model with 1M-token context window
Alibaba's Qwen team has released Qwen3.8-Omni-Flash, an omni-modal AI model capable of processing text, images, audio, and video. The model supports a context window of up to 1 million tokens.

Alibaba's Qwen team has launched Qwen3.8-Omni-Flash, an AI model designed to process text, images, audio, and video within a single workflow. A key feature is its capacity to handle a context window of up to 1 million tokens, enabling the analysis of significantly longer and more complex datasets.
The company reported that the model's performance shows substantial improvement over its predecessor, Qwen3.5-Omni-Plus. Across 30 different evaluations, an average score increase of over 26% was noted. Gains were particularly observed in audio-video agents, coding tasks, long-context processing, and real-time multimodal interaction.
Qwen3.8-Omni-Flash is suited for various applications, including long-video analysis, meeting summarization, video-based research, and multimodal tool utilization. Alibaba has also introduced Qwen-MM-Plugins and Qwen-Live Harness to support extended and real-time workflows.
The model is now accessible through the Qwen AI platform. Additionally, Alibaba has reduced API input pricing to as low as RMB 0.8 per million tokens, aiming to enhance cost-effectiveness for users.