📣 Send us your press release
Site updates every 15 minutes
Technology

Alibaba's Qwen3.8-Max Challenges GPT-5.6 in Agentic Computing

Chinese tech giant Alibaba has unveiled its new flagship large language model, Qwen3.8-Max. The model claims to outperform competitors in agentic computing tasks, including OSWorld benchmarks.

3 August 2026
Alibaba's Qwen3.8-Max Challenges GPT-5.6 in Agentic Computing

Alibaba's AI research team has introduced its new flagship model, Qwen3.8-Max. This 2.4 trillion-parameter multimodal large language model (LLM) targets autonomous software engineering and long-horizon enterprise tasks.

According to the company's published benchmark results, Qwen3.8-Max surpasses several competing models in key agentic computing areas. Notably, in the OSWorld benchmark, Qwen3.8-Max achieved a score of 86.1, exceeding GPT-5.6 Sol Max (83.2) and Fable 5 (85.0). The model also scored highest on PaperBench and demonstrated strong competitiveness in software engineering, research reproduction, and multimodal reasoning.

The release may signal a strategic shift for Alibaba, as the company plans to release open weights for Qwen3.8-Max and Qwen3.8-27B next week. If released under a permissive license, this would enable self-hosting in enterprise environments, potentially significantly impacting adoption.

Precise licensing terms have not yet been disclosed, leaving open the possibility of a more restrictive custom license, similar to rival Chinese firm Moonshot AI's Kimi K3 model. This could limit the model's widespread use.

The model aims to combine several strengths for enterprise automation. Alibaba describes it as an autonomous coworker capable of executing multi-day projects. It is claimed to independently complete software development projects exceeding ten days, reproduce research papers with thousands of lines of code, and use multimodal feedback for continuous plan adjustment. These demonstrations are company-produced and have not yet been broadly verified by independent evaluators, but they reflect an industry trend towards models that complete entire workflows rather than individual prompts.

Original source: venturebeat.com