AI Models

Alibaba's Qwen3.8-Max Opens the Floodgates for 2.4-Trillion-Parameter Autonomous AI

Alibaba's Qwen team unveiled Qwen3.8-Max, a 2.4-trillion-parameter open-weight model for long-horizon autonomous tasks, with weights releasing next week. The model demonstrated autonomous coding, research reproduction, chip design, and e-commerce management, outperforming human teams in a competition. It arrives amid intense competition in China's open-model ecosystem, following Moonshot AI's Kimi K3 release.

Neura News

Neura News

Neura Market Editorial

August 3, 20266 min read
Alibaba's Qwen3.8-Max Opens the Floodgates for 2.4-Trillion-Parameter Autonomous AI

Alibaba's Qwen team has unveiled Qwen3.8-Max, a 2.4-trillion-parameter open-weight language model built for long-horizon autonomous tasks, with its weights slated for release next week. The model, which marks the first time Alibaba has opened the weights of its flagship Qwen-Max class, is already available through QwenCloud. It arrives amid a heated race in China's open-model ecosystem, just days after rival Moonshot AI released its own giant, Kimi K3.

The announcement, published on August 3, 2026, comes with a suite of case studies showing the model coding entire tools, reproducing research papers, running online stores, and even designing chips. The Qwen team says the model is designed to work independently over days, not just answer single prompts. Its weights will land on Hugging Face and ModelScope next week, giving developers outside Alibaba's cloud direct access.

A Giant With a Light Touch

Qwen3.8-Max packs 2.4 trillion total parameters but activates only 95 billion per query, a design that keeps inference costs manageable. It builds on the Qwen3.5 architecture and is the first model in the Qwen-Max class to ship with public weights. The model supports both the OpenAI Chat Completions format and the Anthropic API protocol, allowing it to plug directly into tools like Claude Code, Codex, Qoder CLI, Qwen Code, and OpenClaw.

A preview was announced mid-July through Alibaba's Token Plan, Qoder, and QoderWork at 10% of the standard price. In that preview, the model ranked just behind Fable 5, a Western frontier model, on internal evaluations. The Qwen team claims the model performs similarly across different agent harnesses and is not tuned specifically for its own environment.

The model includes a reasoning_effort parameter with three levels, letting users dial in how much thinking the model does per task. During reinforcement learning, the team expanded training environments to include multi-day workflows, nested directory structures, and various agent harnesses. That shift paid off: the internal score index across more than 10 benchmarks rose from 0.474 to 0.725. The model performed best at around 4,000 training environments, with scores dipping slightly beyond that point.

Building a Tool Without Human Hands

The Qwen team's first case study shows Qwen3.8-Max building a complete command-line tool, oh-my-cli, in 16 days. The model made 265 commits, opened 127 pull requests, and filed 151 issues by July 30, 2026. The team says no human touched the development process.

The second case study is arguably more striking. The model reproduced a research paper, "Unified Data Selection for LLM Reasoning," without any starter code. It took roughly 5 days and about 125 hours of compute time. The model wrote 7,600 lines of code and ran 33 GPU training jobs. It reproduced all six of the paper's main results, then tested 18 of its own ideas across four rounds. One of those ideas beat the paper's method on the AIME24 math benchmark by 2.7 points.

The team says the model kept making deep structural changes even after hundreds of iterations, a sign of sustained autonomy rather than shallow tweaks. That kind of behavior is exactly what the model was trained for, the team argues.

Beating Human Teams at Their Own Game

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

In a third case study, Qwen3.8-Max entered the WWW2025 Multimodal Dialogue Intent Recognition Challenge, hosted on Alibaba's Tianchi platform. The challenge drew 526 human teams. The model fine-tuned several Chinese language models and Qwen2.5-VL-7B, then combined them into a voting system. Across 45 submissions, accuracy climbed from 0.60 to 0.853. That was enough to outperform 458 of 526 human teams.

The fourth case study moves into hardware. The model designed a cryptographic building block, starting with 8,298 logic gates and reducing them to 678 gates over roughly 500 iterations. An automated layout pass using OpenROAD, the open-source chip layout tool, shrank the chip area from 106x106 micrometers to 46x46 micrometers. That is an 81 percent reduction in chip area.

The fifth case study simulates a full fiscal year of e-commerce. Using anonymized data from Taobao and Tmall, the model started with 100,000 yuan in capital and ran multiple online stores in parallel. It bought products, negotiated with suppliers, adjusted prices, managed returns, and dealt with crises like typhoons and supply chain disruptions. Hidden in the supplier pool were 152 scammers. The model ended with a balance of 416,252 yuan, quadrupling its starting capital. That was 38 percent more than runner-up GLM 5.2 and more than 2.5 times what predecessor Qwen3.7-Max managed. The model invested aggressively early in the year and posted a net profit of over 100,000 yuan during the holiday season.

Multimodal Reach and New Benchmarks

Beyond the case studies, Qwen3.8-Max handles documents over 200 pages and videos longer than 100 hours. The team also introduced RecreationBench, a new benchmark for rebuilding apps without source code, covering Ubuntu, macOS, Windows, Android, and web. A companion library, Qwen-MM-Plugins, adds image and video processing, visual tool use, and multimodal memory.

On self-reported internal runs, the model scored 93 on PaperBench, the highest in the comparison. It scored 86.6 on TerminalBench 2.1, while GPT-5.6 Sol scored 88.8. The Qwen team notes these are internal runs and that independent verification is pending. The model's benchmark comparisons place it near or above Claude Opus 4.8, Claude Fable 5, and GPT-5.6 Sol, though those numbers come from the team's own testing.

The Open-Model Arms Race

Qwen3.8-Max does not exist in a vacuum. On July 27, Moonshot AI released Kimi K3 with open weights on Hugging Face. That model is even larger, with 2.8 trillion parameters and a one-million-token context window. Moonshot also published its attention kernels, MoE communication library, and agent tools. But independent testing tempered those claims: K3 fell well short of top Western models in cyber capabilities and complex math.

Alibaba is pushing in both directions at once. Alongside the giant Qwen3.8-Max, the company recently released Qwen3.6-35B-A3B, a small open model with 35 billion total parameters and just 3 billion active. The company also introduced Qwen-Image-3.0 a few weeks ago, an image generator that handles inputs up to 4,500 tokens and renders readable text as small as ten pixels.

The contrast between Qwen3.8-Max and Kimi K3 highlights a broader trend. Chinese labs are no longer just matching Western models; they are opening up their largest systems for anyone to download. Qwen3.8-Max's release next week will give researchers and startups a chance to run a frontier-scale model on their own hardware, something that was unthinkable just a year ago. Whether the model's self-reported benchmarks hold up under independent scrutiny remains to be seen, but the case studies suggest a level of autonomy that goes well beyond chat.

Related on Neura Market

More from Neura News

AI Tools

CFOs Turn AI Budgeting Into an Infrastructure Discipline for 2026

Chief financial officers are shifting AI spending from experimental funding to disciplined, infrastructure-like management for 2026. The change comes as AI costs escalate rapidly across departments, with pilots expanding into complex, multi-vendor systems. CFOs are now prioritizing high-ROI areas like operational automation and governance, while consolidating fragmented AI infrastructure to maintain financial control.

Aug 7·6 min read