Qwen 4 Enters Training as Alibaba Plans 5–10 Trillion Parameter Models

Qwen 4 AI model training with large-scale model architecture and recursive optimization.

Alibaba has confirmed that its next-generation Qwen 4 model is currently in training, while outlining a long-term roadmap that could take future Qwen models to between 5 trillion and 10 trillion parameters.

The announcement came on September 22 at Alibaba Cloud’s annual Apsara Conference in Hangzhou, where the company detailed progress across AI models, chips and cloud infrastructure. Alibaba said Qwen 4.5 and Qwen 5 are projected to scale to the 5–10 trillion-parameter range as the company works toward models capable of handling more complex and longer-horizon tasks.

The company also disclosed progress on recursive self-improvement (RSI), an approach in which AI models use feedback from their own performance to identify weaknesses, construct experiments or data, and iteratively improve parts of their training and inference processes.

A next-generation Alibaba video-generation model is also planned for November, adding another major development to the company’s expanding multimodal AI roadmap.

Quick Summary

  • Qwen 4 is currently in training, according to Alibaba’s latest roadmap.
  • Alibaba says future Qwen 4.5 and Qwen 5 models could scale to 5–10 trillion parameters.
  • The Qwen team is researching recursive self-improvement (RSI).
  • RSI work is being applied to model training, inference and chip-model co-optimization.
  • Alibaba has also announced a next-generation video-generation model planned for November.
  • Specific Qwen 4 variants, pricing and public release specifications remain unconfirmed.
  • Alibaba is expanding Qwen beyond language models into video, speech, image, agents and other multimodal systems.

Qwen 4 Is in Training, Not Yet a Public Model Release

The most significant distinction in the latest announcement is that Qwen 4 has entered training rather than being released to users.

Alibaba described Qwen 4 as its next-generation model and said training is underway. The company has not, in the materials reviewed for this report, published the model weights, API specifications, pricing, or a consumer release date for QWEN 4.

That means claims circulating online that specific QWEN 4 variants have already been officially launched should be treated separately from the confirmed announcement.

In particular, the names Qwen-4-Max, Qwen-4-Flash, Qwen-4-Plus, and Qwen-4-27B were not independently verified in Alibaba’s official conference materials reviewed for this article. They should therefore not be presented as confirmed specifications or released models without a subsequent official model announcement.

The distinction is important for developers because a model being in training does not establish its final architecture, parameter count, availability, API pricing, or product lineup.

Alibaba Plans 5–10 Trillion Parameter, QWen Models

The most significant long-term announcement concerns model scale.

Alibaba says future versions including Qwen 4.5 and Qwen 5 are projected to reach between 5 trillion and 10 trillion parameters. Alibaba CEO Eddie Wu also described the company’s broader objective as developing models capable of increasingly complex and long-running tasks.

Parameters are numerical values learned during model training and are commonly used as a rough indicator of model size. However, parameter counts alone do not determine practical performance, efficiency, or cost.

The proposed 5–10 trillion parameter range would nevertheless represent a substantial increase in scale compared with today’s QWen systems.

Alibaba’s announcement places scaling effort alongside research into model architecture and data optimization rather than presenting parameter growth as the only route to performance.

The company is also building infrastructure to support large AI workloads, including its newly announced Zhenwu V900 AI accelerator. Reuters reported that Alibaba says the chip delivers three times the performance of its predecessor and is intended to support large-scale AI clusters.

Qween’s Recursive Self-Improvement Research

Another major part of the Apsara announcement is Alibaba’s work on recursive self-improvement.

The basic idea is to allow models to participate more directly in parts of the development loop. This includes identifying weaknesses, designing experiments, generating or validating data, evaluating results, and using the resulting feedback to improve subsequent iterations.

Alibaba said the Qwen team has already begun applying RSI techniques to areas including model training, inference optimization and chip-model co-optimization.

One example involved Qwen3.8-Max autonomously constructing parts of a training process, generating training data, designing experiments and diagnosing problems. Alibaba reported that the system ran for more than a month and completed 33 effective iterative cycles without human intervention in that particular experiment.

Alibaba also reported an inference optimization experiment in which Qwen3.8-Max adapted an inference framework for a new GPU architecture, increasing single-instance throughput by 96%.

These are company-reported research demonstrations, rather than evidence that future Qwen models will automatically improve themselves without human involvement in general deployment.

RSI Is Also Used in Chip Design

Alibaba’s RSI work extends beyond model training.

The company described an experiment in which Qwen3.8-Max worked through the chip design process using electronic design automation tools. Alibaba said the system operated for more than 60 hours and made more than 10,000 EDA tool calls while iterating on a chip bus-module design.

Alibaba says the resulting optimization reduced the physical area of the design by 42% without reducing performance.

This is notable because it illustrates Alibaba’s broader attempt to use AI across its entire development stack. This is rather than limiting models to generating text or code.

The company has successfully established a feedback loop between models, software tools, and computing hardware. AI is participating in portions of both model development and chip engineering.

A New Alibaba Video Model Coming in November

Alibaba also confirmed that its next-generation video-generation model is planned for November.

According to reporting from the Apsara Conference, the upcoming system is being developed around longer and more controllable video generation. Alibaba’s stated direction includes maintaining consistency over long sequences and providing creators with increased control over characters, environments, and camera movements.

The company also described a shift from generating individual shots toward understanding broader narratives. This suggests that the model is designed for more structured video-production workflows.

Specific model name, pricing, context limits, and release specifications were not established in the announcement reviewed here, so those details should not be treated as confirmed until Alibaba publishes them.

Qwen’s Roadmap Goes Beyond Text

The announcements at Apsara also point to a broader Alibaba Qwen ecosystem expansion.

Alibaba highlighted updates across video, speech, image, world-model and music systems, alongside its foundation-model work. The company is also developing agent technologies designed to bring Qwen capabilities into more practical workflows and devices.

That broader direction matters because Qwen is increasingly positioned as more than a conventional language-model family. Alibaba’s current strategy connects foundation models with multimodal systems, AI agents, specialized applications, chips, and cloud infrastructure.

For developers, however, the immediate Qwen 4 story remains one of development rather than availability. The model is in training, while many final specifications and release details remain undisclosed.

The long-term roadmap is clear: Alibaba is targeting larger Qwen models, experimenting with recursive self-improvement across the AI stack, and preparing additional multimodal systems such as its next-generation video model.

For now, Qwen 4 should therefore be understood as an upcoming model under active development. This is rather than a released model family with confirmed public variants or pricing.

Also Read –

Qwen3.8-LiveTranslate Launches for Real-Time Translation

Qwen-Image-2.1 Launches as Open-Weight AI Image Model

Source

Alibaba Cloud – 2026 Apsara Conference

Alibaba Cloud – 2026 Apsara Conference Homepage

Reuters – Alibaba plans AI model with 5 trillion to 10 trillion parameters

IT之家 – Qwen 4 training and next-generation video model

Qwen – Official model and research publications

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top