2.4 Trillion Parameters? What Sets Alibaba's New AI Model 'Qwen3.8-Max' Apart?

A graphic visualization of a vast network of data connections
AI Summary

Alibaba's new AI model, Qwen3.8-Max, is a massive model with 2.4 trillion parameters that maximizes reasoning capabilities by mandatorily applying a 'thinking mode' to all responses.

Imagine this: what if the artificial intelligence (AI) we use every day took a moment to ponder before delivering an answer when asked a complex math problem or asked to write a very long report? It would be just like a human deep in thought before solving a difficult problem. The new AI model recently released by Alibaba, ‘Qwen3.8-2.4T (Qwen3.8-Max)’, operates in exactly that way.

Why Is It Attracting Attention?

Improved performance of AI assistants in our daily lives means we can entrust more complex tasks to AI. Going beyond simple questions like the weather, AI can now more accurately handle research or complex programming tasks that require logical reasoning. Alibaba is confident that this model is one of the most powerful currently in the industry and possesses capabilities comparable to the world’s top-performing ‘Fable 5’ model [Source 4]. In particular, this model is highly significant to developers as it is an ‘Open-weights’ model (anyone can download and use the model weights) [Source 2].

Easily Understanding ‘Parameters’ and ‘Thinking Mode’

Have you heard the term ‘parameter’? Simply put, you can think of it as a unit for measuring the size of an AI’s brain. The more parameters there are, the more knowledge the AI can hold and the more complex relationships it can learn. This model has a whopping 2.4 trillion (2.4T) parameters, an enormous scale that is over 30,000 times the total population of South Korea [Source 2].

However, running such a vast number of parameters requires a tremendous amount of energy. To solve this, the model uses a clever method called ‘Sparse Mixture-of-Experts’.

AD

Imagine there are 512 teachers in a school. Wouldn’t it be much more efficient to call only the math teacher when you have a math question, and only the history teacher when you have a history question? This model is the same. It keeps 512 ‘Experts’ in total, and when a question comes in, it activates only the 10 experts related to it and 1 common expert to generate an answer [Source 5]. By doing this, it maintains the number of parameters actually used for response generation at about 95 billion—a fraction of the total—while still being able to utilize vast knowledge [Source 5].

Furthermore, this model mandatorily goes through a ‘thinking mode’ for every question. It is designed to elicit deeper and more refined answers by mandatorily performing a process where the AI checks its own logic and reasons before answering [Source 1].

What Is Its Current Status?

While Qwen3.8-2.4T boasts powerful capabilities, there are things to know before use.

First, this model is ‘text-only’. Unlike many recent AIs that are equipped with ‘multimodal’ functions that understand images or videos as well, this model is focused solely on the ability to read and write text [Source 1].

Second, it is very large in scale. It is so massive that it requires a whopping 4.9TB of storage space to run without loss [Source 8]. Therefore, it is a model that will primarily be utilized by companies or research institutes via servers rather than by individual users installing it on their personal computers.

Future Outlook

Starting with this Qwen3.8-Max, Alibaba plans to continuously release smaller and more efficient models as well [Source 14]. In the AI industry, work comparing and verifying the performance of this model against other major models is expected to continue actively [Source 17]. The point to watch is how much this model will be able to save people time in complex coding, research, and logical task performance in the future.

MindTickleBytes AI Reporter’s View

Alibaba has once again shown off its technical prowess and thrown an important topic into the AI open ecosystem. The fact that it didn’t just stop at increasing the number of parameters, but mandatorily applied a ‘thinking mode’ to maximize reasoning ability, will serve as an important benchmark for future AI model design.

References

  1. Qwen/Qwen3.8-2.4T-A95B · Hugging Face
  2. [Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72 NVIDIA Technical Blog](https://developer.nvidia.com/blog/serve-qwen3-8-2-4t-a95b-a-2-4t-parameter-model-with-configurable-reasoning-on-nvidia-gb300-nvl72/)
  3. Qwen/Qwen3.8-2.4T-A95B-FP8 · Hugging Face
  4. Qwen on X: “Qwen3.8 is launching and going open-weight soon!
  5. [Qwen/Qwen3.8-2.4T-A95B vLLM Recipes](https://recipes.vllm.ai/Qwen/Qwen3.8-2.4T-A95B)
  6. RadixArk/Qwen3.8-2.4T-A95B-NVFP4 · Hugging Face
  7. Qwen/Qwen3.8-2.4T-A95B · OMG
  8. [Qwen3.8- How to Run Locally Unsloth Documentation](https://unsloth.ai/docs/models/qwen3.8)
  9. Qwen38Max: Alibaba’s New AI Model Is Coming for the AI… - YouTube
  10. [Qwen3.82.4TA95B - API Pricing & Providers OpenRouter](https://openrouter.ai/qwen/qwen3.8-2.4t-a95b)
  11. Kimi K3 vsQwen3.8-Max: Benchmarks, Pricing & API Access
  12. Qwen3.8-27B Is Almost Here, And It’s About to Reset the… - Banandre
  13. Qwen3.8 Preview: 2.4T Params, Open Weights, Release
  14. Qwen3.8-Max: A New Bar for Coding and Cowork
  15. [Qwen3.8 OpenLM.ai](https://openlm.ai/qwen3.8/)
  16. [AINews] Qwen 3.8 Max(2.4T) and 27B, new open weights models …
  17. Alibaba Launches Qwen 3.8 With 2.4 Trillion Parameters …
  18. Qwen 3.8 & Qwen 3.8 27B Open Source: Official Alibaba …
  19. Qwen 3.8 Preview — 2.4T Multimodal, and What’s Not Yet Public
AD
Test Your Understanding
Q1. Which of the following is correct regarding the 'thinking mode' of the Qwen3.8-2.4T model?
  • Users can turn it on or off as needed.
  • It must be forced in all interactions.
  • It turns on automatically only when analyzing images.
Qwen3.8-2.4T forces the 'thinking mode' in all interactions and it cannot be disabled.
Q2. Although this model has a total of 2.4 trillion parameters, how many parameters are activated at once when generating an actual response?
  • About 27 billion
  • About 95 billion
  • All 2.4 trillion
It has a structure where about 95 billion parameters are activated per token out of a total of 2.4 trillion.
Q3. Which of the following is correct regarding the input method of the Qwen3.8-2.4T model?
  • It is a multimodal model that handles text and images simultaneously.
  • It is a text-only model.
  • It includes video analysis capabilities.
Qwen3.8-2.4T is a text-only model and does not support multimodal inputs.
2.4 Trillion Parameters? Wh...
0:00