Alibaba on August 3 released Qwen3.8-Max, which it called its largest and most capable AI model to date, and said its scores on several industry tests now rival those of Anthropic's Claude at a fraction of the price. The announcement pushed Alibaba's Hong Kong shares up more than 6%.
Qwen3.8-Max has 2.4 trillion total parameters but uses a "mixture-of-experts" design that routes each request through about 95 billion of them, which keeps running costs down. It also works across formats, handling text, images and video, and can read up to one million tokens of input at once, roughly a long book's worth of text.
The headline number is price. Alibaba is charging $2 per million input tokens and $6 per million output tokens. That is about 40% of what Anthropic charges for input on Claude Opus 5 and roughly a quarter of its output price of $25. On LMArena, a public leaderboard where users vote for the answers they prefer, the model placed second in the vision category and fifth for text. That made it the highest-ranked Chinese text model, though still behind Anthropic's Claude family.
On its own published tests, Alibaba said Qwen3.8-Max led on some coding and agent tasks, scoring 86.6 on Terminal-Bench 2.1 (a test of AI agents working at a command line) against 84.6 for Claude Opus 4.8. It fell behind on others, such as SWE-bench Pro, a software-engineering benchmark, and a hard reasoning exam. Those figures came from Alibaba's own internal runs, with no independent testing available at launch, and the company had not yet published a model card or methodology for the model. Alibaba plans to release the model's weights for free download the following week, alongside a much smaller 27-billion-parameter version that can run on ordinary hardware.
The release lands a week after Moonshot AI put out Kimi K3, a 2.8-trillion-parameter open-weight model. Read together, the two launches point to a shared strategy among Chinese labs: match Western models closely enough on capability, then win developers on price and open access rather than on topping any single leaderboard.
That is also how analysts read the market reaction. Citi told clients the share gain "likely reflects positive cloud revenue read-throughs" and the model's improved scores, but warned that frequent releases are pushing companies toward a "model-agnostic" approach, meaning they pick the cheapest model that does each job. Citi said that erodes any one model's edge.
Outside researchers will be able to run their own tests once those open weights ship. For buyers, the shift worth watching is not who wins a benchmark this month, but how far the price of frontier-grade AI keeps falling.