Independent tests put Qwen3.8-Max top of agentic index
Artificial Analysis published its full evaluation of Alibaba's 2.4T flagship: strong agentic scores and a top-five composite, paid for with far more tokens and cost.
The numbers
Artificial Analysis released its complete evaluation run for Qwen3.8-Max on Thursday, several days after Alibaba opened the 2.4-trillion-parameter model through its API. The composite Intelligence Index — a blend of nine evaluations including GDPval-AA, Terminal-Bench, SciCode and Humanity's Last Exam — puts the model at 56, well above the 32 median for comparable reasoning models. Alibaba's Qwen team said the result places it fifth overall on the index and first on the agentic sub-index, which would make it the first Chinese model to lead that category.
On GDPval-AA, the evaluator's proxy for real knowledge work, Qwen3.8-Max recorded 1,739 Elo. That puts it ahead of Moonshot's Kimi K3 at 1,685, effectively level with Claude Fable 5 at 1,743 and GPT-5.6 Sol (max) at 1,730, and behind only Anthropic's top configuration.
The trade-offs
The gains are bought with effort. The model averaged 64 turns per GDPval-AA task against 14 for its predecessor, and generated roughly 150 million output tokens across the index versus a 66 million median — unusually verbose even by reasoning-model standards. Cost follows: $1.14 per Intelligence Index task, more than double Qwen3.7-Max's $0.53 and about 1.3x the open-weights leader Kimi K3 at $0.86.
One result moved the wrong way. On AA-Omniscience, which nets correct answers against confident wrong ones, the score fell about ten points, from +14 to +4. Raw accuracy stayed roughly flat near 31%, meaning the drop comes from the model asserting more when it should abstain — reversing a calibration gain its predecessor had made.
Why it matters
Vendor benchmarks are cheap; third-party runs with published cost and token accounting are not. This one confirms that Alibaba has closed most of the capability gap with US frontier labs on agentic work, while showing the bill has moved with it — a Chinese flagship is no longer automatically the cheap option. The hallucination regression is the more durable warning: agentic systems that chain dozens of turns compound confident errors, and calibration is now a harder constraint on deployment than raw score.