Maker

The model infrastructure Alibaba operates: what is in service

Alibaba operates the Qwen line through a service called Model Studio, and the newest deployed unit on that line is qwen3.8-max: a one-million-token context window, an output ceiling of 131,072 tokens, and a rate of $2 in and $6 out in the international region. This catalogue documentation is the deepest in the group, yet it has simultaneously fallen behind the models it has deployed.

  • one-million-token context window
  • pricing tied to deployment region
  • separate embedding and reranking models

Last checked: This is a reference page. It is re-checked against the vendor sources and updated when a new version ships. Change log

Model families

The Qwen Max line, where catalogue documentation has fallen behind actual model deployment.

Versions

Versions Deployed Status Model families
Qwen3.8 Max current Qwen3 Max

The nominal window is one million tokens; your usable capacity is not

The qwen3.7-max specification publishes four separate figures where the rest of this group registers only one: a context-window capacity of 1,000,000 tokens, a maximum input of 991,808 tokens, a maximum output of 65,536 tokens, and in thinking mode an input capacity of 983,616 tokens with a chain-of-thought ceiling of 262,144 tokens.

This breakdown has direct engineering value. The phrase "one million token window" looks identical across every vendor marketing, but the usable throughput you can actually send is lower, and lower again once thinking mode is engaged. Alibaba is the only vendor that documents this constraint, which is why we keep the maximum-input figure in the text.

Pricing is a function of the calling region

The pricing page registers a rate of $2.5 in and $7.5 out for the international region, with no tiering applied up to the one-million-token ceiling. But two other units on the same line carry a tiered structure: qwen3.7-plus runs $0.4 and $1.6 below the 256K-token threshold and reaches $1.2 and $4.8 above it, and qwen3.6-flash follows the same pattern, jumping from $0.25 and $1.5 to $1 and $4. The same request, once it grows longer, is billed at three times the rate.

The model dedicated page also publishes a price table broken down by deployment region, and the China (Beijing) row shows a lower rate. That row does not state its currency on the page, so we do not register it and treat the international figure as the reference. We note this because any round-up quoting a single rate for Qwen leaves it unclear which deployment region it read.

A catalogue that extends past the conversational model

Alongside its three text models, Model Studio also carries embedding models (text-embedding-v4 and tongyi-embedding-vision-plus) and a reranking model (qwen3-rerank) in its catalogue, plus the multimodal qwen3.5-omni-plus. None of the other three Chinese vendors in this reference register this level of catalogue diversity. If you are building a semantic search or retrieval pipeline, that architectural difference matters more than a few leaderboard points.

A catalogue that has fallen behind actual model deployment

The Alibaba row on the Arena leaderboard is registered under the id qwen3.8-max, and the English Model Studio catalogue does not list a model of that name. The last-updated date on both the catalogue and the central pricing page is 15 July 2026.

But the model itself is deployed. Its dedicated page sits on the same URL pattern as every other model, was updated on 3 August 2026, and publishes the full specification: context window, maximum input and output, a capability table and pricing for six deployment regions. The issue was never that the model does not exist; the issue is that Alibaba catalogue pages have fallen behind the actual deployment of its models.

And the two units are not the same. qwen3.8-max accepts image and video input, where 3.7 defines itself as a pure text interface; its output ceiling is double, and Batch Inference, which was active on 3.7, is not supported on it. The full specification comparison is documented on the family page.

Documented strengths

  • The broadest catalogue among the Chinese vendors in this reference: text, multimodal, embedding and reranking models in a single service
  • The only vendor that documents maximum input and a chain-of-thought ceiling separately from context-window capacity
  • Its integration surface supports the OpenAI format, the Anthropic format and its own proprietary DashScope protocol
  • A one-million-token window on qwen3.8-max at $2 and $6 in the international region, plus image and video input support

Known weaknesses and limits

  • The official catalogue and the central pricing page still have not registered a unit named qwen3.8-max, even though the leaderboard row and the model own dedicated page carry exactly that id. A buyer checking only those two pages concludes the model does not exist.
  • The output ceiling on qwen3.7-max is only 65,536 tokens, half the qwen3.8-max ceiling and the lowest output capacity in this group against DeepSeek 384K tokens.
  • Batch Inference is not supported on qwen3.8-max, where it was active on qwen3.7-max, and the 3.7 batch-processing rate was half the standard rate.
  • Pricing is a function of deployment region, and the model page publishes separate tables for Beijing, Frankfurt, Virginia, Tokyo and Hong Kong. There is no single rate for the price of Qwen.
  • Two other units on the same line triple in price above the 256K-token threshold, so the cost of a long request does not scale linearly.
  • It publishes no supported-countries list, so the access-from-Iran column stays in unverified status.

Technical verdict

If you are building a pipeline that needs embedding and reranking alongside a conversational model, Alibaba is the only integrated option in this group. If you are after the exact model you saw on the leaderboard, read the id off the model dedicated page rather than the catalogue, because the catalogue lags behind deployment.

Frequently raised questions, with documented answers

What does the qwen3.7-max pricing structure look like

In the international region the list rate is $2.5 per million input tokens and $7.5 per million output tokens, with no tiering up to a million tokens, and a note beside it reads limited-time 50 percent discount. The model dedicated page registers $1.65 and $4.951 for other regions, so the amount you pay depends on the region you call.

Where is qwen3.8-max registered in the catalogue

It is not registered in the catalogue, but its own dedicated page exists and was updated on 3 August 2026. The catalogue and the pricing page were both last updated on 15 July, meaning they trail the model deployment date. One technical wrinkle: the URL slug uses a hyphen, while the model id uses a dot.

Sources

  1. Model Studio supported modelsvendor sourcewww.alibabacloud.comread on 12 August 2026
  2. qwen3.8-max model infovendor sourcewww.alibabacloud.comread on 12 August 2026
  3. qwen3.7-max model infovendor sourcewww.alibabacloud.comread on 12 August 2026
  4. Model Studio model pricingvendor sourcewww.alibabacloud.comread on 12 August 2026