The Qwen3 Max family
Against 3.7, Qwen3.8 Max gained three parameters, image and video input, structured outputs, and a doubled output ceiling, and lost one that is rarely documented: Batch Inference is not supported on it. For a batched workload, this upgrade is a regression.
- image and video input on 3.8
- output ceiling doubled
- Batch Inference removed from the architecture
Last checked: This is a reference page. It is re-checked against the vendor sources and updated when a new version ships. Change log
Which version belongs in deployment
If your input is text only and your workload ships in batches, hold on 3.7 Max. If you feed images or video, require structured outputs, or need an answer longer than 65,000 tokens, 3.8 Max belongs in deployment. This is the only family in this reference where the answer to which version is not a single identifier.
The data is extracted from each model dedicated documentation: both carry a 1,000,000-token context window and both cap input at 991,808 tokens, so this parameter creates no distinction. The output ceiling is 65,536 tokens on 3.7 and 131,072 tokens on 3.8, exactly double. The chain-of-thought ceiling is 262,144 tokens on both.
The capability table is where the real distinction lives. 3.7 Max accepts text input only, and its documentation describes it as a pure text interface; 3.8 Max accepts image, text, and video input. Structured outputs are inactive on 3.7 and active on 3.8. And Batch Inference runs the exact opposite: active on 3.7, inactive on 3.8.
A capability removed from the architecture during the upgrade
This last parameter is worth examining separately. The 3.7 Max documentation does not simply log Batch Inference as active; it carries an independent rate for it: batch file input priced at $0.825 and batch output at $2.475, half the standard rate. In the 3.8 Max documentation, Batch Inference is logged as inactive across all six regions, and no batch rate row exists at all.
For any architecture processing several thousand documents overnight in batches, this means migrating to 3.8 doubles the unit cost, with no figure changing in the headline rate column. This class of regression does not surface in standard comparison lists, because those lists check only the rate column, not the capability column.
Rate depends on deployment region, and a temporary discount sits in the middle of it
The 3.8 Max documentation prices six deployment regions independently. Beijing, Frankfurt, Virginia, Tokyo, and Hong Kong all carry $1.65 input and $4.951 output. Singapore, logged as the international region, carries $2 and $6. The figure recorded in our tables is that international row.
On the central pricing page, 3.7 Max shows a list rate of $2.5 and $7.5 in the international row, with a 50 percent limited-time discount label beside it. So deployment today may cost half that rate, while budgeting six months out may not carry the discount at all. We always use the list rate as the basis for calculation, since a temporary discount is not a figure a long-term contract can be built on.
How this model was located, and why that process is itself part of the content
Until a few days ago, the rate and context-window columns for this model were logged as unverified in our table, and the reason was documented explicitly: the English Model Studio catalogue and the central pricing page, both last updated 15 July 2026, did not log the identifier qwen3.8-max, while the Arena leaderboard row carried exactly that identifier. Assuming 3.8 was simply 3.7 would have been the easy assumption, but it was a guess, and once the two documents are placed side by side, that guess is confirmed incorrect.
What further investigation found: a dedicated page exists for this model, updated 3 August 2026. So Alibaba index pages had fallen behind their own models, rather than the model not existing. One detail makes this hard to locate: the URL slug is written with a hyphen while the model identifier uses a dot.
The operational conclusion for any analysis checking the specifications of a Chinese model: the vendor index is not the final source. The model dedicated page address must be checked, even when it is absent from the index.
Specifications
Every number here comes from the vendor's own page, with that page linked beside it.
Release timeline
Each version with its own release date, and what changed against the one before it.
- Qwen3.8 Max current
Using it from Iran
This column carries a date and says how it was checked, because most listicles guess it. Where we have only read a vendor policy page, the note below says exactly that.
- Reachable
- not checked
- Payment
- not checked
- Free tier
- no
How we checked: Alibaba Cloud publishes no blocked-countries list for this service in any source accessible to us, and no test has been run from an Iranian connection. So the access field is logged as unverified.
Documented strengths
- Multimodal input is active on 3.8 Max: image, text, and video, while 3.7 accepts text only
- A 131,072-token output ceiling on 3.8, double that of 3.7
- Structured outputs are supported on 3.8
- The same model deploys at a $1.65 input rate across five global regions, below the international row
- The model dedicated documentation reports maximum input length and thinking-mode ceiling separately, a parameter no other vendor in this group publishes
Known weaknesses and limits
- Batch Inference is inactive on 3.8 Max while active on 3.7, and the 3.7 batch rate was half the standard rate. For a batched workload, this upgrade is a regression.
- Fine-tuning is unsupported on both versions in this family.
- The Alibaba catalogue and pricing pages still do not log qwen3.8-max, so any analysis checking only those two pages concludes the model does not exist.
- The 3.7 Max international rate carries a 50 percent limited-time discount label, so today real figure differs from the list figure, and no end date is published.
- Access from Iran has not been measured, and Alibaba publishes no list in any source accessible to us.
Technical verdict
If your input is text only and your workload runs in batches, hold on 3.7 Max and skip the migration, because the one parameter you would lose is the one halving your cost. For any other workload, 3.8 Max belongs in deployment, and its identifier should be extracted from the model dedicated documentation, not the catalogue.
Frequently raised questions, with documented answers
Is Qwen3.8 Max the same model as Qwen3.7 Max
No. Alibaba official documentation differs on five parameters: input modality (text only against image, text, and video), structured outputs, output ceiling (65,536 against 131,072), Batch Inference, and the international rate row. The context window and maximum input length are identical on both.
Why is Qwen3.8 Max missing from the Alibaba catalogue
The catalogue and the central pricing page were both last updated 15 July 2026, while the model dedicated documentation was updated 3 August. So the indexes have fallen behind the model. The model page follows the same URL pattern as the others; only its slug is written with a hyphen, while the model identifier uses a dot.