The GLM-5 family
GLM-5.2 and GLM-5.1 carry exactly the same rate, $1.4 in and $4.4 out, but the 5.2 context window is one million tokens against 200,000 for 5.1. So the version that belongs in an active deployment is 5.2, and 5.1 has no technical justification for a new build.
- four active versions on the price table
- 5.2 window is five times 5.1
- rate increased along the production line
Last checked: This is a reference page. It is re-checked against the vendor sources and updated when a new version ships. Change log
Which version belongs in your architecture
GLM-5.2. Z.ai price table logs a single rate for 5.1 and 5.2, not two close figures: $1.4 in and $4.4 out. The documentation page for both versions also states a 128K-token output ceiling. The difference sits in one parameter, context window: 5.1 holds at 200,000 tokens while 5.2 reaches one million. When two models carry the same rate and one accepts five times the input capacity, the choice between them stops being an open engineering decision.
A rate that rose along the production line
GLM-5 shipped at $1 input and $3.2 output. 5.1 and 5.2 arrived at $1.4 and $4.4, a 40 percent increase on input and roughly 37 percent on output. This is the exact inverse of the pattern documented on the Anthropic Opus family, where the rate held flat across five consecutive versions.
The operational consequence for any deployed service is this: migrating from GLM-5 to 5.2 raises the invoice, and should not be mistaken for a no-cost upgrade. In exchange, that same migration multiplies the context window by five. Not a bad trade, but still a trade.
Turbo is not what its name implies
GLM-5-Turbo, priced at $1.2 and $4, sits between 5 and 5.1, and its name implies one parameter: higher speed. Its own documentation states something else. It is presented as a "ClawBench Enhanced Model," "deeply optimized for the OpenClaw scenario," and no latency or throughput data appears anywhere on that page. Its context window is the same 200,000 tokens as GLM-5, while its rate runs 20 percent higher.
So deploying Turbo for speed means deploying against a parameter the vendor has never advertised for it. If your workload runs inside OpenClaw, the calculus changes, and Turbo becomes the only version in this family making a specific claim about that workload.
A timeline with no dates logged
The release order of these four versions can be extracted from the documentation index. Release dates cannot. None of Z.ai guide pages logs a release date, so it cannot be stated how many months separate 5 from 5.2, or how long any version has been in service. For a section where every other parameter is dated, this gap in the data record is notable.
The same documentation also leaks a staleness signal: the 5.1 page still describes itself as "the latest flagship model at Z.ai," while 5.2 carries a HOT label in the index and is described as flagship on its own page. Both pages are live today and contradict each other.
Release timeline
Each version with its own release date, and what changed against the one before it.
- GLM-5.1 current
- GLM-5.2 Max current The second cheapest output rate in this coding table, behind DeepSeek only.
- GLM-5.3-Flash current
- GLM-5.3 current
Using it from Iran
This column carries a date and says how it was checked, because most listicles guess it. Where we have only read a vendor policy page, the note below says exactly that.
- Reachable
- not checked
- Payment
- not checked
- Free tier
- no
How we checked: Z.ai publishes no supported-countries list, so the access field is logged as neither open nor blocked. No test has been run from an Iranian connection either.
Documented strengths
- A one-million-token context window on GLM-5.2, at the same rate as the 200,000-token version
- Cached input priced at $0.26 on 5.2, roughly a fifth of the standard rate
- Four versions run active simultaneously, so no migration is forced on the architecture
- The cheapest version in this family still ships at a $1 input rate
Known weaknesses and limits
- Z.ai publishes no release date for any of these four versions, so this family has a release order but no timeline.
- The rate rose along the production line, so migrating from GLM-5 to 5.2 directly raises the invoice.
- The cached input storage field on the price table is logged as free for a limited time, meaning a parameter that reads zero today is scheduled to change.
- Z.ai publishes no comparative benchmark data between 5.1 and 5.2, so the difference in answer quality is not measurable, only the context window is.
- Access from Iran is unverified, and Z.ai publishes no country list.
Technical verdict
For a new deployment, select GLM-5.2 and do not evaluate 5.1 at all, since their rates are identical. If you already run a service on GLM-5, calculate the 40 percent input-rate gap against your own real volume before migrating, not against the unit rate.
Frequently raised questions, with documented answers
How much better does GLM-5.2 perform than 5.1
On paper the main difference is the context-window parameter: one million tokens against 200,000. Z.ai publishes no comparative benchmark data between the two, so any claim about answer quality between these versions is unsourced. Their rates are identical.
Is GLM-5-Turbo faster
Its own documentation makes no such claim. It is presented as a "ClawBench Enhanced Model" optimized for the OpenClaw scenario, and no latency or throughput data is published. Its context window is the same 200,000 tokens as GLM-5, while its rate runs 20 percent higher.