Best AI for coding
In today configuration, Claude Opus 5 takes the highest score: it holds first place on the WebDev Arena leaderboard at 1692 while its price is half that of the most expensive model on the market. The table below shows the architecture of the criteria and each one weight, and every number links directly to its source.
- eight models with traceable evidence
- criterion weights on the page
- no manually placed positions
Last checked: This is a reference page. It is re-checked against the vendor sources and updated when a new version ships. Change log
Where each tool's score comes from
The ring below is the same set of weights printed in the table headers further down.
- WebDev leaderboard 40
- Score per dollar 20
- Tooling infrastructure maturity 15
- General text leaderboard 10
- Context window capacity 10
- Access from Iran 5
We chose these weights, and that is the only judgment call in the table. Weight them differently and the order changes.
Today ranking, built from verified evidence
A model that outranks our pipeline on the leaderboard is entered into the table even where its dedicated page has not been written yet; otherwise, first place would only mean first among the candidates we had capacity to document.
An empty cell means we could not verify that number, not that the tool scored zero.
Behind each number
Every judged score in this table carries a written reason, and the Iran column says how it was checked. Those two are open here. The measurement trail behind each number, which row of which leaderboard and on how many votes, opens under the model it belongs to.
1 Claude Opus 5
Tooling infrastructure maturity 4/4 The integration surface is complete: API, official app, command-line tool, agent SDK, batch-execution pipeline, and deployment across the infrastructure of three major clouds
Access from Iran blocked / no working route Iran is registered on neither of Anthropic two supported-countries lists, at the API layer or the claude.ai layer. That status is extracted from Anthropic own policy document, not from a direct network measurement; an independent network test has not been run yet, and its result will be logged here once it is.
Measurement trail (3)
WebDev leaderboard 1,692 Source data, the claude-opus-5-max configuration row: rank one, derived from 567,446 votes, updated 11 August 2026
Score per dollar 67.7 WebDev Arena score divided by the price per million output tokens
General text leaderboard 1,494 Source data, the claude-opus-5-high configuration row: rank seven, derived from 7,765,220 votes
2 Claude Fable 5
Tooling infrastructure maturity 4/4 Built on the same deployment infrastructure as Opus 5, with full coverage across all five delivery routes: Anthropic's own API, Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry.
Access from Iran blocked / no working route Iran is not registered on Anthropic supported-countries list. Extracted from Anthropic own policy document, not from a network measurement.
Measurement trail (3)
WebDev leaderboard 1,628 Source data, the WebDev Arena leaderboard, updated 11 August 2026, derived from 567,446 votes
Score per dollar 32.6 WebDev Arena score divided by the price per million output tokens
General text leaderboard 1,506 Source data, the Arena text leaderboard, updated 11 August 2026, derived from 7,765,220 votes
3 Qwen3.8 Max
Measurement trail (3)
WebDev leaderboard 1,670 Extracted from the WebDev Arena leaderboard, logged update 11 August 2026, 567,446 votes
Score per dollar 278.3 WebDev Arena score divided by the price per million output tokens
General text leaderboard 1,490 Extracted from the Arena text leaderboard, logged update 11 August 2026, 7,765,220 votes
4 Kimi K3 Max
Tooling infrastructure maturity 1/4 The integration surface includes an OpenAI-compatible API plus one official client. The platform documentation lists no official command-line tool or editor extension, so no higher score is assigned.
Measurement trail (3)
WebDev leaderboard 1,674 Extracted from the WebDev Arena leaderboard, per the 11 August 2026 update and a total of 567,446 measurement votes
Score per dollar 111.6 WebDev Arena score divided by the price per million output tokens
General text leaderboard 1,487 Extracted from the Arena text leaderboard, per the 11 August 2026 update and a total of 7,765,220 measurement votes
5 GPT-5.6 Sol
Tooling infrastructure maturity 4/4 The top tier is confirmed through the maker own technical documentation: an API, SDKs, an official CLI, the Codex editor extension, an Agents SDK, a batch execution lane, and a first-party deployment guide for Amazon Bedrock that records the id openai.gpt-5.6-sol. This is the one column that reached full completion for this model without openai.com ever becoming reachable.
Measurement trail (3)
WebDev leaderboard 1,623 The gpt-5.6-sol-xhigh (codex-harness) entry on the WebDev Arena leaderboard, position 6 of 113 recorded models, per the 11 August 2026 update and 8,254 measurement votes on this entry
Score per dollar 54.1 WebDev Arena score divided by the price per million output tokens
General text leaderboard 1,481 The gpt-5.6-sol-xhigh entry on the Arena text leaderboard, position 18 of 389 recorded models, per the 11 August 2026 update and 15,226 measurement votes on this entry
6 DeepSeek V4 Flash
Tooling infrastructure maturity 1/4 Integration runs through a documented API and one official client. No official agent tooling or editor extension is listed.
Measurement trail (3)
WebDev leaderboard 1,585 Extracted from the WebDev Arena leaderboard, the build updated 11 August 2026, based on 567,446 recorded votes
Score per dollar 1,200.8 WebDev Arena score divided by the price per million output tokens
General text leaderboard 1,435 Rank 83 out of 389 models on the Arena text leaderboard, the build updated 11 August 2026, with 49,111 recorded votes for this row
7 GLM-5.2 Max
Tooling infrastructure maturity 1/4 The integration surface includes a documented API plus one official client. No official command-line tool or editor extension is recorded in the technical specification.
Measurement trail (3)
WebDev leaderboard 1,588 Extracted from the WebDev Arena leaderboard, per the 11 August 2026 update and a total of 567,446 measurement votes
Score per dollar 360.9 WebDev Arena score divided by the price per million output tokens
General text leaderboard 1,470 Position 33 of 389 recorded models on the Arena text leaderboard, per the 11 August 2026 update and 26,703 measurement votes on this entry
8 Grok 4.5
Tooling infrastructure maturity 3/4 The integration surface includes an OpenAI-compatible API, official Python and JavaScript SDKs, and an official command-line agent named Grok Build. The top tier is out of reach because no first-party editor extension is built into the architecture, and the batch execution lane does not accept this specific model.
Measurement trail (3)
WebDev leaderboard 1,554 Position 12 of 113 recorded models on the WebDev Arena leaderboard, per the 11 August 2026 update and a total of 567,446 measurement votes
Score per dollar 259.0 WebDev Arena score divided by the price per million output tokens
General text leaderboard 1,469 Position 34 of 389 recorded models on the Arena text leaderboard, per the 11 August 2026 update and a total of 7,765,220 measurement votes
Not enough verified data to rank
These are real tools we track, but more than half of their criteria have no verified number yet, so ranking them would be a guess.
- Gemini 3.1 Pro
How this ranking is calculated
Every criterion below has a weight and a source. Change a weight and the whole table recomputes. There is no hand-placed position anywhere in this hub.
| Criterion | Weight | Evidence |
|---|---|---|
| WebDev leaderboard | 40 | WebDev ArenaMeasured on real web-development tasks with tools across multiple execution steps, closer to daily workload than a single-step exam |
| Score per dollar | 20 | computed from two numbers aboveFor a team carrying a monthly bill, provisioning a costlier model for one percent more score is a poor economic decision |
| Tooling infrastructure maturity | 15 | defined scale, with a written reason per assignmentA model without a command-line tool and an agent SDK runs slower in daily deployment, regardless of raw model capability |
| General text leaderboard | 10 | Arena (Text)Weighted low because general conversation is not a coding metric, though it is not fully irrelevant either |
| Context window capacity | 10 | vendor stated specificationOn a large repository, a small window means chunking input and losing execution context |
| Access from Iran | 5 | our access column, with its method statedWeighted low because this axis has its own dedicated page under AI in Iran |
Why Opus 5 leads
Two factors operate at once. First, on WebDev Arena, which measures real web-development work with tools across multiple execution steps, the claude-opus-5-max row holds first place at 1692, ahead of Kimi K3 Max at 1674 and Qwen3.8 Max at 1670. Second, its output cost is $25, half of Fable 5, which itself scores lower at 1628. A model with both higher throughput and lower cost lands first under any reasonable weighting configuration.
Where this table falls short
The primary criterion here is a human-preference leaderboard, not an automated benchmark. WebDev Arena publishes an absolute score, a date, and a vote count, which is what publication on this page requires, but human preference is also sensitive to readability and answer shape, not only to whether the code is correct. SWE-bench Verified would be a sharper instrument, but its leaderboard renders in JavaScript and we have not read its figures directly. A number we have not read is not published on this page, even where another source has quoted it.
One candidate on this page still carries no rank: Gemini 3.1 Pro. It holds a real score on the general text leaderboard, but its WebDev Arena row has not been read, and Google has not published a context window for it at all, leaving more than half the criteria weight empty for it. Its name is listed below the table; its rank is not. GLM-5.2 Max and DeepSeek V4 Flash sat in the same state for a while, and now that their price and context window are verified from their own documentation, they carry a rank.
One more point, so the numbers do not read as inconsistent: the scores are relative. Every column is normalized between its own minimum and maximum, so onboarding a new candidate shifts every row score without any model having changed. Adding Grok 4.5 on 12 August 2026 produced exactly that effect: the order held and the numbers rose. If today figure does not match a screenshot from last week, this mechanism is the reason, not manual intervention.
Two more data changes shifted the scores the same day. First, the text leaderboard was re-read in full this time, and it turned out the GLM-5.2 Max and DeepSeek V4 Flash rows had been there from the start; the earlier pass had only read the top of the page. The general-text column was filled in for both. Second, GPT-5.6 Sol entered as an eighth candidate: openai.com still does not respond to our infrastructure requests, but its developer documentation publishes the price and context window, and its leaderboard row exists. One note on that row, since this is the dedicated coding page: it is named gpt-5.6-sol-xhigh (codex-harness), meaning the score belongs to the model inside OpenAI own coding agent, not to the bare model. GPT landed between Kimi and Qwen, and no other row order changed relative to the rest.
28 August: DeepSeek pricing rose and three rows shifted
This is the only number in this reference that has risen to date. Through 12 August, DeepSeek V4 Flash cost $0.14 in and $0.28 out. Its official pricing page now carries two columns, peak and off-peak, and even the cheaper column sits above the old price: $0.22 and $0.66 off-peak, $0.44 and $1.32 at peak. We record the peak column, because DeepSeek peak window overlaps the middle of the Iranian working day.
DeepSeek itself did not lose a rank from this change, since it remains the cheapest row in the table and still holds the value column. But its margin over the rest narrowed, and because every column is normalized between its own minimum and maximum, that narrowing lifted three other rows: Qwen3.8 Max from fifth to third, Kimi K3 Max from third to fourth, and GPT-5.6 Sol from fourth to fifth. None of these three models changed anything in their own specification. This is the operational definition of a "relative ranking," and it is why the arithmetic is published on this page.
What only shows up in daily deployment
The three top leaderboard scores in this table sit inside a 22-point band, 1692, 1674, and 1670, a spread that is felt far less in practice than the number suggests. What actually changes a working day is tooling infrastructure maturity: whether the model has a command-line tool, supports batch execution, and is deployed on a cloud your organization already has under contract. That is why this criterion carries fifteen weight points, not five.
When our top pick is not the right fit
- If your team infrastructure runs on a cloud that does not host Opus 5, the second model in the table with working tooling is more usable than the first model without it.
- If your workload is general conversation rather than coding, do not use this table as your reference; the text leaderboard orders things differently.
- If budget is genuinely constrained, the open-weight models lower in the table trail by less score than they trail in price.
Known weaknesses and limits
- The primary criterion in this ranking is a human-preference leaderboard, not an automated test against real repositories.
- The SWE-bench Verified figure is not entered in the table, because we could not read that leaderboard directly.
- Anthropic own Frontier-Bench claims are relative and carry no absolute figure, so that column stays empty.
- Qwen3.8 Max now has a verified price and context window, but its tooling-maturity column and Iran column still carry no evidence, meaning 20 of its 100 weight points remain unmeasured. That is a gap in our data, not a weakness in the model.
Technical verdict
If you must pick one model and access from Iran is not a constraint, deploy Opus 5. If cost is the constraint, audit your own token usage before switching models: in the workloads we have observed, trimming execution context saves more than switching model does.
Frequently raised questions, with documented answers
What is the best AI for coding
In today configuration, Claude Opus 5, on the strength of its first-place rank on the WebDev Arena leaderboard and a price half that of the market most expensive model. This answer carries a timestamp and any new release can change it.
Are open-weight models deployable for coding
Yes. On the same leaderboard, Kimi K3 Max scores 1674 and Qwen3.8 Max scores 1670, both close behind the leader. Both now carry a verified price. Qwen sits fifth today rather than fourth, without anything in its own specification changing: GPT-5.6 Sol entered the table above it. Qwen own score also rose the same day.
Why does this ranking differ from other lists
Because the criteria and the weight of each component are documented and open to review. Most lists do not explain their ordering, and where there is no explanation, any ordering is possible.
If you don't have the time to wire these tools into your own project, that work is part of
our app and custom software service