Job

Best AI chatbot

This table instruments the model layer, not the product layer. The primary criterion is the Arena text leaderboard: Claude Fable 5 holds first place among 389 models at a score of 1506, and Opus 5 runs twelve points lower at 1494 for half the operating cost. Because the unit of measurement is the model rather than the product, ChatGPT, a product layer, carries no row in this table.

  • nine models, each with a documented rank
  • every criterion weight logged on this page
  • unit of measurement: the model, not the product

Last checked: This is a reference page. It is re-checked against the vendor sources and updated when a new version ships. Change log

Where each tool's score comes from

The ring below is the same set of weights printed in the table headers further down.

100Weights
  • Performance on the general text leaderboard 45
  • Score-per-dollar yield 20
  • Context window capacity 15
  • Tooling infrastructure maturity 10
  • Access status from Iran 10

We chose these weights, and that is the only judgment call in the table. Weight them differently and the order changes.

Today ranking, built from a leaderboard with documented score, date, and vote count

No candidate is left without a rank below this table. An empty cell indicates the vendor has not published a figure, not a score of zero.

Best AI chatbot
Rank Tool Score Performance on the general text leaderboard 45 Score-per-dollar yield 20 Context window capacity 15 Tooling infrastructure maturity 10 Access status from Iran 10
1 Claude Fable 5 Current pick 78.6 1,506 Source: arena.ai 30.1 Source: arena.ai 1 million tokens Source: platform.claude.com 4/4 blocked / no working route
2 Claude Opus 5 71.6 1,494 Source: arena.ai 59.8 Source: arena.ai 1 million tokens Source: platform.claude.com 4/4 blocked / no working route
3 GPT-5.6 Sol 54.5 1,481 Source: arena.ai 49.4 Source: arena.ai 1 million tokens Source: developers.openai.com 4/4 not verified
4 Qwen3.8 Max 52.6 1,490 Source: arena.ai 248.3 Source: arena.ai 1 million tokens Source: www.alibabacloud.com not verified not verified
5 Kimi K3 Max 49.2 1,487 Source: arena.ai 99.1 Source: arena.ai 1 million tokens Source: platform.kimi.ai 1/4 not verified
6 Gemini 3.1 Pro 44.1 1,486 Source: arena.ai 123.8 Source: arena.ai not verified not verified blocked / no working route
7 GLM-5.2 Max 41.6 1,470 Source: arena.ai 334.1 Source: arena.ai 1 million tokens Source: docs.z.ai 1/4 not verified
8 DeepSeek V4 Flash 33.6 1,435 Source: arena.ai 1,087.1 Source: arena.ai 1 million tokens Source: api-docs.deepseek.com 1/4 not verified
9 Grok 4.5 32.3 1,469 Source: arena.ai 244.8 Source: arena.ai 500k tokens Source: docs.x.ai 3/4 not verified

An empty cell means we could not verify that number, not that the tool scored zero.

Behind each number

Every judged score in this table carries a written reason, and the Iran column says how it was checked. Those two are open here. The measurement trail behind each number, which row of which leaderboard and on how many votes, opens under the model it belongs to.

1 Claude Fable 5

Tooling infrastructure maturity 4/4 Built on the same deployment infrastructure as Opus 5, with full coverage across all five delivery routes: Anthropic's own API, Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry.

Access status from Iran blocked / no working route Iran is not registered on Anthropic supported-countries list. Extracted from Anthropic own policy document, not from a network measurement.

Measurement trail (2)

Performance on the general text leaderboard 1,506 Source data, the Arena text leaderboard, updated 11 August 2026, derived from 7,765,220 votes

Score-per-dollar yield 30.1 Arena (Text) score divided by the price per million output tokens

2 Claude Opus 5

Tooling infrastructure maturity 4/4 The integration surface is complete: API, official app, command-line tool, agent SDK, batch-execution pipeline, and deployment across the infrastructure of three major clouds

Access status from Iran blocked / no working route Iran is registered on neither of Anthropic two supported-countries lists, at the API layer or the claude.ai layer. That status is extracted from Anthropic own policy document, not from a direct network measurement; an independent network test has not been run yet, and its result will be logged here once it is.

Measurement trail (2)

Performance on the general text leaderboard 1,494 Source data, the claude-opus-5-high configuration row: rank seven, derived from 7,765,220 votes

Score-per-dollar yield 59.8 Arena (Text) score divided by the price per million output tokens

3 GPT-5.6 Sol

Tooling infrastructure maturity 4/4 The top tier is confirmed through the maker own technical documentation: an API, SDKs, an official CLI, the Codex editor extension, an Agents SDK, a batch execution lane, and a first-party deployment guide for Amazon Bedrock that records the id openai.gpt-5.6-sol. This is the one column that reached full completion for this model without openai.com ever becoming reachable.

Measurement trail (2)

Performance on the general text leaderboard 1,481 The gpt-5.6-sol-xhigh entry on the Arena text leaderboard, position 18 of 389 recorded models, per the 11 August 2026 update and 15,226 measurement votes on this entry

Score-per-dollar yield 49.4 Arena (Text) score divided by the price per million output tokens

4 Qwen3.8 Max

Measurement trail (2)

Performance on the general text leaderboard 1,490 Extracted from the Arena text leaderboard, logged update 11 August 2026, 7,765,220 votes

Score-per-dollar yield 248.3 Arena (Text) score divided by the price per million output tokens

5 Kimi K3 Max

Tooling infrastructure maturity 1/4 The integration surface includes an OpenAI-compatible API plus one official client. The platform documentation lists no official command-line tool or editor extension, so no higher score is assigned.

Measurement trail (2)

Performance on the general text leaderboard 1,487 Extracted from the Arena text leaderboard, per the 11 August 2026 update and a total of 7,765,220 measurement votes

Score-per-dollar yield 99.1 Arena (Text) score divided by the price per million output tokens

6 Gemini 3.1 Pro

Access status from Iran blocked / no working route The Gemini API available-regions specification defines the set of countries where the API is operational, and Iran is not on that list. This status was extracted from the specification itself rather than a network measurement; an independent test has not yet been run.

Measurement trail (2)

Performance on the general text leaderboard 1,486 Extracted from the Arena text leaderboard, updated 11 August 2026, an aggregate of 7,765,220 recorded votes

Score-per-dollar yield 123.8 Arena (Text) score divided by the price per million output tokens

7 GLM-5.2 Max

Tooling infrastructure maturity 1/4 The integration surface includes a documented API plus one official client. No official command-line tool or editor extension is recorded in the technical specification.

Measurement trail (2)

Performance on the general text leaderboard 1,470 Position 33 of 389 recorded models on the Arena text leaderboard, per the 11 August 2026 update and 26,703 measurement votes on this entry

Score-per-dollar yield 334.1 Arena (Text) score divided by the price per million output tokens

8 DeepSeek V4 Flash

Tooling infrastructure maturity 1/4 Integration runs through a documented API and one official client. No official agent tooling or editor extension is listed.

Measurement trail (2)

Performance on the general text leaderboard 1,435 Rank 83 out of 389 models on the Arena text leaderboard, the build updated 11 August 2026, with 49,111 recorded votes for this row

Score-per-dollar yield 1,087.1 Arena (Text) score divided by the price per million output tokens

9 Grok 4.5

Tooling infrastructure maturity 3/4 The integration surface includes an OpenAI-compatible API, official Python and JavaScript SDKs, and an official command-line agent named Grok Build. The top tier is out of reach because no first-party editor extension is built into the architecture, and the batch execution lane does not accept this specific model.

Measurement trail (2)

Performance on the general text leaderboard 1,469 Position 34 of 389 recorded models on the Arena text leaderboard, per the 11 August 2026 update and a total of 7,765,220 measurement votes

Score-per-dollar yield 244.8 Arena (Text) score divided by the price per million output tokens

How this ranking is calculated

Every criterion below has a weight and a source. Change a weight and the whole table recomputes. There is no hand-placed position anywhere in this hub.

Criterion Weight Evidence
Performance on the general text leaderboard 45 Arena (Text)Human preference in open-ended conversation. It is the closest available proxy for correctness in this use case, which is why it carries the largest weight allocation
Score-per-dollar yield 20 computed from two numbers aboveFor provisioning chat on an API, paying double the cost for one percent more score is a poor economic decision. For a product subscription, this column has no operational effect
Context window capacity 15 vendor stated specificationA long conversation loses its memory first and its cost rises second, once the window fills
Tooling infrastructure maturity 10 defined scale, with a written reason per assignmentA model with no official client, SDK, or cloud deployment path lags behind a slightly weaker model in daily operation
Access status from Iran 10 our access column, with its method statedWeighted at ten of a hundred because this axis has its own dedicated page; today three of nine rows carry evidence in this column

The layer architecture this table instruments

The unit of measurement in this table is the model, not the product. A model is a component the vendor publishes numbers about: a leaderboard score, a cost per million tokens, a context window capacity. The product is the layer above it: a client that swaps its underlying model without notice. Opus 5 simultaneously powers the infrastructure behind the Claude app, the API, and Claude Code; treat "Claude" as a single row and the same number gets counted again in the coding table. That is why the video category also ranks the model version rather than the tool.

The operational cost of that decision is explicit: ChatGPT, the Claude app, and the Gemini app carry no row in this table and are linked to their product pages instead. For the study category we made the inverse call, because Google documents NotebookLM operating limits as figures: fifty sources, five hundred thousand words. For a chat app, no vendor publishes a comparable specification.

Table inputs: five criteria and their weights

The weights are fixed and documented on this page: text leaderboard 45, score per dollar 20, context window 15, tooling maturity 10, access from Iran 10. Only ten of the hundred points sit on an ordinal scale. Any weight change logs a row in the changelog.

Two columns in this architecture depend on vendor data disclosure, and that dependency is what produces empty cells. Qwen3.8 Max and Gemini 3.1 Pro each carry two empty columns; that is a disclosure gap, not a model weakness, but it genuinely pulls the final score down.

Today output from the ranking architecture

Fable 5 holds first place on the Arena text leaderboard among 389 models at 1506, and it captures the entire forty-five-weight column. The same model loses the score-per-dollar column outright: output cost of $50 per million tokens, the most expensive row in the table, for a score-per-dollar ratio of 30.1, the floor of the pricing architecture.

Opus 5 sits twelve points lower and runs at half the operating cost. On a leaderboard whose entire top tier fits inside a roughly fifty-point band, a twelve-point gap is negligible. The optimal operating point for building on the API is therefore Opus 5; for a product subscription, the price column has no effect and the top of the table is the only criterion that matters.

One operational constraint has no column defined for it in this architecture: Fable 5 is the only model Anthropic itself flags as "slower" in its own table, because its adaptive thinking is always active with no lever to disable it. That latency is felt in real conversation, but the table does not measure it.

The boundary of this architecture: ChatGPT

ChatGPT is a product, so it was excluded from model rows under any configuration. The harder constraint is that it cannot even be documented: every path on openai.com returns a 403 status code to this server, re-tested on 12 August 2026, while that domain robots.txt does not forbid crawling.

One integration path stays open: developers.openai.com responds in full. GPT-5.6 Sol is listed there at $5 input and $30 output, a 1.05M-token context window, a 128K-token output ceiling, and a knowledge cutoff of 16 February 2026, and it carries a defined Arena row. On that basis, it occupies the third row of this table today.

Its WebDev row is registered exactly as gpt-5.6-sol-xhigh (codex-harness): the score belongs to the model inside OpenAI own dedicated coding agent, not the bare model. Cost and context window attach to the base model, and the score attaches to whatever configuration was actually tested. What remains unknown is which model powers the ChatGPT infrastructure today, what its plan pricing is, and whether it is reachable from Iran; all three live on the domain that does not respond. The OpenAI page documents this operational boundary in more detail.

Tolerance of the measurement method

The nine rows in this table fit inside a 71-point Elo band, from 1506 down to 1435. Min-max normalization expands that same 71-point band into a range from 78.6 to 29.1. As a result the score column overstates the real distance, and the last row is not "weak" — it is the last position in an architecture composed entirely of top-tier models.

A sharper example: GPT-5.6 Sol captures the context-window column at 1,050,000 tokens, with Kimi K3 Max close behind at 1,048,576. The 1,424-token gap, about a tenth of a percent, has no practical effect. Both absorb nearly the full weight of that column; but provisioning a model with a two-million-token window would drop both out of that column simultaneously, with nothing about their own specification having changed.

Two rows previously recorded as "unverified" were on this leaderboard from the start: GLM-5.2 Max at 33rd place with 1470, and DeepSeek V4 Flash at 83rd place with 1435. That was a gap in our measurement process, not in vendor disclosure. The deepseek-v4-pro row at 1458 on the same leaderboard belongs to a different model in the same family, not a configuration variant, so its figure was not borrowed.

When our top pick is not the right fit

  • If your operational unit of need is the product rather than the model, this table does not answer it; the Claude and Gemini product pages answer the same question at the correct layer.
  • If your workload is coding, the ordering here does not match the coding table, because that architecture runs on entirely different criteria.
  • If your workload is high-volume with short responses, disregard the first column and read score-per-dollar instead; DeepSeek holds first place there by a wide margin.
  • If you are reaching these services from Iran with no official access route, model ordering is the least significant variable in your problem.

Known weaknesses and limits

  • The primary criterion in this architecture is a human-preference leaderboard: it measures which model responses users prefer, not which model is technically more correct.
  • The entire table fits inside a 71-point Elo band, and normalization expands that into a range from 78.6 to 29.1, so the score column overstates the real distance between rows.
  • Six of the nine rows carry no evidence in the Iran column, because their vendors do not publish a country list; only Anthropic and Google make that list available.
  • ChatGPT, the Claude app, and the Gemini app carry no row in this table, because the unit of measurement in this architecture is the model, not the product.
  • Qwen3.8 Max and Gemini 3.1 Pro each carry two empty columns in their data profile. That is a disclosure gap on the vendor side, not a model weakness, but it genuinely pulls their final scores down.
  • DeepSeek, at an output cost of $0.28, stretches the score-per-dollar column so far that the fourfold gap between Opus 5 and Qwen in that column compresses to less than one point.

Technical verdict

For provisioning chat on an API, and if forced to a single choice: Opus 5, twelve points behind Fable 5 at half the operating cost. For a product subscription, the price column in this table has no effect and the top of the table is the only criterion. And if the question is which model answers your customers, the honest answer is none of them; that is a different operational task entirely.

Frequently raised questions, with documented answers

What is the best AI for chat and conversation

On the Arena text leaderboard, Claude Fable 5 holds first place today at a score of 1506. Opus 5 sits at 1494 for half the operating cost, and on a leaderboard whose entire top tier fits inside a fifty-point band, that twelve-point gap is barely felt in practice.

Why does ChatGPT carry no row in this ranking architecture

Because the unit of measurement in this table is the model, and ChatGPT is a product. Its underlying model is present: GPT-5.6 Sol occupies the third row. The app itself is not documented, because every path on openai.com returns a 403 status code to our server.

Why does this order differ from the coding table

Because the criteria architecture differs. There, the WebDev leaderboard carries a weight of forty; here, the general text leaderboard carries forty-five. A model stronger at coding does not necessarily produce more preferred conversational responses, and Opus 5 and Fable 5 swap positions across exactly these two tables.

Which of these work from Iran

None of them has an official access route. Anthropic and Google publish country lists and Iran is absent from both; the remaining vendors publish no list at all, leaving their columns empty. We neither sell accounts nor propose ways to bypass the restriction.

The top of this table is instrumented to answer you, not your customers. For that second task, we deploy a person rather than a bot, and that is exactly

our virtual assistant service
This page in the Seemerce family

Comparing chat assistants and what each one is genuinely good at.

One rule throughout: See fuses onto a word and makes a single new one, never two words side by side. Exactly as See and commerce became Seemerce.