AI technical reference

This reference's change log

Every change to this reference is logged here with its date and its source document, from a new model release to a correction of ranking-criterion weights.

  1. new page

    The Claude Fable 5.1 record was logged; deployment date 1 September 2026. Audit result: the base price parameter is unchanged, $10 in and $50 out, matching Fable 5 exactly. The only parameter modified in Anthropic pricing table is the cache read, from one dollar to twenty-five cents. Anthropic flags this in a footnote under that same table as an exception in the pricing structure, because the cache multiplier is fixed at 0.1x input price across the entire Claude line and drops to 0.025x only for Fable 5.1 and Mythos 5.1. The benchmark jump is not uniform either: agentic scientific research moves from 24.7 to 52.6, terminal coding from 42.0 to 55.8, CursorBench only from 70.5 to 73.4. The model was not entered into any ranking table on this site; this is an operational decision: the coding rubric allocates 70 of 100 weight units to two Arena leaderboards, and this one-day-old model has not yet been scored on either, so entering it into the category would have produced only an unranked row under the table. Mythos 5.1 was added to the catalogue and Fable 5 was moved to legacy status.

    Source: www.anthropic.com

  2. new page Translation

    The translation category was deployed, the first of tranche two. Audit result: two parameters have come apart here. The model that registers your language by name in its own list is not procurable, and the model that is downloadable and runnable does not name your language. Command A Translate is the only model whose maker both classifies it as a translation model and lists all four of this site languages in one 23-language list; Llama 4, which carries a commercial licence, ships a closed twelve-language list with neither Persian nor Turkish on it. Licence-openness weight in this category is 30 of 100, a figure not repeated in any other category, because in Iran exactly that hosted API is the link that is down. One criterion was tested and then removed: context window, because min-max normalisation with a floor of 8,000 and a ceiling of ten million reported a thirty-two fold difference as 0.4 against 0.2.

    Source: docs.cohere.com

  3. new page Cohere

    The Cohere record was logged; a maker that until this deployment held only a few rows in the catalogue. Audit result: Cohere is the only maker in this reference that registers Persian item by item in a model language list, in Command A Translate and in Aya Expanse 32B, and all three access routes to that Persian are blocked separately. The translation model carries no published price, the open-weight model holds a CC BY-NC licence meaning commercial use is prohibited, and the model licensed under Apache 2.0 with commercial use permitted does not cover Persian at all. Separately, Cohere own commercial agreement registers Iran by name in its Restricted Location definition, differing from every other maker practice here: the others omit Iran from a list, Cohere enters it explicitly.

    Source: docs.cohere.com

  4. new page Amazon

    The Amazon record was logged. Two findings from this audit: first, the Nova specification table prints "+200" in the Supported Languages cell, and that cell footnote lists fifteen languages; Arabic and Turkish are among the fifteen and Persian is not. Same pattern as ElevenLabs, this time surfacing at footnote level. Second, Nova access capacity is a data-center list rather than a country list, and neither list registers a Middle East region; the nearest deployment point for a Nova model is Mumbai or Frankfurt. Nova pricing was also not recorded, because the Bedrock pricing page renders its figures in JavaScript.

    Source: docs.aws.amazon.com

  5. new page Voice

    The voice category was deployed, with six ranked models and one finding the entire page is built around: at ElevenLabs, Persian is registered only in the v3 language list. Flash v2.5 carries a 32-language list and Multilingual v2 carries 29, and Persian is on neither. Flash v2.5 is exactly the model users default toward, because it is faster, cheaper and carries eight times the character ceiling; as a result many inject Persian text into it, receive a defective output, and conclude ElevenLabs does not support Persian. In this category, the Persian criterion carries a weight of 35 of 100, because when the output is itself audio, a model lacking Persian scores zero rather than a somewhat lower score.

    Source: elevenlabs.io

  6. new page ElevenLabs

    The ElevenLabs record was logged. Apart from the Persian parameter, there is a condition on their own pricing page that is rarely documented: the free plan carries neither a commercial licence nor voice-cloning capability. So an entire project can be tested for free, the audio output extracted, and only then does it surface that no right existed to use that output in paid work. The commercial licence activates on the Starter plan at $6.

    Source: elevenlabs.io

  7. new page ElevenLabs v3

    The ElevenLabs v3 record was logged, the model that took first rank in the voice table. A constraint to know before a project starts: its character ceiling is 5,000, an eighth of Flash v2.5. So the model that reads Persian is the one that requires long text to be split into separate segments in the pipeline, with those segments stitched back together manually.

    Source: elevenlabs.io

  8. new page

    Records for Llama 4 Scout and Mistral Large 3 were logged, two models their maker pages could not link to directly. Scout carries the largest context window recorded in this reference, ten million tokens, ten times the nearest rival. Large 3 is the largest open-weight model recorded here, 675 billion parameters under an Apache 2.0 licence. Both records document the same technical point, rarely stated: active parameter count determines compute cost, and total parameter count determines required memory footprint; anyone reading only the first figure provisions the wrong hardware.

    Source: huggingface.co

  9. new page Mistral AI

    Records for Mistral and Meta were logged, two makers that until the prior deployment carried only data-level entries. Mistral audit result: the term "open" carries four distinct meanings at this company, and one of them prohibits commercial use. Large 3, Small 4 and Ministral 3 hold an Apache 2.0 licence, Medium 3.5 holds a modified MIT licence, and Voxtral TTS holds a CC BY-NC 4.0 licence. Anyone who assumed "Mistral means open" and built their product voice on Voxtral is in licence violation. Also, Le Chat and La Plateforme no longer appear in their current documentation; Vibe and Studio have replaced them.

    Source: docs.mistral.ai

  10. specification correction Meta

    Meta lists exactly twelve languages for Llama 4 and Persian is not registered among them. Arabic is registered. The list is closed and carries no "including" qualifier, so unlike the Mistral case, the absence of Persian here is not an inference but a documented fact. Two more parameters were recorded: both models carry an August 2024 knowledge cutoff, the oldest figure among the active models in this reference, and Scout at 109 billion parameters carries a ten-million-token window while Maverick at 400 billion carries a one-million-token window.

    Source: huggingface.co

  11. new page

    A full model list was added to this reference, along with ninety-two newly logged entities, including four makers that had no presence in this reference until this deployment: Mistral, Cohere, Amazon and Meta. Catalogue volume moved from eighteen models to eighty-nine, and every parameter in it was extracted directly from the maker own page. These new models were deliberately not entered into any category, meaning they affected no ranking table: injecting thirty new rows into six published rankings in one operation would have invalidated the prose on those pages the instant it ran.

  12. price change

    DeepSeek V4 Flash pricing increased, and this is the only pricing parameter in this reference to have moved upward to date. On 12 August the rate was $0.14 in and $0.28 out; its own pricing page now carries two columns, $0.22 and $0.66 off peak and $0.44 and $1.32 at peak. We record the peak column, because DeepSeek peak window falls squarely in the middle of the Iranian working day. DeepSeek itself lost no position and remains the cheapest row, but because its price lead narrowed, three other rows moved up in the coding table: Qwen from fifth to third, Kimi from third to fourth and GPT from fourth to fifth, none of which changed any parameter of their own.

    Source: api-docs.deepseek.com

  13. specification correction

    Sora exited the "no published data" state. This section had decided twice to keep Sora unrecorded, because openai.com returns a 403 to our server and its developer surface did not cover Sora either. It now does: the model page and the price list both respond, so the id, resolutions, input and output modalities, and per-second rate are now recorded. Sora 2 Pro still carries no recorded price, because OpenAI publishes its rate as a range, and a range is not a recordable figure.

    Source: developers.openai.com

  14. new page Veo

    The Veo family record was logged, documenting a contradiction inside Google own documentation: the model table flags Veo 3 and Veo 2 as Stable, and the pricing page on the same infrastructure declares both deprecated with a shutdown date of 30 June 2026, a date already past at the time this record was logged. The only version worth deploying, Veo 3.1, carries the Preview flag. This record does not claim the models were disabled; it claims the Stable label in this family carries no signal about which one is correct to build on.

    Source: ai.google.dev

  15. new page Seedance

    The Seedance family record was logged, and the resolution regression in this family now carries its third and firmest piece of evidence: on the Runway pricing list, Seedance 2.0 registers four resolution rows reaching up to 4K, while Seedance 2.5 registers only two, 480p and 720p. A pricing table is not a marketing surface. The family timeline is extracted from ByteDance own model ids, which carry the deployment date embedded in their structure.

    Source: docs.dev.runwayml.com

  16. new page Runway

    Records for Runway, ByteDance and Kuaishou were logged, closing the remaining incomplete entries in this section. Runway audit result: its own Gen-4.5 model is priced at twelve credits per second, and the Seedance it resells runs a hundred and fifty credits at 4K, so the most expensive video output on this platform costs twelve times its own cheapest model. Runway also carries a silent-audio rate for Veo 3.1 that does not exist at all on Google own pricing page.

    Source: docs.dev.runwayml.com

  17. new page Music generation

    The music category was deployed, with five candidates and four of them ranked. One column in this table is not recorded in any other comparison in this reference: output file delivery. The parameter "which produces higher audio quality" has a clear answer and Suno wins it, but the second parameter is whether the final file actually reaches the user, and that is where the ranking order shifts. This is the only category here whose quality figures are not extracted from Arena, because Arena carries no music leaderboard.

    Source: artificialanalysis.ai

  18. new page Udio

    The Udio record was logged, and one clause in their own help center is the technical reason for logging it: since 29 October 2025, coinciding with the Universal agreement, audio, video and stem download capability has been disabled. So the paid subscription generates songs but delivers no file to the user. No dollar price is available, because udio.com renders its entire content in JavaScript and the page body returns empty to curl, so that column was recorded empty in the table.

    Source: help.udio.com

  19. new page Suno

    The Suno record was logged and it took the top of the music table: a score of 1171 on vocals and 1188 on instrumental, and the longest output at eight minutes. A parameter buyers typically miss, and it is documented in this record: credits carry over neither from one day to the next nor from one month to the next, and the download cap is a separate ceiling from the credit cap. The Pro plan allocates 2,500 credits and twenty downloads a month.

    Source: suno.com

  20. new page Chat and assistants

    The chat category was deployed, with nine candidates, all ranked. The architectural decision documented in this record: this table ranks the model, not the application, because all five of its columns are figures published only about a model. The cost of this decision is that ChatGPT, the Claude app and the Gemini app carry no independent row.

    Source: arena.ai

  21. new page

    GPT-5.6 Sol was added to the dataset as a candidate in both the chat and coding tables: $5 and $30, a 1.05M-token window, 128K max output. This record also corrects a prior assumption of ours. It had previously been recorded that nothing from OpenAI is extractable; the accurate statement is that openai.com is not extractable while developers.openai.com is. So the model can be ranked while the app cannot be documented. This model has no dedicated page, because a complete page is impossible without addressing ChatGPT.

    Source: developers.openai.com

  22. ranking shift Coding

    Two data corrections restructured the parameters in the coding table. First, the text leaderboard was re-extracted in full, and the GLM-5.2 Max row (1470) and the DeepSeek V4 Flash row (1435) turned out to have been present there from the start, with an earlier extraction run having covered only the top of the page. Second, GPT-5.6 Sol was entered as an eighth candidate. Scores: Opus 5 moved from 76.0 to 77.5, Kimi from 49.9 to 52.4, Qwen from 49.3 to 51.3, GLM from 20.1 to 25.0, Grok from 10.8 to 15.6. The only order change: GPT landed between Kimi and Qwen.

  23. new page

    The xAI record, the Grok 4 family and Grok 4.5, was added to the dataset. The parameter most worth recording: from 200,000 tokens up, every token in that request is billed at double, not just the tokens above the threshold. And the newest member of this family carries half the context window of its predecessor, at more than double the output rate.

    Source: docs.x.ai

  24. ranking shift Coding

    With Grok 4.5 entered as a seventh candidate, the score of every row in the coding table rose without any model changing a parameter of its own: Opus 5 from 64.1 to 76.0, Fable 5 from 46.1 to 60.6, Kimi from 44.1 to 49.9, Qwen from 34.6 to 49.3, DeepSeek from 20.0 to 38.1 and GLM from 2.3 to 20.1. The technical cause is relative normalisation: Grok became the floor of three columns, and a column floor scores zero. No row changed position.

  25. new page

    The Fable 5 record was logged. The parameter that surfaced the sharpest inconsistency: the knowledge cutoff of Anthropic priciest widely released model is January 2026, while Opus 5, at half the price, reads through May 2026. The two are 64 points apart on the code leaderboard and 12 apart on the text one, in opposite directions.

    Source: arena.ai

  26. new page

    The Kimi K3 Max record was logged, with the technical note that K3 Max is not an independent model in the Moonshot catalogue. Its context window is 1,048,576 tokens, two to the twentieth power, and that sub-five-percent difference from the rest hands it the entire ten-point weight of the context-window column. This is documented in the record, because those ten points do not represent real advantage.

    Source: platform.kimi.ai

  27. new page

    The GLM-5.2 Max record was logged, and it carries the most solid evidence for the claim that max is a parameter, not a model: Z.ai own documentation names the model glm-5.2 and records max as a value of the reasoning_effort parameter. Its total table score is 2.3, with fifteen of a hundred weight units lacking sourced evidence.

    Source: docs.z.ai

  28. new page

    The DeepSeek V4 Flash record was logged. Output rate at $0.28, close to ninety times cheaper than Opus 5, and a 384,000-token output ceiling, three times the Anthropic models. Its total table score is 20 and that entire score is extracted from a single column; it scores zero in the forty-weight leaderboard column.

    Source: api-docs.deepseek.com

  29. new page

    The Gemini 3.1 Pro record was logged, including an Iran parameter extracted from the Gemini API available-regions list, in which Iran is not registered. It is the only candidate the coding table cannot rank: Google publishes no context window for it and its web-dev leaderboard row has not been extracted. Its pricing structure is also tiered, with the tier boundary at 200,000 tokens.

    Source: ai.google.dev

  30. specification correction Coding

    Three clauses in the coding record no longer matched the current state and were corrected in all four languages. GLM-5.2 Max and DeepSeek V4 Flash now carry a rank, and the only unranked candidate is Gemini 3.1 Pro; of the four Chinese models present in the table, only Qwen3.8 Max still lacks a verified price. No criterion weight and no rank was changed.

    Source: arena.ai

  31. new page Video generation

    The video category was deployed with four models, and ranking runs on the version, not the product, because one model is resold across several tools. Criterion weights: clip length 30, resolution 25, reference inputs 20, scene control 15, Iran 10. Seedance 2.5 took first rank. Generation time was not entered as a criterion because no vendor publishes that parameter, and price was excluded for the same reason: Kling is absent from the only dollar-denominated list available.

    Source: www.volcengine.com

  32. new page

    Seedance 2.5 was recorded: up to 30 seconds per generation pass and up to 50 reference inputs, comprising 30 images, 10 videos and 10 audio clips. A parameter absent from the product page and recorded only in the API documentation: this model resolution ceiling is 720p, where Seedance 2.0 reached 4K.

    Source: www.volcengine.com

  33. new page

    Kling VIDEO 3.0 was recorded: a three-to-fifteen-second range, native 4K since 23 April, and native audio across five speech languages. Persian is not registered among the five, and for a Persian-speaking user this is the most consequential parameter in the record.

    Source: kling.ai

  34. new page Kling

    The Kling record was logged. Five plans with prices extracted from their own credit guide, along with the note that advertised prices are first-cycle subscription rates: the Standard plan is purchased at $6.99 and renews at $8.80 from the following month.

    Source: kling.ai

  35. new page

    Veo 3.1 was recorded: a four, six or eight second range, audio permanently enabled, and up to three reference images. 4K and 1080p are available only on the eight-second clip, and extended output is capped at 720p.

    Source: ai.google.dev

  36. new page Runway Gen-4.5

    Runway Gen-4.5 was recorded, with figures extracted from their own OpenAPI schema: a two-to-ten-second range, a 720p ceiling, no audio parameter. No reference-input ceiling is published, so that cell remained unverified and cost it twenty weight units.

    Source: docs.dev.runwayml.com

  37. new page Studying and learning

    The studying category was deployed with four candidates and, unlike coding, ranking runs on the product rather than the model version. Criterion weights: processing output formats 35, source capacity 20, workspace 20, output languages 10, Iran 15. Gemini Notebook took first rank. ChatGPT is entered in the table without a rank, because OpenAI pages return a 403 to our server.

    Source: support.google.com

  38. new page Gemini Notebook

    Gemini Notebook was recorded, and the act of recording it surfaced that Google has dropped the NotebookLM name: both the official page and the help center now use Gemini Notebook. Free-plan specifications: 100 notebooks, 50 sources per notebook, nine output formats, and Persian registered on the output-language list.

    Source: notebooklm.google

  39. new page Gemini

    The Gemini record was logged with a focus on Guided Learning mode. The ten-files-per-prompt ceiling and the seventy-language coverage are extracted from Google own pages, as is the functional boundary with Gemini Notebook: Studio panel outputs cannot be generated within Gemini.

    Source: blog.google

  40. new page Claude

    The Claude record was logged. Learning mode was confirmed from Anthropic own education page, and the Pro plan price ($20 monthly, $17 on the annual plan) was extracted from the pricing page. Project capacity carries no published figure, so that column is recorded empty in the studying table.

    Source: claude.com

  41. new page Coding

    The coding category was deployed with six candidates and sourced evidence. The primary criterion is the WebDev Arena leaderboard at weight forty, and Opus 5 took first rank at 1692. No SWE-bench figure was entered, because that leaderboard could not be extracted directly.

    Source: arena.ai

  42. new version Claude Opus 5

    Opus 5 (deployment date 24 July 2026) was recorded with full specifications at $5 and $25. The Frontier-Bench and ARC-AGI 3 claims in the vendor announcement are stated in relative terms and carry no absolute figure, so they were not entered into the table.

    Source: www.anthropic.com