Category

Best AI for voice generation

In today configuration, ElevenLabs v3 holds the top score in this table, because it is the only frontier model whose published language specification lists Persian while also supporting new-voice provisioning under a commercial licence. The failure mode that catches most users: the same vendor faster, cheaper model carries no Persian coverage at all.

  • Six models with traceable evidence
  • Persian criterion weighted 35 of 100
  • An Iran access column

Last checked: This is a reference page. It is re-checked against the vendor sources and updated when a new version ships. Change log

Where each tool's score comes from

The ring below is the same set of weights printed in the table headers further down.

100Weights
  • Persian coverage 35
  • Voice control 30
  • Language count 25
  • Access from Iran 10

We chose these weights, and that is the only judgment call in the table. Weight them differently and the order changes.

Today ranking, built from each vendor published language specification

Each model Persian coverage was read from the official language list the vendor publishes, not from a listening test. A model whose vendor publishes no language list at all is marked unverified in that column.

Best AI for voice generation
Rank Tool Score Persian coverage 35 Voice control 30 Language count 25 Access from Iran 10
1 ElevenLabs v3 Current pick 99.0 1/2 3/3 74 languages Source: elevenlabs.io blocked / no working route
2 Gemini 3.1 Flash TTS 80.0 1/2 1/3 77 languages Source: ai.google.dev blocked / no working route
3 GPT-4o mini TTS 55.0 1/2 1/3 not verified blocked / no working route
4 ElevenLabs Flash v2.5 49.4 0/2 3/3 32 languages Source: elevenlabs.io blocked / no working route
5 ElevenLabs Multilingual v2 48.3 0/2 3/3 29 languages Source: elevenlabs.io blocked / no working route
6 Amazon Nova Sonic 0.0 0/2 0/3 5 languages Source: docs.aws.amazon.com not verified

An empty cell means we could not verify that number, not that the tool scored zero.

Behind each number

Every judged score in this table carries a written reason, and the Iran column says how it was checked. Those two are open here. The measurement trail behind each number, which row of which leaderboard and on how many votes, opens under the model it belongs to.

1 ElevenLabs v3

Persian coverage 1/2 In ElevenLabs official specification, Persian is registered as fas within this model 74-language list. Because no dedicated voice for Persian is defined in this model voice architecture, it does not qualify for the level-two score.

Voice control 3/3 The voice-cloning subsystem operates from a single input sample, and the commercial-use license activates from the Creator subscription tier upward. The pricing specification defines instant cloning at the Starter tier and professional cloning at the Creator tier.

Access from Iran blocked / no working route ElevenLabs terms-of-use document names Iran explicitly in its export-control clause: the user must warrant they are not located in a country under US economic sanctions, and Iran is listed there. We extracted this constraint from the text of that clause itself, not from a direct network measurement.

2 Gemini 3.1 Flash TTS

Persian coverage 1/2 Persian is registered on the documented language list of Google's speech-generation specification. Google names no dedicated Persian voice and performs automatic input-language detection; on that basis this component does not reach the second scoring tier.

Voice control 1/3 The technical specification for this deployment covers thirty preconfigured voices with tone control delivered through a text instruction. Cloning a new voice from an audio sample is not documented in this specification.

Access from Iran blocked / no working route The Gemini API available-regions documentation specifies the countries within its operating envelope, and Iran does not fall within the declared operational scope. This status is drawn from the documentation rather than direct network measurement.

3 GPT-4o mini TTS

Persian coverage 1/2 Persian is explicitly recorded among the supported languages in the OpenAI text-to-speech technical documentation. No Persian-specific voice is defined in this specification, so tier two is out of reach.

Voice control 1/3 Thirteen preconfigured voices are available, along with tone control through a text instruction. Building a new voice from an audio sample is not built into this specification.

Access from Iran blocked / no working route Iran is not recorded on the supported-countries list in OpenAI developer technical documentation. This status is derived directly from that page, not from a network measurement.

4 ElevenLabs Flash v2.5

Persian coverage 0/2 This model specification does not cover Persian in its supported-language surface; the 32-language list in ElevenLabs official documentation includes Arabic and Turkish, but Persian is not registered in it.

Voice control 3/3 The voice-cloning subsystem operates from a single input sample, and the commercial-use license activates from the Creator subscription tier upward. The pricing specification defines instant cloning at the Starter tier and professional cloning at the Creator tier, with commercial authorization provisioned on the same tiers.

Access from Iran blocked / no working route ElevenLabs terms-of-use document names Iran explicitly in its export-control clause: the user must warrant they are not located in a country under US economic sanctions, and Iran is listed there. We extracted this constraint from the text of that clause itself, not from a direct network measurement.

5 ElevenLabs Multilingual v2

Persian coverage 0/2 This model language-coverage surface, per ElevenLabs official specification, does not register Persian within its 29-language list.

Voice control 3/3 The voice-cloning subsystem operates from a single input sample, and the commercial-use license activates from the Creator subscription tier upward. The pricing specification defines instant cloning at the Starter tier and professional cloning at the Creator tier, with commercial authorization provisioned on the same tiers.

Access from Iran blocked / no working route ElevenLabs terms-of-use document names Iran explicitly in its export-control clause: the user must warrant they are not located in a country under US economic sanctions, and Iran is listed there. We extracted this constraint from the text of that clause itself, not from a direct network measurement.

6 Amazon Nova Sonic

Persian coverage 0/2 Amazon specification for this unit records five languages: English US and UK, French, Italian, German and Spanish. Persian is not recorded among them.

Voice control 0/3 Preset voices provisioned by the maker. No path for constructing a new voice from a sample is stated in the Nova documentation.

Not enough verified data to rank

These are real tools we track, but more than half of their criteria have no verified number yet, so ranking them would be a guess.

How this ranking is calculated

Every criterion below has a weight and a source. Change a weight and the whole table recomputes. There is no hand-placed position anywhere in this hub.

Criterion Weight Evidence
Persian coverage 35 defined scale, with a written reason per assignmentIn this category the output itself carries a language, so a model without Persian coverage scores zero for a Persian workload, not slightly lower
Voice control 30 defined scale, with a written reason per assignmentVoice production requires a specific, deployable voice. A handful of preset voices is not equivalent to new-voice provisioning, and the commercial licence is part of this axis
Language count 25 vendor stated specificationCounted only where the vendor publishes the full list or states a definite figure
Access from Iran 10 our access column, with its method statedWeighted low because this axis has its own dedicated page under AI in Iran, though it is not zero

The common failure mode in this category

ElevenLabs runs three text-to-speech models in its stack. Persian is documented in the specification of exactly one of them.

v3 carries a 74-language list with Persian entered as fas. Flash v2.5 carries 32 languages, and Multilingual v2 carries 29, and Persian is on neither. The operational problem is that Flash v2.5 is the model users default to: it is the fastest at 75ms latency, its character ceiling is eight times v3 (40,000 against 5,000), and it is also the cheapest option.

As a result, a user who signs up, selects the fast model, feeds it Persian input and gets a broken output typically generalizes the failure to the entire ElevenLabs stack. The more accurate diagnosis is that this specific model is not provisioned for Persian.

Why the Persian criterion carries a weight of 35

In the coding category, output language is not part of the technical specification. Here the output is itself a voice, and a voice has a language. A model that covers seventy languages without Persian among them scores zero for a Persian workload, not slightly lower. No other weighting represents that fact honestly, and any table missing this axis is the wrong reference for a Persian-speaking audience.

Each model coverage level was read from the vendor own official language list, not from listening to sample output. So this column only reports what the vendor has formally declared, not accent quality. Judging accent quality requires a listening test that has not yet been run.

Three criteria deliberately excluded from the table

Latency. Only ElevenLabs publishes a figure (75ms for Flash, 280ms for the conversational v3). Google and OpenAI publish nothing. A criterion with data for two candidates out of six drags the rest down without a fair basis.

Price. ElevenLabs sells monthly credits, Google and OpenAI bill per token, and Mistral publishes model weights at no usage cost. These units do not convert onto one column. Any table that forces them into one is converting one figure into another on assumptions it never documents.

Preset voice count. Google ships thirty and OpenAI thirteen, but ElevenLabs supports building an entirely new voice. Once new-voice provisioning is possible, counting presets stops being a meaningful metric.

No candidate has reached level two

The Persian coverage scale has three levels, and level two requires the vendor to also publish a Persian-specific voice. No candidate meets that bar today. Google states that it auto-detects input language and names no separate Persian voice; OpenAI has not tied its thirteen voices to specific languages. So every model here reads Persian through a voice that was not engineered for Persian.

A contradiction that has to be stated plainly

The model at the top of this table names Iran directly in its own terms of use. The ElevenLabs export-control clause requires the user to warrant that they are not located in a country under US economic sanctions, and Iran is entered on that list.

We document this because it is the actual state of the infrastructure, and omitting it helps no one. What we do not document is a route around it. This reference does not sell accounts, does not broker intermediaries, and does not describe a method for bypassing sanctions. If your workload is Persian voice production and you have no official access path, this is a direct statement that no official access path exists.

One model with open access that still fails your deployment

Voxtral TTS from Mistral is the only model in this table with published weights, meaning it can be deployed on your own hardware with no country-list gate. Its licence is CC BY-NC 4.0, and the NC clause prohibits commercial use.

In a category whose entire function is commissioned voice production, that is precisely the deployment most users need. So the model that scores best on access is the least usable under licence. Mistral has also not published a language list for it, so its Persian column remains unverified and it carries no rank.

When our top pick is not the right fit

  • If your workload is English-only and low latency is the priority, Flash v2.5 from the same vendor at 75ms is the better pick; its lower rank in this table is entirely a function of missing Persian coverage.
  • If you need to self-host the model, none of the top three rows in this table is usable, and Voxtral is the option to evaluate, with its licence restriction factored in.
  • If your workload is speech-to-text rather than voice generation, this table is not the reference you need. Scribe and Voxtral Transcribe cover a different function.

Known weaknesses and limits

  • Persian coverage was read from the published language list, not from a listening test. Accent quality has not been assessed.
  • There is no price column in this table at all, because the three vendors billing units are not comparable on one scale.
  • Latency figures are published only for the ElevenLabs models, so latency is not part of the ranking.
  • The Iran access column was read from each vendor official policy page, not measured from a network connection inside Iran.

Technical verdict

If your workload is Persian and you need a specific, repeatable voice, ElevenLabs v3 is the only option today that covers both requirements. Set the model to v3 manually, because the faster default configuration is the one without Persian coverage.

Frequently raised questions, with documented answers

Does ElevenLabs support Persian

It depends on the specific model. v3 carries a 74-language list with Persian entered as fas. Flash v2.5 and Multilingual v2 carry 32 and 29 languages, and Persian is on neither. For a Persian workload you must set the model to v3 manually.

What is the best AI voice generator for Persian

By the table on this page, ElevenLabs v3. Three models list Persian in their specification (v3, Gemini Flash TTS and GPT-4o mini TTS), and of those three only v3 supports new-voice provisioning under a commercial licence.

How good is the Persian accent on these models

That data does not currently exist and we make no claim about it. This table only records whether the vendor lists Persian among its languages. Assessing accent quality requires a listening test that has not been run, and it will be documented here once it is.

Is there a free or open model for Persian voice

Voxtral TTS from Mistral ships open weights, but its licence is CC BY-NC 4.0 and prohibits commercial use. Mistral has also not published a language list for it, so whether it covers Persian is unverified.

We can produce the voice track for your video

Motion graphics and voice services