Voice clone comparison

Best voice clone in 2026: ElevenLabs vs Cartesia vs MiniMax

Searching for the best voice clone usually returns vendor benchmarks and affiliate listicles. This comparison is built the other way round: from what people actually say after shipping ElevenLabs, Cartesia and MiniMax in production, cross-checked against two blind listening leaderboards and every provider price we could verify on a provider-owned page.

Best voice clone at a glance

Three engines dominate the best voice clone conversation in 2026, and each wins a different category. The final column is what the same engine costs through Cheap.dev, where cloning is a one-time charge and speech is billed per ten seconds of output audio instead of per character.

CategoryBest voice clone pickWhy users choose itClone fee on Cheap.dev
Best overall realismElevenLabsTop blind voteLeads the same-voice blind listening test; the long-standing quality reference$2.00
Best for real-time agentsCartesia SonicFastest time to first audio; currently #1 on the Artificial Analysis Speech Arena$1.00
Best value and multilingualMiniMax SpeechStrongest reported results for Chinese, Japanese and other tonal languages, and for long passages$2.00

Clone fees are the live Cheap.dev one-time charge per cloned voice, published on the Cheap.dev pricing page. Leaderboard positions move; both boards are linked in the sources below so you can re-check them.

How we judged the best voice clone

Almost every page that ranks for “best voice clone” is written by a provider or an affiliate. We ignored those and used four kinds of evidence instead: unpaid practitioner comments on Hacker News, customer reviews on Trustpilot, open bug reports on the support forums of platforms that resell these engines, and two blind listening benchmarks that publish their method. Every claim below links to where it came from.

One theme showed up immediately, and it is the single most useful finding in this comparison: almost nobody abandons a voice cloning provider because of how it sounds. They leave because of billing surprises, because prosody drifts across a long script, or because a locale is handled badly.

ElevenLabs: the best voice clone for realism, the worst for billing

What users praise

ElevenLabs has been the quality reference since 2023 and still is under the strictest test available. On the Vapi Humanness Index, which clones one conversational voice onto every model so listeners judge the model rather than its demo reel, Eleven v3 leads with a humanness score of 97 against a human baseline of 100.

Practitioner comments match that. One Hacker News commenter called ElevenLabs “the highest quality models IMO” for voice quality, emotional awareness and inflection. Another, surveying the open-source field, wrote that GPT-SoVITS, StyleTTS2 and RVCv2 were still the open-source state of the art yet remained “really far behind Elevenlabs’ offerings”. A Vietnamese developer team documenting their English voice-over pipeline described the cloned result as having “more natural pauses; better rhythm and breathing; a speaking style that closely matched my own.”

What users complain about

Two complaints repeat far more than any other. The first is price. Hacker News comments include “Their pricing is ridiculous $200/1M is way too expensive”, “The Elevenlabs pricing to me makes it completely useless for audiobooks”, and a blunt value verdict: “if we put the price to punch ratio of kokoro at 100, elevenlabs is probably a 35”.

The second is credit expiry. Trustpilot reviewers report that unused credits, including credits that already rolled over from a paid month, are wiped on cancellation or downgrade, with one reviewer describing the loss of 400,000 credits after downgrading. Refund outcomes are genuinely bimodal in the same review set: some customers were refunded within minutes, others spent weeks chasing it.

There is also a newer, model-specific complaint. Eleven v3 audio tags fire inconsistently, with users reporting that only roughly one attempt in six delivers the tagged effect, and that short prompts make it worse. The same practitioner writeup found long-form drift: past about 1,000 characters the output showed “faster reading, uneven rhythm, or a drop in tone.”

Cartesia: the best voice clone for latency-bound voice agents

What users praise

Cartesia is bought for one number, time to first byte, and it delivers on it. As of August 2026 Sonic 3.6 also took the top position on both Artificial Analysis Speech Arena boards, which its team announced and independent coverage confirmed. Cartesia also ships an on-device build; an engineer from the company described it on Hacker News as retaining the full model capabilities while running locally.

What users complain about

The complaints are never about speed. A developer running Sonic in a video pipeline put the real constraint plainly: “the thing that actually matters for content creation isnt raw speed - its whether you can get consistent emotional delivery” across scenes.

The clearest documented defect is locale handling. An open thread on the Retell AI community forum reports Sonic 3.5 pronouncing Spanish words such as “Jornada” and “virtual” with an English accent, described as “unnatural and confusing to native Latin American Spanish speakers.” The cause was a generic es language code producing neutral pronunciation; a per-agent es-MX override was escalated to engineering and, as of the last message in that thread, had not shipped.

Cartesia also expects you to hand-tune pronunciation, a real operational cost that a per-second latency figure never shows.

MiniMax: the best voice clone for price, long passages and tonal languages

What users praise

MiniMax attracts the most concrete “we shipped this and stayed” reports. A developer building Japanese tutoring content wrote on Hacker News: “Minimax's new model is quite good. We use their voices for some of our Japanese tutors. The pitch accent is almost perfect. There are incorrect reading or Chinese readings occasionally, but you can tell when that happens due to the furigana being different.” On the Vapi Humanness Index, MiniMax appears twice in the upper rows, with Speech 2.8 at 91 and Speech 2 HD at 89, at noticeably lower latency than the model ranked above them.

The most useful account is the Meta Box team's migration writeup. They left ElevenLabs when credit reductions turned their Starter plan from “several videos per month” into “one or two”, found MiniMax quality “slightly more natural”, and stayed because of stability on long input: no degradation generating 3,000 to 4,000 characters in a single run, against ElevenLabs faltering after roughly 1,000. Their honest caveat is worth repeating: both engines performed poorly in Vietnamese.

What users complain about

MiniMax complaints are commercial and operational rather than acoustic: credit-based pricing that has to be actively monitored, monthly credits that expire at the end of each calendar month with no refund on mid-period cancellation, a web interface that new users find cluttered, a thinner curated voice library, and default API rate limits that require contacting the business team to raise.

The two leaderboards disagree, and that is the finding

Any page that cites one leaderboard to crown the best voice clone is selling something. The two most-cited boards return different winners because they test different things.

BenchmarkMethodResult as of August 2026
Artificial Analysis Speech ArenaElo from human preference votes; providers present their own voicesCartesia Sonic 3.6 ranked #1 on both the Provider Voice and Controlled Voice boards
Vapi Humanness IndexBlind same-voice battles; one cloned voice is applied to every model, with a real human baseline of 100ElevenLabs Eleven v3 leads at 97; MiniMax Speech 2.8 at 91; Cartesia's best model does not appear in the leading rows

The gap is methodological, not random. When every model is forced to speak with the same cloned voice, the thing being measured is exactly voice cloning fidelity, and the ordering changes. If your use case is cloning a specific person, weight the same-voice test more heavily. If you are picking a stock voice, the provider-voice board is closer to what you will ship.

Best voice clone pricing compared

Every number below comes from a provider-owned page, linked on the number itself. Derived figures are labelled as derived.

ProviderPublished configurationClone feeSpeech price
Cheap.devOursElevenLabs engine, pay as you go, no plan$2.00 once$0.03 / 10 sec
Cheap.devOursMiniMax engine, pay as you go, no plan$2.00 once$0.02 / 10 sec
Cheap.devOursCartesia engine, pay as you go, no plan$1.00 once$0.02 / 10 sec
ElevenLabsStarter plan, instant voice cloning, 30k credits$6 / mo1 credit / char
ElevenLabsCreator plan, professional voice cloning, 121k credits$22 / mo1 credit / char
CartesiaPro plan, instant voice cloning, 100k credits$5 / mo$0.06 / agent min
CartesiaStartup plan, 2 professional clone slots, 1.25M credits$49 / mo$0.06 / agent min
MiniMaxPay as you go, rapid voice cloning, speech-2.8-turbo$1.50 / voice$60 / M chars
MiniMaxPay as you go, voice design, speech-2.8-hd$3.00 / voice$100 / M chars

Prices exclude taxes and enterprise discounts. Cheap.dev bills speech by output audio, so ten seconds of speech costs the same regardless of how many characters produced it; ElevenLabs and MiniMax bill by character, and Cartesia bills plan credits whose character conversion is documented behind a login rather than on the public pricing page.

What that means per minute

Normalized to a minute of speech, Cheap.dev charges $0.12 for Cartesia or MiniMax and $0.18 for ElevenLabs, directly from the per-ten-second rates above, with no monthly plan. Buying direct can be cheaper per minute at volume: assuming roughly 900 characters per spoken minute, MiniMax pay-as-you-go works out near $0.05 per minute on turbo and $0.09 on HD, and an ElevenLabs Creator plan works out near $0.16 per minute if you consume the whole allowance. Those two per-minute figures are derived, not published, and the ElevenLabs one only holds while the subscription is active, because the allowance does not survive cancellation.

Which is the best voice clone for your use case?

Run all three behind one key

The practical answer to “what is the best voice clone” is usually “test them on your own script”, and the thing that stops teams doing that is three separate accounts with three separate subscriptions. Cheap.dev exposes the ElevenLabs, Cartesia and MiniMax engines through one pay-as-you-go API key: cloning is a one-time charge per voice, speech is billed per ten seconds of output, and nothing expires at the end of the month. You can browse every voice and audio endpoint in the model catalog alongside the video and image models.

Best voice clone FAQ

What is the best voice clone tool in 2026?

There is no single winner. ElevenLabs leads the same-voice blind listening test, Cartesia Sonic 3.6 leads the Artificial Analysis Speech Arena and is fastest to first audio, and MiniMax is the value pick with the strongest reports for tonal languages and long passages.

What is the best voice clone for a real-time voice agent?

Cartesia Sonic, because time to first byte decides whether a caller hears a pause. The trade-off users report is emotional consistency across long scripts and weaker non-English locale handling, not speed.

What do users complain about most?

Billing, not audio quality. The most repeated ElevenLabs complaint is losing paid, unused credits on cancellation or downgrade. Cartesia users report locale pronunciation bugs. MiniMax users report credits expiring at the end of each calendar month.

How can I use these models without a monthly subscription?

Through Cheap.dev, where all three engines are pay as you go: $1.00 one-time to clone a Cartesia voice, $2.00 for MiniMax or ElevenLabs, and per-ten-second speech billing with no expiring allowance.

How much reference audio does a voice clone need?

Instant or rapid cloning works from a short sample on all three engines, and MiniMax advertises high-precision cloning from about ten seconds of audio. Professional cloning, which needs far more audio and a training pass, sits behind higher paid tiers at both ElevenLabs and Cartesia.

Methodology and sources

Provider prices were checked on provider-owned pages on August 23, 2026, and Cheap.dev prices come from its live catalog on the same date, using the same source-linking rule as our other API price comparisons. User feedback was taken from unpaid practitioner comments, customer reviews and open support threads; nothing in the quality sections comes from a vendor blog or an affiliate ranking. Leaderboard standings move, so both boards are linked rather than reproduced as settled fact.

Clone a voice on Cheap.dev

Compare ElevenLabs, Cartesia and MiniMax on your own script behind one pay-as-you-go key, with no subscription and no expiring credits.