Retell AI vs Bland AI vs Vapi vs ElevenLabs


Every platform in this category publishes a latency number, and none of them measure the same thing.

ElevenLabs documents roughly 75ms, which covers the speech synthesis layer alone while full conversational latency runs far higher (Cekura platform review). An independent harness measured Retell at 1.96s per conversational turn on platform defaults (Cekura fixed-stack study). Both numbers were accurate. Neither helped a buyer, because neither said where the clock started and stopped.
So this comparison reports measured results from tests Retell did not run, names the boundary behind every number, and links the run reports so you can check the work.
Retell owns the turn-taking model rather than chaining public APIs, so response times stay consistent at concurrency and the agent handles barge-in, pauses and background noise on real telephony. Test it on your own accents and line quality during the pilot.
Retell AI publishes this article. We lead some of these measurements and lose others. The losses stay on the page.
Retell leads repeatable reliability at 75.61% pass³ and scores 5.00 out of 5 on interruption handling in the current independent cohort (Cekura Bench). ElevenLabs sounds best at 4.47 out of 5 for voice tone and clarity and answers fastest at 1.27s mean response. Vapi completes the highest share of the calls it connects, 97.56%, and connects the fewest, with 41 of 246 calls failing to connect. Bland has no independent benchmark result at all, so every Bland figure here comes from vendor documentation.
Cekura builds evaluation harnesses for voice agents and publishes per-run reports. Its current study deploys one matched agent across seven configurations, runs 82 scenarios three times each, and retains 246 calls per configuration (Cekura Bench).
The headline metric is pass³. A scenario counts only when all three runs pass, so one lucky call earns nothing. Platforms compared in this article appear in bold.
| Configuration | pass³ | Task completion | Calls with no provider or connection issue | Interruption /5 | Voice tone and clarity /5 | Mean response |
|---|---|---|---|---|---|---|
| Retell | 75.61% | 93.88% | 98.37% | 5.00 | 4.36 | 2.21s |
| LiveKit | 70.73% | 95.12% | 99.19% | 4.97 | 4.36 | 2.59s |
| ElevenLabs | 69.51% | 91.46% | 100.00% | 4.96 | 4.47 | 1.27s |
| GPT Realtime | 64.63% | 92.68% | 95.53% | 4.98 | 4.25 | 1.58s |
| Pipecat | 63.41% | 94.21% | 97.15% | 4.97 | 3.74 | 1.97s |
| Vapi | 59.76% | 97.56% | 82.93% | 4.73 | 4.08 | 3.08s |
| Gemini Live | 30.49% | 87.80% | 72.36% | 4.97 | not scored | 3.05s |
All figures come from Cekura Bench: 82 scenarios, three repeats, 246 retained calls per configuration. Bland does not appear in this cohort.
Read two of those columns together or you will misread them. Vapi posts the highest task completion in the field at 97.56%, and that rate covers the 205 calls that produced outcome evidence. The other 41 calls never connected, which is why its infrastructure column sits at 82.93% (Cekura Bench). A platform that finishes almost every call it answers, and answers fewer of them, carries a different risk from one that answers everything and finishes less.
A second Cekura experiment took a harder path. One byte-identical Medicare third-party marketing organization agent went onto six platforms, with 23 evaluators aimed at the moments regulated insurance calls break, three runs each, 414 calls in total (Cekura Experiment 02).
| Platform | Workflow pass³ | Strict end-to-end pass³ |
|---|---|---|
| Retell | 95.7% (22/23) | 95.7% (22/23) |
| ElevenLabs | 91.3% (21/23) | 91.3% (21/23) |
| Pipecat | 82.6% (19/23) | 73.9% (17/23) |
| LiveKit | 73.9% (17/23) | 73.9% (17/23) |
| Synthflow | 69.6% (16/23) | 8.7% (2/23) |
| Vapi | 65.2% (15/23) | 65.2% (15/23) |
Source: Cekura Experiment 02. You can also read our write-up of the Medicare benchmark.
The Synthflow row shows why two scores exist. Workflow accuracy asks whether the agent chose the right action and called its tools correctly. Strict end-to-end accuracy also asks whether the caller could finish. During tool use, Synthflow callers sat through long silences with no signal that the agent was still working, so task accuracy held at 69.6% while the strict score fell to 8.7%
Two experiments, two cohorts, and the ordering of the three measured platforms in this article stays the same. That consistency matters more than either result on its own.
Voice quality collapses three separate measurements into one word, and vendors quote whichever one flatters them. Separate them and the comparison becomes straightforward.
ElevenLabs documents roughly 75ms, and that figure covers text-to-speech synthesis alone, with full end-to-end conversational latency running substantially higher (Cekura platform review). Measured as a complete turn on the same harness as everyone else, the same platform records 1.27s mean response (Cekura Bench). Both numbers describe ElevenLabs correctly. Only one of them describes a phone call.
Coval benchmarks 24 text-to-speech models on time to first audio and word error rate, and 31 speech-to-text models on time to final segment and word error rate, against one pinned dataset, refreshed roughly every 30 minutes, with open Apache-2.0 methodology. The gap between the fastest and slowest text-to-speech model in that set exceeds 1,000ms, and no provider document reveals it. Check the live TTS leaderboard and STT leaderboard before you blame a platform for a component you picked yourself.
Coval states plainly that perceived naturalness resists automation and recommends a structured listening test on your own audio. Cekura scores it as voice tone and clarity, where ElevenLabs leads the current cohort at 4.47 out of 5, with Retell and LiveKit at 4.36 (Cekura Bench). ElevenLabs winning that column is the expected result, and it is the reason Retell offers ElevenLabs voices inside its own platform.
If a vendor will not tell you where their clock starts and stops, treat the number as decoration.
Human conversation turns over with a mean gap of about 208ms, which sets the perceptual bar rather than a target any platform currently reaches (Cekura latency analysis). Measured per turn on platform defaults, the six platforms in Cekura archived fixed-stack study ranged from 1,730ms to 3,160ms at the median (Cekura latency analysis). Any published figure in the low hundreds of milliseconds measures a component, not a conversation.
The tail matters more than the median, and it reorders the field. In that study ElevenLabs led the median at 1,730ms but reached 3,194ms at the 95th percentile, while Vapi sat third at the median with 2,340ms and first at the 95th with 2,950ms. Vapi held the tightest spread in the group, a 1.26x multiplier from median to 95th, against 1.85x for ElevenLabs and 1.93x for Retell (Cekura latency analysis). Callers experience the tail, so a shortlist built on medians picks a different vendor from one built on the 95th percentile.
One caveat deserves repeating rather than burying. These are platform defaults, and the numbers move once a team tunes an agent for a specific use case (Cekura latency analysis).
Interruption handling separates a demo from a production agent, and almost no comparison measures it. Cekura scores barge-in and turn changes out of 5. In the current cohort Retell scores 5.00, five configurations land between 4.96 and 4.98, and Vapi scores 4.73 (Cekura Bench).
Those scores cluster tightly, so treat the column as a floor check rather than a ranking. On your own calls, measure the two timings the research literature separates: how long the agent keeps talking after a caller cuts in, and how long it then takes to answer what the caller said. Full-Duplex-Bench v1.5 formalises both and publishes its code. Fast yielding and floor holding trade against each other, so a system that drops the floor instantly also drops it when a dog barks.
These rows hold their value for longer than a quarter. Everything faster moving than that, including review scores and monthly call volumes, belongs on the vendor pages linked in each cell.
| Dimension | Retell AI | Bland AI | Vapi | ElevenLabs |
|---|---|---|---|---|
| Base price | Published rate card, varies by model and voice. | Plan tiers plus per-minute rate | $0.05/min orchestration | Free to $99+/mo plan tiers |
| What the base covers | Voice engine, with LLM and telephony passed through at cost | LLM, voice and telephony bundled in the per-minute rate | Orchestration only, every component billed separately | Agent minutes, with the reasoning model and telephony billed separately |
| Telephony | Twilio, Vonage, Telnyx, SIP, web SDK | Own numbers or bring your own Twilio | Twilio, SIP, WebRTC | Twilio, Vonage, SIP |
| HIPAA | Standard plans, self-service BAA | Enterprise tier, signed BAA | Paid add-on | Available |
| No-code builder | Yes | Pathways graph builder | Flow Studio | Workflow builder |
| Bring your own LLM | GPT, Claude, Gemini, custom | Plan-gated | OpenAI, Anthropic, Google, custom | Third-party or custom model |
| Self-hosted option | Yes | Yes | No | No |
| Free to start | $10 in credits | Free plan with an inbound number | $10 in credits | Free plan with 15 agent minutes |
Prices and plan structures change often. Each cell links the vendor page it came from, so confirm before you budget.
The tables above show where the four platforms diverge. This section explains what those differences mean once an agent takes live calls.
Measured: 75.61% pass³, 93.88% task completion, 98.37% of calls free of provider or connection issues, 5.00 out of 5 on interruption handling, 4.36 on voice tone and clarity, and 2.21s mean response (Cekura Bench). On the regulated Medicare workflow Retell passed 22 of 23 evaluators on all three runs, first in that field (Cekura Experiment 02).
Retell runs a no-code builder and a developer SDK in the same product, so an operator and an engineer work side by side. A streaming knowledge base keeps answers current, call transfer hands a caller to a human with the context already gathered, and post call analysis scores every call afterwards. HIPAA ships on standard plans.
Retell publishes its full rate card with no platform fee and no feature gating, so every plan opens the entire platform. Where you land depends on the model and voice you pick rather than on a sales conversation. Telephony is $0.015 a minute, SIP trunking carries no extra charge, and every account includes 20 concurrent calls.
Retell is not the fastest platform in the cohort, and three configurations answer sooner. Cekura also recorded a run where our transcript captured a phone number correctly and a different number reached the tool (Cekura Bench). Advanced multi-step flows still reward prompt tuning, which G2 reviewers note alongside high marks.
Measured: nothing. Bland does not appear in Cekura published cohorts, and we found no repeated-run independent benchmark covering it while writing this. Hold every Bland figure below to a lower standard than the rows above.
Bland targets high-volume outbound, runs self-hosted models on dedicated GPUs, and offers the cleanest deterministic flow builder in this group through Pathways, which matters when a script must run the same way every time. Its billing documentation lists plan tiers with a per-minute rate that falls as the plan rises (Bland billing docs).
Hands-on review from Cekura puts Bland latency in the 700 to 900ms range and notes Discord-based support with no dedicated account management. That is a reviewer observation rather than a harness result, and the difference matters. If you run or know of a repeated-run benchmark that includes Bland, send it to us and we will add it here.
Measured: 59.76% pass³, the highest task completion in the cohort at 97.56% across the 205 calls with outcome evidence, 82.93% of calls free of provider or connection issues, 4.73 on interruption handling, and 3.08s mean response. On the Medicare workflow Vapi passed 15 of 23 evaluators.
Vapi is the orchestration layer for teams that want to own every component, connecting more than 14 providers for speech-to-text, language models, voice and telephony behind one API. Squads let developers chain specialised agents inside a single call. The advertised $0.05 per minute covers orchestration alone, so the real rate lands higher once the stack is attached, and HIPAA sits behind a paid add-on.
Vapi earns one genuine measurement win. In the archived fixed-stack study it held the tightest latency spread of the six platforms, a 1.26x multiplier from median to 95th percentile, so its slow turns stay closest to its typical turns. Vapi rewards teams with engineers and punishes teams without them.
Measured: 69.51% pass³, 91.46% task completion, 100% of calls free of provider or connection issues, the highest voice tone and clarity score at 4.47, and the fastest mean response at 1.27s (Cekura Bench). On the Medicare workflow ElevenLabs finished second at 21 of 23 evaluators (Cekura Experiment 02).
ElevenLabs makes the most natural voices in the category with broad multilingual coverage, which is why other platforms, Retell included, integrate its voices. Teams stand up a basic agent quickly, and telephony still runs through Twilio, Vonage or SIP, with the reasoning model and telephony billed on top of the plan.
Two cautions. The widely quoted 75ms figure measures synthesis, not a conversational turn. And Cekura recorded a run where the agent narrated a tool call and then continued with an invented result, which is the failure mode that matters most when a voice sounds convincing.
Bland and Retell both target outbound at scale, so the decision comes down to control against evidence. Bland Pathways gives engineers deterministic, node-by-node command over a script, which suits collections and compliance-heavy campaigns where every branch must be predictable.
Retell covers the same outbound ground with a published independent record behind it. Built-in batch call handling drives reminders, surveys and lead follow-up at volume, branded call ID lifts pickup on cold outbound, and the no-code builder lets a non-engineer adjust a lead qualification flow that Bland would route through code.
Neither of us can settle the speed question with data, because no independent harness has tested Bland. Choose Bland when a team has engineers who specifically want Pathways-style determinism, and choose Retell when you want a measured reliability record and an operator who can change the flow.
Vapi and Retell appeal to overlapping developer audiences and assign the integration burden differently. Vapi hands you maximum control and asks you to assemble and maintain the stack, which is the right trade for a team building voice as a core product.
The measurements split along the same line. Retell leads repeatable reliability at 75.61% pass³ against 59.76%, and keeps 98.37% of calls free of connection issues against 82.93%. Vapi leads task completion among the calls it connects and holds the tightest latency spread in the archived study (Cekura latency analysis).
Retell also ships the connectors most teams need, including a pre-built HubSpot connector and n8n for event-driven flows after each call. Vapi wins when an engineering team wants to swap providers per call stage. Retell wins when the goal is a working production agent this week.
ElevenLabs wins on voice quality, and the independent number says so: 4.47 out of 5 on voice tone and clarity against 4.36 for Retell, plus the fastest mean response in the cohort at 1.27s against 2.21s. For a branded consumer product or a voice companion, ElevenLabs sounds better, and Retell offers ElevenLabs voices for exactly that reason.
Retell leads where a phone agent has to hold up. It finished ahead on repeatable reliability, 75.61% against 69.51% and on the regulated Medicare workflow, 22 of 23 evaluators against 21 . Telephony, monitoring, simulation testing and warm transfer ship in the product rather than bolted on.
Compliance widens the gap for regulated work in healthcare, where Pine Park Health reported a 38% increase in scheduling NPS after deploying Retell for patient scheduling. HIPAA ships on Retell standard plans with a self-service BAA, and you can read the federal rules behind that requirement on the U.S. Department of Health and Human Services HIPAA portal. If voice fidelity is the single most important variable, choose ElevenLabs. If a deployable, compliant phone agent is the goal, the measurements favour Retell.
Headline rates mislead in this category, because every platform bills some components separately. The table below models realistic all-in cost per minute rather than the advertised floor.
| Cost component | Retell AI | Bland AI | Vapi | ElevenLabs |
|---|---|---|---|---|
| Platform or base fee | None | $0–$499/mo plan | $0.05/min orchestration | $0–$99+/mo plan |
| LLM | Pass-through $0.003–$0.08/min | Bundled in base | Separate, $0.06–$0.10/min | Separate, passed through |
| Voice (TTS) | $0.015–$0.040/min | Bundled, cloning add-on | Separate, $0.04–$0.08/min | Included in agent minutes |
| Telephony | $0.015/min or own SIP | Bundled or BYO Twilio | Separate, ~$0.015/min | Separate at cost |
| Compliance | HIPAA included | HIPAA on Enterprise | HIPAA add-on ~$1,000/mo | HIPAA now available |
| Realistic all-in per minute | Published rate card, varies by model and voice. | $0.11–$0.18 | $0.12–$0.33 | $0.08 overage plus LLM and telephony |
Bland is often the cheapest predictable rate at high outbound volume because its base bundles the stack. Retell stays lowest among the unbundled platforms because it charges no platform fee, and it avoids the five-invoice complexity and paid HIPAA add-on that come with Vapi.
Read cost next to the reliability data before you decide. A cheaper minute on a platform where 41 of 246 test calls never connected is not a cheaper conversation, and a retried call bills twice.
ElevenLabs sounds better and answers faster. It leads voice tone and clarity at 4.47 against our 4.36, posts the fastest mean response at 1.27s against our 2.21s, and was the only configuration with 100% of calls free of provider or connection issues (Cekura Bench).
Vapi finishes what it starts. It leads task completion at 97.56% among the calls that produced outcome evidence.
Retell is not the fastest. Three configurations in the cohort answer sooner than we do.
Our agent made mistakes in the run. Cekura provider notes record a call where the transcript captured a phone number correctly and a different number went to the tool.,
If voice quality is the product, pick ElevenLabs. It leads the independent naturalness score and the response-time column, and those are the two things a branded voice experience lives on.
If repeatable reliability and interruption handling decide it, pick Retell. It leads pass³ in the current cohort, scores 5.00 out of 5 on turn changes, and led the regulated Medicare workflow on both the workflow and the strict measure.
If your team wants component-level control and will maintain the stack, pick Vapi. Its tight latency spread and per-stage provider swaps reward engineers who want to own the pipeline.
If deterministic outbound flows at volume matter most, look at Bland, and go in knowing no independent benchmark covers it yet.
The honest way to settle it: build the same agent on two of these platforms with free credits, run twenty real calls, measure the median and the 95th percentile yourself, and keep the one your team still wants a week later.
How do I compare voice quality between AI platforms?
Measure three things separately: synthesis speed, word error rate on recorded phone audio, and naturalness through a blind listening panel. Vendor latency figures often describe the synthesis layer alone, so check what the clock measures before you compare two numbers. Coval publishes live TTS and STT leaderboards with open methodology as a neutral starting point.
Which platform has the best voice quality in independent testing?
ElevenLabs leads the Cekura voice tone and clarity score at 4.47 out of 5, with Retell and LiveKit at 4.36.
Which platform is most reliable?
Retell leads repeatable reliability at 75.61% pass³ across 82 scenarios run three times each, and led the Medicare workflow test at 95.7%
What does pass³ mean?
A scenario counts only when all three runs pass, so the score measures consistency rather than a single good call.
Which platform is fastest?
ElevenLabs, at 1.27s mean response in the current cohort, ahead of GPT Realtime at 1.58s, Pipecat at 1.97s and Retell at 2.21s
Why do published latency numbers disagree so much?
They measure different boundaries. Some cover speech synthesis, some cover a full conversational turn including the model and any tool calls, and most do not say which. Human conversation turns over at about 208ms, so any platform figure in the low hundreds is measuring a component (Cekura latency analysis).
Is a lower median latency always better?
No. Vapi placed third on the median and first at the 95th percentile in the archived study, because its spread was tightest. Callers experience the tail.
How do I test interruption handling?
Cut in mid-sentence and time two things: how long the agent keeps talking, and how long it then takes to answer you. Full-Duplex-Bench v1.5 defines both metrics and publishes the code.
Is Bland covered by these benchmarks?
No. Bland does not appear in the published cohorts, so its figures on this page come from vendor documentation and hands-on review rather than repeated-run testing.
Did Retell run these tests?
No. Cekura ran them and publishes the per-run reports. Retell wrote this article, and every number links back to the source so you can check it.
See how much your business could save by switching to AI-powered voice agents.
Total Human Agent Cost
AI Agent Cost
Estimated Savings
A Demo Phone Number From Retell Clinic Office

Start building smarter conversations today.




