Voice AI Benchmark: Retell Ranks #1 for Medicare Workflow Accuracy

Voice AI Benchmark: Retell Ranks #1 for Medicare Workflow Accuracy
BACK TO BLOGS
Add Retell AI as a preferred source on Google
ON THIS PAGE
Back to top

A new voice AI benchmark took one Medicare insurance agent, put it on six different platforms without changing a line, and measured which one actually followed the workflow. Retell finished first.

The test came from Cekura, an independent company that builds evaluation harnesses for voice agents. Its Medicare experiment ran 414 calls in total and scored each platform on how reliably it handled a regulated insurance conversation.

Retell led the field at 95.7% workflow accuracy, clearing 22 of the 23 checks. Every other platform landed between 65.2% and 95.7%, even though they all ran the identical agent.

TL;DR

  • Cekura deployed one byte-identical Medicare agent across six voice platforms and ran 414 calls.

  • Retell ranked #1 at 95.7% workflow accuracy, passing 22 of 23 evaluators.

  • Scores across the field ran from 65.2% to 95.7%, so the platform, not the agent, drove the difference.

  • Retell also held first place under strict end-to-end grading, which counts runtime failures, again at 95.7%.

  • On a regulated Medicare call, that gap is the difference between a compliant handoff and a dropped or non-compliant one.

What Cekura's voice AI benchmark tested

Cekura built one Medicare TPMO agent and deployed a byte-identical copy on each of the six platforms. Same prompt, same tools, same voice. The only variable was the platform running underneath.

A TPMO is a third-party marketing organization, the kind of licensed agency that sells Medicare Advantage and Part D plans. These calls are tightly regulated, so the agent has to get a lot right, and in a fixed order.

The benchmark used 23 evaluators, each targeting a point where insurance calls tend to break. They fell into four groups: opening the call and getting permission, working out what the caller actually needs, qualifying and capturing the lead, and staying inside consumer-protection limits.

Every evaluator ran three times on each platform, and a platform earned credit only if all three runs passed, a measure Cekura calls pass3. Twenty-three evaluators, three runs each, across six platforms works out to the full 414 calls. You can read the methodology and per-platform runs on Cekura's benchmark page.

The results: Retell ranked #1 at 95.7%

Here is how the six platforms scored, ordered by workflow accuracy:

PlatformWorkflow pass3Strict end-to-end pass3
Retell95.7% (22/23)95.7% (22/23)
ElevenLabs91.3% (21/23)91.3% (21/23)
Pipecat82.6% (19/23)73.9% (17/23)
LiveKit73.9% (17/23)73.9% (17/23)
Synthflow69.6% (16/23)8.7% (2/23)
Vapi65.2% (15/23)65.2% (15/23)

Retell was the only platform to clear 95% on workflow accuracy, and it kept that score under the stricter measure explained below.

Two ways to measure reliability

Cekura reported two scores because a voice agent can fail in two different ways.

Workflow pass3 asks whether the agent chose the right action and called its tools correctly. It uses two signals Cekura calls Expected Outcome and Mock Tool Accuracy.

Strict end-to-end pass3 adds a second question: did the caller actually get through the task? It counts infrastructure failures, like the agent stalling mid-call, that stop a caller from finishing even when the intended action was right.

Retell scored 95.7% on both, so it did not lose ground to runtime problems. Some platforms did. Synthflow passed 16 of 23 evaluators on workflow accuracy but only 2 of 23 under strict grading, because during tool use callers often sat through long silences with no sign the agent was still working.

Why workflow accuracy matters for Medicare calls

On a Medicare sales call, the agent is not just answering questions. It has to deliver the required disclosure, earn consent before discussing plans, avoid giving personalized advice, and route the caller to a licensed human at the right time.

Those steps are set by federal TPMO rules, and skipping one is a compliance problem, not a rough edge.

That is what the 23 evaluators probed. A caller interrupts the disclosure and demands plan details. A family member calls on behalf of a beneficiary. A caller asks the agent to promise that a plan covers their medication. The right move in each case is specific, and a platform that stalls or improvises fails the caller and the compliance record at once.

The benchmark shows that running the same agent is not enough. The platform underneath decides whether that agent holds up on the highest-risk calls.

Where Retell fits

Retell is a platform for building, testing, deploying, and monitoring AI voice agents that make and take phone calls. The benchmark rewarded two things Retell is built around: following a multi-step workflow and calling tools correctly under pressure.

  • Warm transfer to a licensed human: when a caller needs a person, the agent hands off with the context it already gathered, using call transfer.

  • Answers from a knowledge base, not off-script advice: the agent pulls factual answers from a connected knowledge base and stays inside what it is allowed to say.

  • Every call reviewed after the fact: teams use post-call analysis and AI quality assurance to see where a flow held up and where it slipped.

  • Built for regulated industries: Retell runs voice agents for teams in healthcare and insurance, where the order of operations on a call is not optional.

This is one independent test on one workflow, and results vary by how an agent is built and maintained. Even so, on a hard, regulated task, Retell came out ahead of the field.

Frequently asked questions

What is a voice AI benchmark?

A voice AI benchmark is a controlled test that runs the same task across different voice-agent platforms and scores how well each one handles it. Cekura's version uses repeated calls and fixed evaluators so the comparison stays fair.

Who ran this benchmark?

Cekura, an independent company that builds evaluation tools for voice agents. Retell did not run the test or grade itself.

What does pass3 mean?

A platform had to pass an evaluator on three separate runs to get credit for it. One lucky run does not count, which makes the score a test of consistency rather than a single good call.

Did every platform run the same agent?

Yes. Cekura deployed a byte-identical Medicare agent to all six platforms, with the same prompt, tools, and voice, so any difference in score comes from the platform, not the agent.

Where can I see the full results?

Cekura publishes the per-platform runs and the full methodology on its benchmark page, including a breakdown of every evaluator in the Medicare experiment.

Build a voice agent that holds up on the hard calls.

Retell lets you build, test, and launch an AI voice agent, then watch how it performs on every call. Try Retell free or talk to sales.

ROI Calculator
Estimate Your ROI from Automating Calls

See how much your business could save by switching to AI-powered voice agents.

All done! 
Your submission has been sent to your email
Oops! Something went wrong while submitting the form.
   1
   8
20
Oops! Something went wrong while submitting the form.

ROI Result

2,000

Total Human Agent Cost

$5,000
/month

AI Agent Cost

$3,000
/month

Estimated Savings

$2,000
/month
Live Demo
Try Our Live Demo

A Demo Phone Number From Retell Clinic Office

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

Read Other Blogs

Revolutionize your call operation with Retell