Stand up a WebSocket server that speaks Retell's custom LLM protocol.
Handle the events Retell sends: transcript updates, interruptions and call metadata.
Stream your model's responses back as they generate, so speech starts without waiting for the full answer.
In Retell, create an agent and select Custom LLM as the response engine.
Enter your WebSocket URL and any authentication headers.
Test in the playground, confirm latency is acceptable, then assign a phone number.
Call connects β Retell streams the transcript to your server β your model decides and responds β Retell speaks it and handles the interruption
β οΈ This is not a partner integration. It has no logo, no vendor site and no install URL, so the card and hero will look wrong in the current template. Recommend either moving it out of the integrations directory into Developers/Docs, or giving it a distinct card treatment.












.avif)
.avif)
.avif)

.avif)
.avif)












.avif)














Start building smarter conversations today.

