
VapiCon is back November 11-12! Tickets now available•Register Now

VapiCon is back November 11-12! Tickets now available•Register Now

Last month we opened a limited beta for GPT-Live, OpenAI's full-duplex speech-to-speech model, and started a waitlist. Today the beta is open to every Vapi customer. If you have a Vapi account, you can build with GPT-Live now.
We spent the first month working hands-on with a small group of teams to get the model into production. That work shaped how GPT-Live runs on the platform. Now we're opening it up to learn from more of you.
Voice agents have historically run on a modular pipeline: a speech-to-text model transcribes the call, an LLM reasons over the text, and a voice model speaks the response. GPT-Live is a speech-to-speech model. It listens to the caller's audio and generates its own speech, all in one model.
On Vapi, GPT-Live runs as two roles: a speaker and a reasoner. The speaker is the model itself. It handles the whole conversation and keeps listening while it talks. When the call needs a lookup or an action, the speaker delegates to the reasoner, a separate model that follows your procedures, runs your tools, and returns the result. You control that handoff in the prompt. The speaker never stops talking, so the conversation stays natural the entire time.
For use cases where sounding human decides the outcome, this is a step change.
GPT-Live runs on the same stack as your existing agents. Phone numbers, tools, tests, and observability all carry over.
What changes is how you set up the agent. A standard Vapi assistant has one system prompt that does everything: tone, procedures, tool rules, edge cases. GPT-Live splits that into two prompts, one for each model.

The speaker prompt shapes what the caller hears. Tone, pacing, how to greet, what to talk about while a tool runs. It also holds the delegation triggers, the specific conditions under which the speaker hands work to the reasoner. "Delegate when you know the service, location, and date." "Don't delegate if you can answer from what was already said."
The reasoner prompt holds the procedures. Tool rules, confirmation steps, what to check before taking an action, and how to format a result so the speaker can say it in a sentence.

In practice, migrating an assistant means pulling your existing prompt apart along that seam: conversation guidance goes to the speaker, procedures go to the reasoner. The transcriber config goes away since the model works from audio, and the voice switches to one of OpenAI's 22. Composer can do the first pass of that split from an existing assistant. If you're running a squad, each specialist assistant becomes a reasoner skill and the front-desk routing becomes delegation triggers, so the caller stays with one voice the whole time.
Everything downstream is unchanged. Function, API request, and MCP tools work as-is. Transcripts, recordings, scorecards, and Voice Simulations run against GPT-Live assistants the same way, so you can test the new version against the same calls before moving a phone number over.
Vapi was built to be model-agnostic from day one. Assistants, tools, telephony, testing, and observability all sit independent of the model layer, so when a new model generation ships, it lands on the platform as a config change. You're never locked into a single provider.
That's what made the beta possible. Teams already running GPT-Live didn't rebuild anything to try it, and nothing is locked in if a different architecture serves the use case better.
New models are arriving faster than most platforms can adopt them. Our job is to make each generation production-ready as soon as it's viable, so you're choosing the best model for the job instead of inheriting whatever your platform is locked into.
Speech-to-speech is an addition to Vapi, not a replacement. You can run a modular pipeline or a speech-to-speech model, and choose per agent.
Modular pipelines give you maximum control: choose your transcription, intelligence, and voice models independently, tune each stage, and swap components as better ones ship. That control is why many production use cases are best served by modular configurations today, and why we'll keep investing in them. On Vapi, you run both side by side, and we're continuing to build infrastructure that tightens how the stages of a modular pipeline work together.
GPT-Live is live in beta for every Vapi customer today. Open the dashboard, create an assistant, pick GPT-Live as the model, and place a call. For a quickstart, specific examples and more details, read the docs and start building →
