
VapiCon is back November 11-12! Tickets now available•Register Now

Today we're opening a limited beta for GPT-Live, OpenAI's next-generation full-duplex speech-to-speech model, on Vapi. We're starting hands-on with a small set of customers, working directly with their teams to bring the model into production. Alongside the beta, we're opening a public waitlist for early access.
Joining the waitlist doesn't guarantee early access, but it puts your team first in line as we expand, and you'll get early insight into the model's availability on Vapi. Through the beta, we're growing deliberately, prioritizing use cases where we can best support teams, with broader availability planned in the coming months. If you want in, join the waitlist and tell us what you'd build with the model.
Voice agents have historically run on a modular pipeline: a speech-to-text model transcribes the call, an LLM reasons over the text, and a voice model speaks the response. GPT-Live is a speech-to-speech model: it does all three in one.
Traditional Modular Pipeline:
GPT-Live Speech-to-Speech:
When paired with the most comprehensive voice agent platform, it unlocks new capabilities and can help orchestrate new possibilities for your voice agents:
It converses, not just responds. The model backchannels, pauses naturally, and handles interruptions cleanly, because it's listening the whole time it's talking.
No dead air. With no handoffs between models, turn latency drops to something callers read as human.
It hears more than words. Tone and emphasis reach the model directly instead of being flattened into a transcript before it reasons.
Accuracy where calls are won or lost. Stronger capture of numbers, names, and addresses, the details that make or break verification and support calls.
Deep work without breaking flow. Complex tool calls and delegated reasoning happen while the conversation keeps moving.
For the use cases where sounding human drives outcomes (patient calls, identity verification, high-stakes support), this is the capability customers have been asking for by name, and we're excited to support it.
Power of GPT-Live compounded on Vapi:
Vapi was built to be model-agnostic from day one. Assistants, tools, telephony, testing, and observability all sit independent of the model layer, so when a new model generation ships, it lands on the platform as a config change, not a full migration, and you're never locked into a single model provider.
That's what makes this beta possible. Teams in the beta run GPT-Live on the same stack as their existing modular agents: same tooling, same phone numbers, same tests. Nothing gets rebuilt to try a frontier model, and nothing is locked in if a different architecture serves the use case better.
New models are arriving faster than most platforms can adopt them. Our job is to make each generation production-ready as soon as it's viable, so you're choosing the best model for the job instead of inheriting whatever your platform is locked into. GPT-Live is the latest example of frontier innovation, and it won't be the last.
Speech-to-speech is an addition to Vapi, not a replacement.
Modular pipelines give you maximum control: choose your transcription, intelligence, and voice models independently, tune each stage, and swap components as better ones ship. That control is why many production use cases are best served by modular configurations today, and why we'll keep investing in them. The right architecture depends on the use case, not on what your platform happens to sell. On Vapi, you'll run both side by side and choose per agent.
The private beta is underway, and the public waitlist is open. Signing up doesn't guarantee access, but it puts you in line, and the more we know about your use case, the better we can prioritize.
