
VapiCon is back November 11-12! Tickets now available•Register Now

The biggest problem in CX is simple to state. By ACSI's estimate, companies spend well over $100 billion a year on customer experience, with no detectable return. The national score sits exactly where it was in 2013.
Complaints hit record levels in the first quarter of this year, up 16%. Corporate profits rose 15% over the same period.
That gap is the whole story. Investments funded IVR, offshoring, self-service, chatbots, and deflection platforms; each one was built to make an interaction cheaper, and each one delivered on what it promised. Cost per contact went down. Average handle time went down. Containment went up. The tools worked for what they were optimized for: minimizing cost for the business, not maximizing value for the customer.
That's why we're launching Vapi for CX: the complete toolkit your CX, operations, product, and engineering teams need to ship fast and keep improving voice agents for customer experience.
The modern contact center was built around one question: how do we get through these callers as fast as possible? Automatic call distribution arrived in the 1970s to answer it, and everything since has been refinements and evolution of the same idea.
IVR came next. Offshoring followed, then web self-service, then chatbots. IVA was supposed to be a big leap forward, but was more of the same.
Each wave was measured the same way. Cost per contact, average handle time, containment rate. Each one worked, in the sense that the number its targeted outcome went down. And satisfaction stayed flat through all of it, because nobody was optimizing for whether the customer's problem got solved.
Large language models changed what's possible. An agent can hold a real conversation, understand an unstructured request, and take an action in another system. Where even the best IVAs could only route you quickly. A voice agent can actually fix your problem on the spot!
Today though, a lot of deployments don't. They point the new technology at the old metric and measure how many calls never reached a human. The result is a more pleasant version of deflection, which is progress of a kind, but nowhere near what the technology can do.
Part of that is habit. A bigger part is that most teams can't reach further even if they want to.
There are a few ways to get a voice agent, each with their own tradeoff.
Rent a managed agent.
An off-the-shelf agent app hands you something that works on day one. Sensible defaults, a dashboard your team can use, a vendor who owns the hard parts. Then you meet the ceiling. The model is chosen for you, the orchestration is closed, and the behaviors you most want to change are the ones the vendor decides. Your request hits the tail end of a roadmap you don't control.
A custom managed service moves the ceiling without removing it. Someone builds you something that fits, which is genuinely better than a template, but it's only editable inside their proprietary sandbox. And the people who can change it don't work for you and have other customers, so every adjustment is a ticket and a wait. You've ended up with a bespoke black box instead of a standard one.
Build the agent yourself.
Now your engineers have everything. The model, the logic, the data, the orchestration. They also have telephony, turn-taking, interruption handling, fallback plans, concurrency, scalability, and compliance to own forever.
Building on a platform takes most of that infrastructure off their plate. The catch is what platforms do next. Most of them abstract away the deep configuration in the name of making it easy to build, and the abstraction is usually good. It's also a set of decisions someone else made for you. Every setting the platform smoothed over is a setting you can't reach when the call goes wrong in a way the vendor didn't anticipate. That caps what you can build, and it's ownership you gave up without being asked.
What we kept hearing from contact center leaders is that this trade is the actual problem. You can't run a contact center where the people closest to the customer need an engineer to change a sentence. And you can't put a black box in front of 100% of your inbound volume and tell your board you're accountable for the experience.
That's the trade we set out to remove. Ease should be a layer, not a replacement. For example, Vapi Composer is a natural language way to build voice agents, and your CX team can build a working agent through it without touching a single model parameter. But what comes out the other side is the same config object your engineers reach through the API, the CLI, and an MCP server, with every setting underneath still there and still editable. Hundreds on a single assistant.
And the part neither team wants to own, the telephony, turn-taking, interruption handling, fallback plans, concurrency, and compliance, is ours to manage and yours to tune.
So the answer to how much you should own is: the parts that are your business. The prompts, the logic, the tools, the integrations, the data, the model choice. Your agent, your stack, your outcomes.
On one axis, how much of the system you control. On the other, how much value a single interaction can produce. The reason? Control dictates how much of the agent you can reach, and how quickly you can tune.
Resolving a problem end to end requires the agent to speak naturally, understand in real time without upsetting the customer, then read entitlement logic, write to your system, and confirm the change. That capability doesn't just come from better models. It comes from being able to define the tools, the schemas, and the behavior yourself.
Vapi CX is full control over the agent lifecycle. Everything you need to ship fast and continually improve voice agents for customer experience.
Two ways in.
Your CX team describes the agent in Composer. It writes the prompt, provisions the number, and wires up the tools. A model preset bundles a transcriber, model, and voice into a single choice, so nobody has to assemble a stack on day one. You have something answering calls the same afternoon.
Your engineers take a different door. The Vapi MCP server and published OpenAPI spec mean a coding agent writes against real API definitions instead of guessing, so the integration into your order system gets built in hours rather than sprints.
Both doors produce one versionable config object. What your CX team edits in the dashboard is what your engineers hit through the API. That's the difference between a fast start and a low-code ceiling.
Underneath it, hundreds of configurable settings on a single assistant, plus:
Amazon Ring went from zero to production in two weeks this way, and now runs 100% of its inbound volume on Vapi. They picked Vapi after evaluating more than 40 vendors specifically highlighting how Vapi enables their CX and Ops teams to tune the agent without depending on engineering. This allows CX to own the experience without giving up any of the deep configurability that makes the Ring agent unique and effective.
Wire up your SIP trunk, and Vapi sits in front of what you already run. Your ACD, routing, reporting, and workforce management stay exactly where they are. Calls the agent can't resolve warm transfer to a human with full context, so the customer doesn't repeat themselves and the rep doesn't start cold. It works for outbound too: reminders, delivery notifications, callback deflection, and renewals.
From there, the question is whether you can deploy safely. Four things make that true:
Scorecard grades 100% of calls against criteria you define. Did the agent verify identity before touching the account? Did it resolve without transferring? Did it offer escalation when policy required it?
Three things make that useful rather than just visible:
Observability only matters if you can act on it. A failing row in the scorecard is a specification.
Turn it into a simulation. AI testers with distinct personalities run your scenarios against a draft, including the caller who interrupts mid-sentence and the one who demands a refund on turn one. Fix the agent, run it again, ship when it passes.
Most of the levers have nothing to do with the model:
When the model is the lever, Model Intelligence publishes latency, cost, and quality on every transcriber, model, and voice, refreshed weekly from medians across production calls. Something cheaper or better ships; you compare it and seamlessly make the switch.
At Kavak, business teams build new versions in about five minutes. NPS is up 20 points, conversion 30%, since launch.
Run the improvement loop for a few months, and something changes that's easy to miss while you're in it. The agent stops being a thing that answers calls and starts being a thing that handles real work.
Deflect the FAQ. The agent answers and routes. Cost per contact. This is table stakes for voice agents today.
Understand, then hand off. The agent knows exactly what the customer needs and transfers with full context. Better than a cold transfer, and a human still does the work.
Resolve it end to end. The agent diagnoses, acts in your systems, and confirms, all in real time. First rung where cost and satisfaction move in the same direction.
Save the cancellation. The agent handles the objection and keeps the account. You're protecting revenue now, not reducing expense.
Sell and expand. The agent books, upsells, and follows through until it's done.
Getting started doesn't take a procurement cycle. Point Composer at your help center and you'll have an agent answering calls this afternoon, or pull the MCP server into your editor and build it from your own call data. Keep the contact center you already run when you're ready to take real traffic. Monitor your performance; what you discover becomes tomorrow's test and next week's deployment.
The best part? You can build it all before you talk to anyone here.
