Vapi raises $50M Series B
Read More →
Vapi raises $50M Series B to power the next generation of enterprise voice AI
Vapi raises $50M Series B
Read More →
Two models added to the Humanness Index™ in the last ten days went straight to the top of the leaderboard. State of the art is now measured in weeks.
A month ago we published the first Humanness Index™ results. xAI's Grok TTS led the models at 95, five points under the real human baseline, and the story was how close synthetic voices had gotten to a real person. Those standings held for about three weeks.
Then they stopped holding. On July 24, we added Speechify's Simba 3.2 to the arena, and it went straight to the top of the board at 99, only one point under the human. Even more recently, we added Fish Audio's S2.1-Pro, and it debuted at 94 with a 141 ms time to first audio, the fastest among the leaders. The success of both models serves as a harbinger of the continued success and development of voice models, and the importance of staying flexible by using a provider that can support them all.
As a reminder, the Humanness Index™ is a blind, crowdsourced benchmark for how human a voice model sounds. We clone the same voice onto every model and have each one read the same script, then listeners hear two clips side by side with no labels and pick the one that sounds more human. A real human recording is included in the comparison set and anchors the scale at 100, so a model scoring 99 means listeners picked it almost as often as the actual person. The newly added Speechify Simba 3.2 currently leads the Index at 99 humanness. In blind battles, listeners picked it over other models at nearly the same rate they picked the actual human recording of the same quote. Its latency benches at 428 ms to first audio and sits among the cheapest paid models on the board.
Fish Audio S2.1-Pro landed at 94 in its first week of voting. The score alone would be news, and the latency makes it more interesting: 141 ms, on a board where most of the leaders sit between 300 and 800 ms. A voice that sounds human and starts speaking in under 150 ms makes it an even more realistic voice model in live deployments.
Both models are new to the arena, so their vote counts are still low (139 and 126 against 500+ for older entries), and their rank ranges are wide. More votes will tighten those ranges, but regardless, the initial reception of these models shows the continued innovation that we are seeing with more human models.
Grok TTS held the top spot for three weeks. Simba 3.2 has held it for ten days, and S2.1-Pro closed to within two points after seven. Since launch, the leading humanness score has moved from 95 to 99 against a fixed human baseline of 100. The room at the top is nearly gone, and the names occupying it continue to change.
This is what a fast-moving field looks like from inside a benchmark. The models people were deploying in June are mid-table in August, and nothing about them got worse. The field moved, and with the current pace of AI innovation, state of the art is accelerating faster than ever.
If you picked a voice vendor in June that locked you into their voice models, this month's leaderboard is a representation of the fact that you might already be falling behind. If your platform treats the TTS model as a swappable parameter, you can innovate faster and change models as the best ones become available.
That's the practical argument for a model-agnostic voice platform. No single vendor is the problem; the problem is that no ranking survives the next release. We built Vapi so the voice layer, along with your transcriber and intelligence models, is a choice you can revisit at any time. The Index itself is independent of the Vapi product, but the lesson still lands on the product side: the best model this quarter wasn't the best model last quarter, and your switching cost decides whether you benefit from that or watch it from a distance. And with Model Intelligence, Humanness Index™ scores are available directly in the platform, so you can compare how human each voice sounds without leaving Vapi.
Nearly 12,000 votes are in across 21 models from 11 providers. The two newest entries have barely 120 votes each, which is why their rank ranges are still wide. A few minutes of blind listening from you narrows them. Listen to a few battles at humannessindex.vapi.ai and pick the voice that sounds more human. Then check back, because the last two weeks suggest the board won't look the same for long.
