The Great Voice-AI Debate: Cascaded vs Speech-to-Speech

The Great Voice-AI Debate: Cascaded vs Speech-to-Speech
Jay has got the mic 👀 🐦
upload in progress, 0

Welcome back to The Bluejay Times !

For the past year, the narrative in voice AI has been clear. Speech-to-speech models are the future. A single model that natively understands audio, processes it, and generates audio output. No pipeline, no handoffs between systems, no cascaded architecture. Just one model doing everything. Simpler, faster, more natural.

Except the people who have been building in this space the longest are starting to push back on that.

The cascaded pipeline has been the standard way to build voice agents for a while now. It combines a speech-to-text model, an LLM, a text-to-speech model, and a set of supporting models like voice activity detection and turn detectors, all working together to create a conversational experience. It has its challenges:

  • Latency
  • Interruption handling
  • Naturalness

These are real problems that speech-to-speech models were supposed to solve. But some of the most experienced teams in the room, the ones handling millions of calls a month, are making a different argument. Speech-to-speech models are trying to solve a problem that does not actually exist. With the right technical optimizations, the right configurations, and the right approach to voice cloning and TTS tuning, you can solve for all of those things within a cascaded setup. The architecture was never the problem. The implementation was.

That does not mean speech-to-speech models have no place. It means the conversation is more nuanced than most people think. And if you are building a voice agent today, the choice of architecture matters a lot less than understanding where your agent is actually underperforming and why.

Faraz spent two days at the AI Engineer Conference last week, in the room with developers building across the voice AI and broader AI agent space. The conference had an entire full day track dedicated to voice and real time AI.

That is exactly what Bluejay was built to help you navigate. Whatever architecture you choose, we help you figure out what is going wrong and how to fix it!

As a recap:

  1. Bluejay is a testing and monitoring platform for Conversational AI agents. Companies ranging from Fortune 10 enterprises to fast-growth startups in the Silicon Valley use Bluejay to make sure their voice and text agents work in production (monitoring) and development (testing) environments.
  2. Our team, now ten strong, works around the clock to make sure your agent behaves when talking to customers.
  3. This newsletter is 100% human written. It always has been, and it always will be. Ask yourself about what you are consuming. If the writer hasn't read it, why should you?

Give Humanity the tools to trust Artificial Intelligence.


upload in progress, 0

Announcements


Heres what happened at Bluejay last week:

  • Faraz attended the AI Engineer Conference last week, connecting with AI developers across the voice space and the broader community.
  • Rohan sat down with Krishna Gupta, CEO of Presto, to talk about what it actually takes to deploy voice AI in one of the most demanding real world environments out there, the drive-thru.
  • The team pushed 33,045 lines of code this week to make Conversational AI more reliable!

Skywatch Podcast Episode

Inside the Drive Thru: Building Voice AI At Scale | Krishna Gupta, Co-Founder and CEO of Presto AI

This week Rohan sat down with Krishna Gupta, co-founder and CEO of Presto AI, to talk about what it actually takes to deploy voice AI in one of the most demanding real world environments there is, the drive-thru.

From spending fifteen years in voice AI before most people knew it existed, to building vertical AI that lives inside the customer's world rather than on top of it, Krishna shares what most teams get wrong when they try to bring AI into quick service restaurants.

A few things from this conversation worth paying attention to:

  • There are so many things we do today that create value driven off of conversations, and they are all suboptimal. Krishna's thesis is that voice AI is the unlock for changing that, starting in the drive-thru and expanding from there.
  • Vertical AI wins because it lives inside the problem. The teams that fail are the ones that try to generalize. The teams that win are the ones that know their customer's world better than their customer does.
  • The neighborhood diner used to know your name, your order, and your story. Krishna believes voice AI is how we bring that experience back at scale.

Full episode on YouTube and Spotify now! 😄

Behind the Build

Engineering shipped:

  • Zero downtime deployments are now live. The platform no longer goes down when the team pushes new code. Every deploy is health check gated so the new version has to pass before the old one is retired.
  • Sequential test delays are now available. You can now delay tests between steps by an amount of your choice, directly in the general settings of any simulation.
  • Agent timeout is now configurable. You can set your own agent timeout directly in the general settings of any sim run.
  • Run comparison is now live for customers letting you compare runs across multiple simulations or within the same simulation.

Feature Spotlight:

A few weeks ago we talked about how real customers never call from quiet rooms. They call from cars, waiting rooms, busy streets, and restaurants. The background noise is always there and your agent needs to be tested against it.

This week we gave you a way to actually measure it.

Bluejay now has a full audio fidelity suite built into every conversation. You get an overall audio fidelity score plus six sub scores that break down exactly what is happening with the audio quality on each call. Loudness, clarity, background noise interference, and more. All of it visible at a glance without having to manually listen through every recording.

The result is that you can now see which calls had audio quality issues, understand exactly what caused them, and catch the patterns before they become a customer complaint. If a particular environment is consistently causing problems, the data will show you. If a recent change introduced an audio regression, you will see it before your customers do.

Head into any simulation run and look for the audio fidelity score. The six sub scores are right there alongside it.

Your agent might sound great in testing. Now you can actually prove it sounds great in production too.

Give it a try and let us know what you think.

0:00
/0:18

Audio Fidelity in action.

That's all for now. I'll see you next time!

Azfar Khan
Storyteller @ Bluejay

upload in progress, 0

Subscribe to The Bluejay Times

Sign up now to get access to the library of members-only issues.
Jamie Larson
Subscribe