Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Production voice agents in 2026 still rely on cascaded pipelines (STT → LLM → TTS) because end-to-end speech-to-speech models lack reliability, accuracy, and guardrails for complex enterprise use cases. The panel of builders from Decagon, Vapy, Retell, Daily, and Smallest AI agrees that while speech-to-speech feels more natural in demos, it fails on factual grounding and interpretability, making the cascaded approach the practical default. The real engineering challenge is managing the inevitable trade-off between intelligence and latency, often by parallelizing pipeline steps, using filler words, and evaluating rigorously with open-source benchmarks.
Key points
The current state-of-the-art for voice agents is a cascaded pipeline: speech-to-text, LLM, then text-to-speech. Speech-to-speech models are not reliable enough for production because they hallucinate facts and lose interpretability.
Speech-to-speech models (like Sesame's demo) feel more natural and can handle emotional tone, but they often fail on accuracy; for example, they may accept an incorrect date without verification.
Turn-taking (knowing when the human has finished speaking) is a non-trivial problem; voice activity detection and smart turn models are used to distinguish pauses from finished turns.
For outbound calls (e.g., debt collection), the bot has more control and can drop calls if the human goes off-script. For inbound calls, the bot has less context and must handle unpredictable queries, requiring more complex workflow logic.
Reliability of frontier LLM APIs (like OpenAI or Anthropic) is a concern; teams build waterfalls of models so that if one provider goes down, another takes over without the user noticing.
Fine-tuning smaller models (e.g., Smallest AI's Electron) can be cheaper, lower latency, and more reliable than relying on large frontier models, especially for repetitive or predictable use cases.
Multilingual support is easier with cascaded pipelines because you can swap individual components (STT, TTS) for language-specific models, avoiding the complexity of training a single speech-to-speech model for all languages.
Evaluation and benchmarking are critical; Smallest AI publishes open-source turn-based benchmarks for STT, LLM, and TTS that teams can run against their own infrastructure to compare performance.
A hybrid architecture is emerging: use speech-to-speech for the main conversational loop (for speed and naturalness) and delegate complex lookups (e.g., return policy) to a cascaded sub-process that runs asynchronously.
Managing latency involves techniques like parallelizing pipeline steps, offloading instructions to specialized agents, using small language models for fast intent classification, and injecting natural filler words while waiting for slow responses.
Tools mentioned
Techniques
- Cascaded pipeline (STT → LLM → TTS)
- Voice activity detection for turn-taking
- Smart turn models to determine completion of speech
- Waterfall of models for provider failover
- Parallelization of pipeline steps (fork STT output to multiple LLMs)
- Context compaction to manage context windows
- Filler words (e.g. 'let me check that') to mask latency
- Hybrid architecture: speech-to-speech for main loop, cascade for complex lookups
- Offloading instructions to specialized sub-agents
- Streaming to small language models for fast intent classification
- LLM-as-judge for evaluation
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
All right, we're in the remote studio. This is a special one because we're launching a fourth podcast on Dayton Space. Uh people don't uh necessarily keep track, but we cover different
things. Basil uh came across my radar because you're doing all these like dinners and gatherings and panels and podcasts on for deployed engineering. You hosted the FD track at AIE and it
did super well. So, welcome to the pod. >> Yeah, thanks. >> So, you've been running your podcast for a while. What what made you decide to focus on FDE and you know what's a
what's your typical self intro? >> Yeah, so I started as a product manager at in I was specifically working on credit karma for a couple years. I worked at a small venture studio after
that and then I ended up starting a consulting business called Exoflop Labs where we were working with like retailers, insurance companies, that sort of thing just building agents
essentially. This was like 2 years ago. So it just felt like a natural extension of a lot of the product work that I was doing earlier. It was like, hey, I'm working with customers. They have a
specific type of project that they want built for some use case and I'm going to understand what they want. I'm going to make trade-offs and I'm going to help them build that. So, yeah, I just felt
like a natural extension of what I was doing. And so, back in January, we started working with a couple private equity firms and a lot of people were just like, "Hey, like I'm following on
what's going on on Twitter, what's going on LinkedIn. I don't know what's marketing BS and what's not." So, I was like, "Why not just bring on people who are working on the cool stuff at cool
companies? we will talk about what they're working on like we'll do a deep dive and we'll record it we'll post it and hopefully people can take learnings away that they can apply to their
businesses and that's why I ended up starting doing a lot of the fireside panels and the podcasts that I've been doing since then >> can you give like rattle off just your
greatest hits like what you've covered >> yeah so our very first episode was on the future of agentic engineering so we brought on companies like factory cognition composio sound grap we've done
panels on voice agents which is actually the one that we're going to be showing here. We've done some on computer use agents. We've done agents in the enterprise. So, we've done a ton of
stuff just to get nitty-gritty into the details of a lot of this stuff. >> Yeah. And I guess you know, you picked the for for our sort of first feature, you picked the voice agents one. What
are we about to listen to and what stood out in particular? >> Yeah. So, we brought on some people from Decagon, Vapy, Retail, Daily, and a company called Smallest AI. Basically,
we just talked about like what is the state-of-the-art in building voice agents cuz that's a very hot use case. I think uh Sierra even talked about this is this is like one of the most
competitive markets in AI right now is like building voice agents. And so I thought it would be useful to talk about how they're actually built and some of the things that you know engineers who
are building in this space still have to contend with. So for example, we talk about how like the state-of-the-art right now in building voice agents is a cascaded pipeline. So we talk about how
it's like a three-step process. you go like speech to text LLM then text to speech and we talk about like why is that versus why don't we just have a voice to voice model like why don't we
use that like the reality is those just are not super reliable right now we also talk about like the trade-offs that you have to make when you're building voice agents so you have to trade off like the
intelligence of responses that you're getting versus the latency so you can get very intelligent responses but you're also going to trade off it's going to be a slower response and so for
different use cases that might be good or bad we also talk about like reliability so there are some LLM companies that may not be like super reliable in terms of like their
infrastructure. And so like what do you do when opus goes down? Like you need to have a waterfall of models to basically pick it up so that your voice agents don't just stop working. We also talked
about like turn taking and how like it seems like a trivial problem where hey like you're talking to a voice agent and let's say I pause. So how does the voice agent know that hey like I should I
should interject and I should actually like respond to you versus oh the person is just taking time to think. So turn taking is actually like not a trivial problem to solve. And so we talk a
little bit about that. And then also we even talk about like how um I think Pipat built a voice agent for the World's Fair, right? So we even talked about that a little bit and how like I
think I mean you can talk about this a little bit but I think like you guys built a voice agent to like if anyone calls the voice agent it'll answer questions about the conference and then
you got a ton of like real world information and turn that into a benchmark. >> I mean to be clear I have no idea actually. I never looked at the
analytics so I have no idea how many people actually called. probably you know people just kicking the tires you know like it's not it's not that serious like we it's a conference people know
what it is yeah but you know I think it was a good deployment of use case daily is a sponsor so why not >> I you know I I would say that vast majority of voice stuff is support and
like yeah oh my god there's you know so many call centers out there hundreds of billions of dollars spent on this stuff and they're all you know bad so hopefully we can sort of raise the the
state of the art which you know I think that's why Decagon and all these things people exist To me, you know, it's not as interesting if a bunch of people who are obviously selling you voice
pipelines telling you that voice pipelines are the state-of-the-art. Like, yeah, duh. But, you know, having having the the people actually focus on customer support and all those things
also basically conclude like, yeah, like the models are not there yet. They may never be. And actually, this is just the way that you got to do it. And the the people that really feel it are not the
researchers cuz researchers just always want bigger models to solve everything. And the product engineers will try to do it, but the FDES, the people actually dealing with customers will be like,
"Dude, they made this like horrible mistake. Like, this can never happen again. How can you guarantee me that?" Well, >> yeah. Yeah. Yeah. So, like actually we
even talked about that a little bit. So, like we talked about like the inbound versus outbound use cases cuz we even did that once. We like ran an outbound use case and like what you find is
whenever somebody like picks up the phone and you do an outbound use case, a lot of the times they just hang up as soon as they realize it's a bot. So we even talked about that like how do you
how do you solve for that? >> All right. Uh that's a good teaser. I you know I admire your work. Uh I'm excited to feature on space. We'll be doing more uh in future together but uh
this is just an intro to FDE for uh at least for at least a l space audience. >> I think this might be a good place to give everyone a 101 lesson on what does a simple voice agent architecture look
like. So maybe that's a good question for Brun because you guys uh run Pipecat. So >> yeah. So there are many ways to build one nowadays because the complexity is
the nature of things. But the simplest one which is called a cascade model is typically uh voice input. Uh it comes through a transport. It could be web RTC, phone call or web sockets. It goes
into the mo into a speech to text. Uh the transcription happens. There can be some additional models right there for background noise removal, voice isolation. So if there are few people in
in the foreground, it will isolate to one person, remove the background noise, it there's a turn detection model, unlike text where you know when you're done, you press enter and the message
goes and the LLM starts to like execute with voice there is no it's not walkie-talkie or pushto talk. So there is no additional uh information apart from like if you're
looking at the face you could actually figure out that the person stopped speaking but typically use something like a voice activity detection and a smart turn model to like figure out that
the turn is complete. Um if I pause mid mids sentence that's different from like if I finish speaking and pause uh you may use a speech to speech model right there. Uh then when you have the text
you uh from the transcription then you send it to an LLM. The LLM may respond with an inference. It may actually tell you it may infer to tool calls. It may do a bunch of things there. You get that
when you have the full output you text to like you do uh text to speech DDS and uh it's usually faster than real time. Uh so then you have to like figure out how to like stream it back to the end
user. It'll go over the same p transport that that you spoke over and then it'll play it back on your speakers. Uh but things are getting more complex as people, you know, this was like 2024 25
people wanted to do this. If you're doing an outbound use case, like you do an outbound call, debt collection is quite commonly a use case for this. Um you get a call, they tell you you have
like an unpaid bill. um they ask you if you would pay it or you know give you a certain set of options. Uh for that the bot has really like if you go off the rails like the human goes and starts
saying random stuff it can just drop the call. It's not obligated to stay on the call in any way. So guardrails are easy. Um it knows what the inputs it can accept. It knows what inputs it can it
doesn't need to respond to. It flips when you have an inbound use case because the bot has no context like has some context of why it exists but doesn't know why you the person who's
calling it is if you're a mom and pop shop like a flower shop or or like a barrier or something like that okay there's only like finite things that you can do there but if you are more complex
let's say you're Amazon right there like a billion products on your platform um and you have different policies for different things is it uh you know is it a physical item is it a cheap physical
item is it uh an expensive item u is it an electronic item like all of that so it needs to figure and then when you said something like you know you can't put all of that inside a single prompt
because it will like LLMs are you know goofy in that sense that always remember the first 4% and the last 4% and everything in between they kind of forget uh so if you have a return policy
right in the middle of that context, it's very likely it'll hallucinate. So you have to do more smart things and that's where you have like more models come in. Uh you can have a compaction
model like we're all using coding agents now, right? So you know like if you're using like a 1 million context model like at 25% you should become nervous at that 25% mark. is nowhere near 70% mark
but at 25% mark you're like okay should I start compacting and like saving my work so that this thing does not go off the rails right the same thing I think voice AI has been like one generation of
ahead of coding agents like we've all in the last two years solved things that felt like so alien but very tangible for us which now when we see coding agents do we like man compaction we we were
doing compaction from day one like We knew like after five tons or 10 tons and your context was only 250 tokens or 250,000 tokens or 50,000 tokens, you know, models were really small like
before and at 10,000 tokens it would like go off the rails. So, so a lot of our work is I think about like making sure that the bot does not hallucinate and trying to like keep track of the
conversation as the the turns progress. And the turns are fairly short because the either way the bot is more likely to speak more than the human did. So you have to like kind of keep track of like
what did the human say? Where are we in the conversation? >> Yeah. So um I guess this is a question for one of you three. Uh so do you guys use this like cascading model for your
guys' voice agents? >> Yes, that is one of the major offerings that we have. Cascading and speech to speech. >> Yeah. I was going to ask like that
sounds too complicated. Why not just go voice to voice? Like why not? Like why do you have to have this like crazy three-step process? >> I feel like the easiest way to to answer
this is you can call into a, you know, really nice voicetovoice demo and you're like, "Wow, it's like listening to me laugh and it's like responding to my tone and it's so snappy. It's so fast."
But then, you know, I I tell it that it asked me what day I want to schedule my appointment for, and I say, you know, next week. And it says, "Great. Uh, is that June 10th?" And I'm like, "No, it's
the year is 2030." And it's like, you're right. It is 2030. So, let's schedule this for June 10th, 2030. And I'm like, great. Sounds good. Um, and so from that perspective, I am actually very curious,
especially on the on the smallest side, um, how that's evolved over time because we are very much keeping our eye on how these speech-to-pech models are performing. Uh but overall I think our
current stance is uh the cascading model just allows you to enforce so many more um rigid guard rails and just tight control over the ability to say hey this input is going to go through the same
supervisor models to detect for you know prompt injection or social engineering. Then it's going to go down the conveyor belt into intent selection. Then we're going to uh optimize context by checking
conditions ahead of time to say, you know, in this complex um process that we're going through with this user, uh maybe this prompt is applicable if they're this type of customer, but this
prompt is applicable if they're that type of customer. So let's, you know, figure all this stuff out ahead of time, compact and compile a good system prompt for our message generation model, and
get a response back. And then we can take that response and we can check a whole bunch of other things, right? uh we can check is this grounded in truth, right? Is it 2030 or is it 2020 20
whatever year it is 2026 um and then it goes back out over the line, right? And so the obvious constraint to that is how do you make that performant, right? How do you parallelize as many of those
steps in the conveyor belt as possible? I think the last six or so months has been uh a really amazing feat from our engineering team at least to find and shave off like 10 milliseconds at a time
across every single part of this pipeline to make it feel snappy even though there's a lot of things going on behind the scenes that uh you know you don't necessarily have to do with a
speechtoech model. So that's at least my take but yeah I am curious for for the rest of the group's thoughts on this. Um you want >> yeah um no I think you brought a
question around like cascaded versus speech to speech like let's talk about like why did people start thinking about speech to speech and so the initial idea was simple that hey if you convert
speech to text it's going to lose the emotional information right so if you say hey um you might be saying it in a sad way or an excited way the bot is going to answer in the same manner right
so that was the obvious reason people started I think uh at some point of time at least at smallest like that has evolved to uh speech to speech is uh a more natural way in which the human
brain operates. Uh so for example uh when you do the cascaded thing you do speech to text then you send the prompt to an LLM and then it responds right so we call that a synchronous architecture
like it's happening one after the other but our brain is thinking while listening so as I'm speaking to you you're already forming your thoughts and if I'm talking for too long you'll
interrupt me right or you might be taking notes in the back end right so you might be essentially doing tool calls while uh I'm speaking to you. And so the whole idea is that if you ever
want to pass the Turing test of how the human brain operates, you need something that is working asynchronously and can not just understand emotions but also operate uh like take in speech natively
and give out speech natively asynchronously and and so that's how why we have been sort of building Hydra. Now Hydra is our speechtoech model. Um now in terms of accuracy and and
interpretability and and all those things I think whenever there is like a new architecture that comes out it's often sort of good in one parameter and then regressed a little bit in other
parameters right like so for example um the speechtoext accuracies for a just a speechto text model might be way better than the encoder of a speechtospech model that that you have and so I think
The challenge is that while you made progress on the um you know uh making it more natural and operate more humanlike etc etc how could you keep the accuracy bar the same and so I think a lot of
that comes down to interpretability of um such models like because you don't want it to be a black box so how can you understand where it is lacking and um so that that's a lot of research that we do
and how to make speechtospech models more interpretable Um the other constraint is I think initially the speechtospech models like I think sesame etc that came in that
they were just speech to speech like they literally took in speech and give out speech ours is like multimodel so it takes in speech and text gives out speech and text both so it can do tool
call and uh take in text parallelly. So if you want to put guard rails if you want to do all those things that constraint does not go away. Um what we are also seeing is there is still at
least in enterprises a lot more cascaded deployments compared to speechtospech model but speech to speech I think will be like an eventual future that that's my opinion.
So my take is that when you have like competing things you end up in a hybrid and uh I think the answer for the midterm will be some form of hybrid because the speech to speech are
improving. There are parts of the conversation loop where you would say like oh my use case of my workflow is complex. Uh I'm going to use multiple LLMs anyway. So for the active loop like
I'm talking to as a human you called me. I'm answering your questions. But whenever I have to do some kind of lookup or something, I delegate to another LLM which will then do the
cascade stuff. So the voice in the first part of the loop keeps running. Uh and then you have like interesting kind of semaphors or like you can think of the as threads. uh you want to interrupt the
the speech-to-pech model because you realize that this is a complex question and you want to before the speechtospech response you say actually you cannot answer this question right and delegate
to the cascade and and so on so forth so typically I kind of say there are many use cases where you know if you wake me up in the middle of the night and ask me a bunch of questions there are
definitely some class of questions that I can answer without thinking and we all do our jobs in a certain way where we can in autopilot like 50% of our time, right? So as these use cases become
emergent and uh you know you're fully deployed with a customer doing high volume, you could essentially train specialized models that fully understand like that 50% use case very well. Uh and
you could you could always ask the model like hey what is my return policy and say like in the simplest case this is my return policy and it applies to like 70%. And I know which SKUs or or
products it belongs to or it applies to and if you ask me outside of that then I have to like do all the crazy work and if I don't then I answer it right. So, >> so on these cascaded pipelines, what
what models are you guys using let's say of the frontier models of the you know GPTs of the of the uh what opus and sonnet are you guys using the latest ones or are you using you know like GPT4
because it's like the right balance between speed and intelligence. If you turn off thinking then you can actually use any of the like if you want a fast model which responds to like the prompt
and does not need to think because if you wanted to think you actually start to think about um paralyzing the work because you want the the first model to be super fast. I still like my Gemini
2.5 really well. Like it's so fast. It's like the 3.5 which they launched is like okay it's slower but the 2.5 is so good. Uh even the haiku is really really good. uh they're still slower than what you
would expect, but yeah. >> How do you parallelize that? >> How do you parallelize that pipeline? Because it feels like, oh, well, I do have to first know what text the person
said and then the LLM needs the text to do anything. And then you can only generate speech once the text has been generated that you want to wanted to send, right? you you just spend the you
know you can send the speech like in the speech to speech model case you would send the speech directly to one place fork it into another place run the st there um and if you want multiple models
the text output from the first forks into multiple LLMs you can think of it as like a waterfall like you know it comes down >> because you don't know what the output
exactly will be >> yeah you put a put a gate at the bottom which XR or whatever fancy thing you want to So which says like the if all of them say we are right then you put
another LLM downstream which one is better like you can go really crazy if you love computer architecture 20 years we had nothing to do now you
can do all this crazy stuff from 20 years ago. >> Yeah. Um so how do like accents and different languages all fit into this? Maybe Sudarian because you guys are
building speech to speech. Um yeah so at least for speech to speech right now we are just fully focused on making it more intelligent in English and you know getting it really good in English like
we don't want to do any other we don't want to introduce any other variables because the technology in itself is I would say quite frontier asynchronous is not yet like mainstream etc right uh but
in terms of like cascaded yeah we've seen like a lot of demand in so we have a lot of presence in India so we see like a lot of demand from India Um we see a lot of demand from Latin
America etc. Um I think US is more like mostly English and then there's some Spanish and then there's a lot of accents to it. Um and I think yeah US at least we have not had any troubles in
terms of um technology. I think it's yeah I mean noise cancellation is probably the last mile of problem that is pending uh in terms of handling those things and but yeah that has nothing to
do with accents I'm curious what you guys have seen >> u maybe I can add uh some examples about that um because on multilingual uh I like to definitely leverage cascade
model because there are multiple levers you can pull uh to talk about an example um it's a customer we were deploying uh for Japan Um, so you could typically go uh you
know to your uh 4.1 on OpenAI or your deepgram for transcriber and that typically worked well. Now the challenge became on the voice uh piece. So it turns out uh what we thought it was
state-of-the-art for um something like English, Spanish, Portuguese um those type of uh that are typical in the US didn't work at all. Um so one of the benefits uh about this is that you
can swap. certain pieces that doesn't quite work for you. Um, at least on Vapy, how we solve it is that we let customers bring their uh custom uh text to speech server. Um, so they actually
have like there's some startup there's a lab um in there and there's typically uh you will find that in every market. Um, Arabic is also really hard to get it right uh in pronunciation of brands,
addresses uh and that really we don't uh cannot build expertise and optimize for every single use case. So that's one of the reasons why I like that as well. Cool. Um, also actually question on the
like sesame model. Yeah, like they came out with this like really insane like demo and then I've never heard of them since like what do you guys know like what happened and what's going on?
>> Yeah, I think uh so they are building hardware is what I understand like they are putting uh those the voice into some sort of a hardware. I'm not sure if it's glasses or what are they working on
exactly but uh I know a few people who got hired there and they're all sort of focused on the hardware and voice sort of backgrounds. Yeah, the answer I got was that the CEO had already made a ton
of money and he just really wants to play around basically. So I guess he can do whatever he wants. Um, cool. Um, so is it standard practice for in this like cascaded pipeline for you to just
give just have one giant like system prompt that you're giving the LLM on these are all the rules for how you should be responding? um or is it more of a um like a workflow approach? How do
you guys think about that? What's the standard best practice right now? >> So I think it highly depends upon the use case and also the size of your prompt, right? So I think I have a
counter approach to Baroon in in terms of that like usually in inbound I find you know you you're usually having dedicated lines and you know specifically where the workflows are
going to go right and so in these cases if you know a little bit more about and you can predict where the conversation's going to go then you can have a more of a node or graph builder kind of approach
but if you don't know what's going to happen typically I think in outbound yes there is a dedicated like message but what the caller says back you never know because they could just be frustrated
they could be angry or they could be I don't know like annoyed and in that case when they're erratic when a person's erratic you never know so having it in a way that is like all in one one prompt
allows you to have an more like a central brain in terms of I can pick up components from my prompt that may not have been in a dedicated flow but actually can be utilized iz to enhance
the response. And then on top of that, you have to think of your knowledge base and knowledge the things at your at your expense, right? Like in a inbound flow, maybe you know exactly when you need to
pull a specific part of your knowledge. So you don't need to extract every single piece of context, right? Whether it's your websites and documents, but sometimes in outbound, you should
utilize all that information to get a better response. then you will rely more on cosign similarity and stuff to make sure that hopefully your retrieval is good enough to respond to them.
I think from my perspective um the limiting factor has been and and will continue to be although it's getting a lot better over time just uh to you know some of the earlier points made uh the
ability for these uh frontier LLMs to be able to take a massive wall of text and not just read the first three lines and the last three lines but actually be able to properly uh reason and like
localize themselves into you know we are halfway through this procedure and these things have already happened and these things haven't happened yet and these things are true and these things are
false about this situation and so therefore like this is the one thing that I should be really zooming in on right now right um so from our perspective I think um we kind of take
an approach that depends on the customer um but we do within our execution engine have the ability to um make some of these rules or prompts or guidelines more situational so that we can you know
check if these are uh relevant right now before we even uh include include them or or uh uh don't include them in the system prompt. And so that was our solution that was especially relevant,
you know, like 7 months ago when the LLMs really were just like consistently skipping step 4 A1. Um but we have found over time now that uh sometimes the reliability of deciding when these
things are relevant or not is actually worse than just giving all of it to the model because the models have improved a lot since that time. And so now we're going back, I think, swinging a little
bit back towards just give the model everything and more or less it will figure it out. Um, but at the end of the day, you know, with nondeterministic systems, it really just ends up being
customer by customer. uh we have to really understand their use cases and at the end of the day just build it and then test it a whole bunch of times and then use LLM as a judge to tell us you
know did it work was it reliable or do we need to add a little bit more um pre-processing and and context optimization ahead of time in order to make it reliable. So the the real answer
is most likely always just test and find out and then test again and then find out again and you know it goes on forever. I'm assuming that yeah maybe the models are able to handle those edge
cases now like just in a giant prompt but that probably also increases latency. So first could you talk about like what is a good latency look like for a voice agent and is that even true
that yeah the models are getting better but they also require more time and so for the end user that's actually a bad thing. >> Yeah that is a really good question. I
mean I I think the definition of what like benchmarking latency stats look like that are good have uh changed a lot and um expectations are you know getting better and better and better from our
customers. Overall, I think it it is also important to mention that um as much as fillers and contextual fillers is something that uh we don't like to hear just as a noun, um it is normal for
in normal conversation for humans to say a few words while they're thinking. And so we've actually found that when we have zero contextual fillers, sometimes it actually feels more rigid than when
we don't. And so we tend to find that uh having the the right amount of fillers that um kick in when you know we are experiencing some sort of latency because of some very large prompt or
some very complex thing that the model is thinking through uh most commonly you know running a tool call where our customer's API becomes the bottleneck and we're waiting you know 5 seconds for
something to come back. Um being able to say hey just give me a sec uh looking at that. Okay, cool. Here's here's your answer, right? There's actually nothing wrong with that. Uh when we've when
we've tested with our customers, um obviously the the limiting factor in that case then becomes how good are your fillers and how natural do they sound. Um and so that's been a place where
we've invested a lot of time as well to say how do we create an execution engine where we really don't need fillers, but in those moments where we do, they're really good. Um, so that's at least my
take, but again, I know folks have probably a lot of opinions here about latency and and how to find that trade-off because in my opinion, I honestly think that that's the hardest
problem to solve in voice deployments is just that trade-off that's always going to be there of like performance versus um stability and reliability in in general. So,
>> um, yeah, I agree with all of that. Some of the ways that we think about it in Vappy, uh, is offloading, uh, some of that instruction. So break it down. Uh sometimes you don't really need um that
you know 10 15 step workflow in the in the prompt. Uh one example I can talk about is um this uh collections uh use case. Um we do allow users to capture um businesses to capture credit card
information. Uh but not every call is uh about that. Some people already have a payment method maybe added or that. So uh one way of doing that offloading into a uh specialized agent uh so we can
start thinking about multi-agents and there are many architecture and patterns in there um is to constrain the uh instructions that you give it to to that particular moment of the conversation
and once that achieves its goal it can go back to that u more free form approach which has loaded like uh your longer uh context. Something we're also been uh experimenting lately um and it's
more about the guar space is um to uh your point where you can have uh think about it as a waterfall you can actually stream to multiple places. So what if you have uh streaming down to a a small
language model which can uh do inference in a very little time. So think about classifiers to maybe collect the intent um while the pipeline is still maybe adding a filler word as well to keep a
consistent experience but uh kicking off a background process that is still intelligent still um thinking. So those are still things that we'd like to invest a little bit more research on. Um
so yeah constantly investing on that. >> Sure. I think of one thing that we haven't also factored in in terms of the latency uh tradeoff is cost right now that you're putting such a big large
prompt into your your LLM. You have to think about those calls where many people just hang up and predominantly calls just hang up 10 seconds into the call. Right now you're having to pay for
all those tokens to just get inputed in. So that like Stephen mentioned splitting out your prompt, you are also playing to that strength of lower latency and lower cost.
I'll speak just to the benchmarks. So we publish u turnbased ST benchmarks, LLM benchmarks and TDS benchmarks. The whole point of that is that as I think today Neumatron 3.5 launched and there's
already an ASR um benchmark out did very well the Nvidia one. So um the idea there is that all our tooling is open source. So with the HTT benchmarks, the TDS benchmarks and the LLM benchmarks
all the turns are actually open. So you can run it against any new model that comes out. Uh you can look run it locally. you can look at the outputs from from our graphs and compare that to
what you're getting on your infrastructure if you're self-hosting or you know elsewhere. Um so that's one important aspect to like uh consider if you know what your turns will look like
you can just update or you know just download our benchmarks update the turns and see how that like goes for your own flows. Uh usually for the LLM stuff, it's really useful because you can throw
out the ST and the TTS and just run the LLM inference loops to like check against uh a customer. So yeah, I think eval are really important. Um, and if you have like a like while talking to
customers, if you can evaluate like what type of conversations or like what type of workflow they're going to have, like those are real direct inputs into like what you can directly start testing even
though you're not live with them or even they've not shown any intent. Uh, so I think that's where like the sees and the forward deployed engineers can actually come to for like you we all have cloud
running in the background, right? So you can just say here's what we had a discussion like can you build an eval like based on our eval suit like can you like just run a smoke test with the
conversation information that we already have and that helps a lot I think. Yeah, just just adding to that like we are seeing so we while we focus on models we also have all the model sort of
orchestrated on our platform. So you could build a voice agent on smallest. Um we are seeing a lot of customers who just pair our models with our own like small language model. It's called
electron. fine-tune one of those to make it work for realtime voice use cases and see a lot of folks using that with better knowledge basis memory etc over GB 40
4.1 you know the popular realtime models um just a from a cost perspective it's lower b from latency perspective it's like way lower u and then just more reliable I think I'm not sure if
everyone faces this but like um open AI APIs spike and in um you have no control on those latencies. Um with if you sort of have a self-hosted model um you could scale up scale down based on your
requirement. So that's one thing we are seeing a lot with our customers.