Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
No summary yet. Click “Regenerate summary” to queue one.
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
[music] >> Hope everyone's having a great conference. Super excited to get started. Uh how many of you all just quick show
of hands have ever gotten a proactive healthcare call from your provider? Yeah, looks like no one. Me neither. Uh I'm uh Vivek of a few hands over there. I'm Vivek. I run uh engineering at
Hippocratic. Uh we've built a product that calls patients and can have clinical conversations. Uh and we're over 200 million conversations in at this point.
Uh here's the reality of like healthcare uh across all of human history, the entire system has been built on scarcity, right? Not enough clinicians, not enough time, not enough money. And
hence the word triage. We're always trying to fig- figure out who amongst us is the sickest and deserves to get care and attention. As an engineer, I like to think about
the math. And for most of human history, the math has just never been in our favor. Until now. We're finally at the point where we've seen enough of technology
progress. And we also have AI systems that can do clinically safe conversations. And the cost for these conversations continues to drop. So the math now flips. We can stop rationing,
and you don't have to have calls just for the sickest 5%, but you can call everyone. So that's the mission. We're trying to build clinically safe abundance for all.
And it's not about replacing clinicians. In fact, what we're trying to do is help clinicians reach everyone that they couldn't do before because there was simply not enough of them to do this
work. So how are we doing this? Um our product and our company is predicated on three things. First, do no harm. Next, patient first, and access
for all. Um and here's a cool thing. Uh this is plastered across our walls in our company, uh but it's not just a poster. Uh this is an oath every employee in the
company takes. I did too on on my first day. Uh and we built our product uh based on this oath. Um and this is working at scale. Uh at this
point, we've had 200 million clinical interactions. We've had zero significant safety incidents. We've deployed in over 60 plus health systems and have an 8.5 on 10 patient satisfaction rating. Now,
uh I could put an architecture diagram and explain how all of this works. And don't worry, we'll do that. But before we get there, uh I want to show you the product in action. So, we'll have one of
our agents interact with the patient so you get a feel for what we're up to. >> I'm your AI care manager. I'm calling to check in on how you're doing with your health monitoring devices. [music] Do
you have time to chat now? >> I I need to know if I'm talking to a human being or a some kind of computer. You sound like either one.
>> I understand your concern, miss. I'm an AI, which stands for artificial intelligence. >> Okay. >> How are you doing today?
>> [laughter] >> I was tired earlier, but I'm somewhat revived. I'm incredibly amused. >> Delivering at this level of care
required building clinical capabilities you won't find [music] anywhere else. On the safety side, accurate drug name recognition, >> Could you tell me the name of the
medication you're currently on? >> Prazosin. And that's one of them, but I don't remember the name. >> Could you help me out by spelling the
name of this medication? [music] >> P r o z o s i n, I guess. >> It seems like you might be talking about prazosin. Does that sound >> Vital sign clarification,
>> Let me say it again it was 121 over 45 91. >> Just to confirm, your blood pressure is 121 over 45 and your heart rate is 91,
right? [music] >> Right. >> Med stoppage detection. >> What about the nifedipine tablet? Were you able to stop taking that one as
instructed? >> Well, I took it this morning. Well, no, I took it last night because I didn't have the other prescription. And I don't think that would keep my blood pressure
down. >> I understand your concern about your blood pressure and it's great that you're thinking ahead. I would recommend that you do talk to your primary care
doctor about the nifedipine and get their guidance on what to do. >> And real-time nurse escalation when clinically necessary. >> How are you feeling?
>> Okay, just a little tired and weak. >> It's quite common to feel this way after being in the hospital. Could you tell me more about it? >> My legs feel weak when I walk.
>> Got it. Along with the weakness in your legs, have you noticed any chest discomfort or difficulty breathing? >> I just was this was having shortness of breath for a while.
>> I'm so sorry to hear that you're experiencing shortness of breath. Given your symptoms, it would be best to have a nurse speak with you immediately. >> Yeah, so that's our sneak peek into our
product and so we'll get into why this is hard and why a generic stack doesn't work for us. Okay, so if I were to strip out all of the complexity around the audio and
speaking back, it boils down to two things, intelligence and latency, right? On the top left, you see some of our ever-evolving smart intelligent models. These models are fantastic at what they
do. They're clinically safe. But unfortunately, many of these models take tens of seconds to respond, sometimes over a minute. And that's completely useless when we're trying to have a
two-way conversation on a telephone. On the bottom right, we have models that are fantastic. They're really fast, but then they're just not at the same level of clinical accuracy for us to have a
safe conversation. When we started building out this product, where we wanted to be was on the top right corner. And in order to accomplish this, what we had to do was
build a vertically integrated stack ground up bit by bit, optimizing every part of the stack stack. When we started out building our own models, we also had several seconds of
like latency, but hundreds of optimizations later, where we landed at was a product that is insanely fast, but also doesn't lose its intelligence. And we continue to benchmark all of this
consistently. All of these results are on our website and more. Um and there's three typical ways in which folks build uh voice systems. One, you can use an ensemble of models, or you use cascaded
models, or you use speech-to-speech real-time models. Again, all of these models and architectures are great at different things. As an example, let me take two specific benchmarks among so
many hundreds that are critical for our workflows, uh lab results check and IVR navigation. Now, these are things that you don't typically hear about or see in the most common benchmarks. And most of
the generic models out there don't perform well at those. And we needed our models and our product to be over 99% accuracy on those specific benchmarks amongst other things. Uh and that's one
of the key reasons why we continued down investing our own stack. And latency is key to this entire system. So, every time we work on an optimization, we buy back some latency,
and we just don't bank that latency. We use that extra gap now to pack more intelligence into the overall system, such that we can have a more reliable conversation with the patient.
Um and then we do that, we go back, work on more optimization, and that flywheel compounds. So, what seems like a tug-of-war between latency and intelligence for us is a compounding
flywheel. So, what does this entire machine look like? Uh this is Polaris. This is our constellation architecture. Left to right, you have the system that hears.
The middle is the brain that reasons. The right is the system that talks back. And all of this round trip needs to be really fast. So, on the left-hand side, what you're
seeing are collection of models uh that are used for speech detection. So, everything from bilingual switching to background noise detection um to contextual understanding of the
conversation. The brain isn't a singular model. We in fact run 31 models at any given point of time for every conversation. So, we have one central model that's hard like have a handling
the conversation, where we have other 30 specialist models, everything from labs to medications to scheduling, uh that are feeding input into this model. And then finally, on the output side, we
have a custom personality, voice, HD quality, and clinical documentation engine that make sure the patient gets back all of the right information, and all of the same information flows back
into the health system. The reason we have the system is because we see a singular model being as like one point of failure, and that's just unacceptable for a patient conversation.
Uh this architecture gives us the redundancy and safety that we otherwise couldn't get to. And the patient never can tell the different. It's pretty seamless from their perspective. Let me
double click into some of these systems for you. Uh first, we'll talk about our uh ASR system. So, the real world, as all of us know,
is fairly loud and noisy, uh but most of the audio benchmarks get recorded in a quiet room. And something we learned the hard way was most of what looks like model reasoning failures end up actually
being model mishearing things, right? So, the Spanish C or yes gets transcribed as the alphabet C, or you could have an Arabic drug name that gets over 30% in inaccuracy in terms of like
word error rate. So, we had to build a system to to combat all of these challenges. A typical speech to text system just takes in the audio and then outputs some
text. So, ours is a decoder only audio LLM system. So, we take in the audio, but in addition to that we're giving it two additional pieces of data. One is the context around the
conversation thus far, and second, it's the domain knowledge around that entire conversation. And that is sort of like the trick that makes this work for us. So, under the hood, what's happening?
So, we have First, we have a encoder, which is So, we took an open-source Whisper V3 large turbo model and fine-tuned that on millions of clinical conversations to
get to the accuracy we need. Next, we have a conformer projector. What this does is takes the audio and compacts it and projects it into tokens that the language model would understand. But,
the key over here is it also maintains all the prosody. So, all the pauses and the stresses are maintained. So, the model hears not just the what, but also the how.
Next, we're passing in the context around, let's say, the medications this patient's taking in, or the task at hand. Are we trying to fill in a form? What kind of like details that need to
go in there, and the policy. And how this helps is Now, when a patient mentions a medication name, we aren't guessing from like an infinite list of medications, but we have the
chance to optimize around a finite list, and that helps in getting the word error rate down. Next, we also have contextual biasing as a part of like the training. So, during
the training process, uh we have millions of these synthetic patients with their addresses, phone numbers, uh and other details uh that are fed into the training process.
So, as a concrete example, let's say someone mentions their address on a conversation as 1100 Boulevard. Now, that specific utterance is ripe for phonetic garbage for most ASR systems,
but given we're feeding that as additional context to the audio LLM, we have a high degree of accuracy and are able to get to the right transcription, which then feeds into the brain
accurately. Further, what we see is for most clinical conversations, many of the patient responses are mono mono words, right? Uh and those often get
transcribed incorrectly. A now becomes a no, or a five becomes a fine. And in a patient conversation, that's catastrophic. So, what we do is a secondary round of scoring, specifically
when we see a patient has said only a singular word. And again, we use the context of the entire conversation to do this. So, all of these like help us bring down
the medical word accuracy rate by over uh 50% um from what we see as standard off-the-shelf models. In addition to that, uh the specific system we've built allows us to uh check in the latency,
and at a P99, it's three times faster than every other system out there. Next, let's look into the constellation, uh the brain. So, as I said, we have 31 models live
running in uh at any given point of time. So, how do we do this without falling apart on latency? Uh the answer is a bit counterintuitive. We actually run every single one of these models in
parallel. Uh but the trick over here is uh every single model does a really quick check to see if they even need to say something on this conversation, right?
So, every specialist first decides, "Hey, do I need to speak?" If not, it's a short circuit and that's what helps us keep us in the budget. So, that's the synchronous part. In addition to that,
we also have asynchronous run models running in the background that are doing verification. And let's take two specific examples to walk through how we use these models.
Again, going back to our medication example, for example, if the patient talks about ibuprofen or is talking about another medication, that's when the med engine specialist will kick in
and it will say, "Oh, I have something to say over here." Main model will pass in that context to the main model. The main model with a lung along with everything else that already knows
around the goal and the conversation will take this input from the specialist and steer the conversation as appropriate. Next,
like many other agentic systems, we heavily use like tool calls and I think pretty much every one of us has experienced tool calls have innumerable failure modes.
So, we have these like verifiers in the background that are running to make sure all the parameters into the tool calls and the responses and the tool calls are accurate.
And for use cases like scheduling, this is what's gotten us to get to a 3949 accuracy. And finally, for certain use cases, we're also running these verifiers offline where possible.
So, again, going back to scheduling, sometimes we have the luxury of time to go back and check to make sure if there was an inaccurate appointment, we can course correct that on the back end or
call a patient back and apologize. All of these changes would be meaningless if we didn't optimize our inference stack and we have a fantastic research and engineering team that lives
and breathes on this problem. And for inference itself, quality is our constraint and speed is the work. So, we can never compromise on the quality of the output. Every speed optimization has
to be lossless. We've worked on like many many of these like optimizations. Let me talk you through three specific ones that we found to be the most meaningful.
One is four-bit quantization. So basically what we've done is we've shrunk down the math for model computation to be from 16-bit to four-bit, again lossless, and this has
like really helped us with overall latency. The next is speculative decoding. So we have a smaller model that's generating the tokens ahead of time, and then the main model just
checks this in one go, and this also has greatly improved latency for us. Finally, a KV cache compression system. So as these conversations get like really
long in length, we've figured out a way to keep a large chunk of these conversations warm on cache, giving us an over 96% hit rate, and then also helps us with the pre-fill portion,
which goes 18 times faster. So this style of across-the-board board optimizations is what's getting us ahead during the product development phase. Now, evals are obviously key as we've
seen across many of these talks. So some quick math, right? For example, let's take the scheduling use case again. With over 10,000 calls a day,
a one 1% failure rate like sounds all right. Most agentic systems would claim 80%, 90% accuracy, and that's great for them. For us, even the 99% is pretty bad because 1% error means 100 people a day
are going to get the wrong appointment type, and it's physically bad because they'll show up on the wrong date, wrong time, but worse off, they could actually be
missing a critical appointment, and that's unacceptable from our standpoint. So it's not just an annoyance, right? And the challenge over here is in order to catch this like statistically, the
math like is like incredibly challenging, right? You need about 450 tests to be 99% sure that you can catch this like 1% error rate and and 1900 tests to be able to see that you've
caught it like 10 times. So you can't purely rely on synthetic data from our experience to be able to get to the scale of accuracy. So we use a combination of human beings and
synthetic data. So we have over 7,000 trained clinicians on our platform who are continuously helping us like evaluate our platform. They've done over 700,000 close to 800,000 clinical
conversations to help us figure out the accuracy of our system and we do this on a continuous basis. Finally, coming to safety. We've shipped five versions of our
product thus far and each one's gotten better compared to the previous one. So we don't just grade the output of our models on the correctness or validity of
it. So we use the same scale that as they do for human beings. So they're graded across correctness, no harm, minor harm, severe harm and death. And across five generations where we're at
with our Polaris system is a 99.89% accuracy with respect to like no harm. And humans on the same rubric are at about 81%. It's not because like we're
terrible, but AI systems don't get tired and unfortunately we do and we also don't have the luxury of having 30 plus supervisors helping us at any given point of time.
And finally, safe is not enough, right? A key part of the product is the empathy. For human beings to open up to AI systems, we needed to make sure the voice and how we actually lead these
conversations are empathetic. And as with most other things during our product journey, we found that there weren't actually accurate or good benchmarks to measure the empathy of
these systems. So we actually built one, which is called HEART. The paper for this is public. Encourage you guys to like read it. And we we had to make sure the end system
was not just safe, but also empathetic, so that people would actually prefer it and use it. And finally, all of this like comes down to the
people. We wouldn't be here if we didn't invest in people from an AI native perspective. Everything in our company across the last couple of years has been transformed our hiring processes to how
we use like tokens. We still have unlimited tokens for everyone at our company. We also have two great programs for anyone who comes into Hippocratic. For early career engineers,
we have an agent employment residency program where they're trained to write some of the best like health care safest agents. And then for experienced software engineers, we have an AI
residency program where we help you start like training models and build the next great system along with us. So, yeah. We're told you got to pick two of these options around
quality, speed, and safety. We didn't and we decided to go with all of them and we're hoping the rest of you should come and build with us. Thank you.