Enjoying this issue?
Get tomorrow's AI & engineering digest in your inbox — hand-picked, summarized, and always spam-free.
TLDR
Abridge has moved from clinical documentation to clinical intelligence by using AI to process medical conversations. The company now serves 300 health systems and is building contextual clinical decision support that acts during live doctor-patient visits. The key technical challenge is maintaining high quality and trust in a high-stakes domain, addressed through rigorous evaluation with clinician-created rubrics and a data flywheel of 100 million conversations.
Key points
- Abridge started with AI-powered clinical documentation (SOAP notes) to reduce clinician burnout and pajama time, and scaled to 300 health systems including Kaiser, Mayo, and Johns Hopkins.
- The company is now building clinical decision support that works pre-visit (suggested discussion topics), during the visit (live conversation analysis), and post-visit (auto-generated notes, orders, patient summaries).
- Quality is ensured through a multi-stage evaluation pipeline: internal benchmarks, staged rollout (alpha, beta, A/B testing), and continuous monitoring with expert-calibrated LM judges and online signals.
- For clinical decision support, the system handles underspecified questions by pulling context from EHR data, the live conversation, and clinical guidelines or medical literature.
- Abridge decomposes the clinical note generation into smaller workflows (e.g., per section) and post-trains smaller models for each, reducing latency and cost while maintaining quality.
- The company claims a unique data advantage with 100 million medical conversations per year, which they use to train models that can outperform frontier models on specific healthcare tasks.
- In-visit order generation uses a gating approach: cheaper models detect order mentions, then trigger heavier models for matching and approval, avoiding constant high-cost processing.
- The talk emphasizes that healthcare AI is a high-stakes domain where quality, latency, and cost are all critical, and that Abridge treats evaluation as the operating system of the company.
Tools mentioned
Techniques
- expert-calibrated LM judges
- rubric-based evaluation with clinician-created rubrics
- decomposing problems into smaller workflows
- post-training smaller models for specific tasks
- gating mechanisms for triggering heavier models
- data flywheel from 100 million medical conversations
Stop scrolling. Start reading smarter.
Receive the day's most important AI & engineering updates in one concise email. No spam.
Transcript (captions)
[music] >> Thank you so much for everyone being here. We're going to get started in a second. Um but before we get started, I am curious, how many of you currently
work in the healthcare industry in some shape or form? Oh, that's amazing to hear. Uh how many of you are clinicians by training? Okay, a couple. How many people in the
room are engineers? Okay, awesome. Um and then how many people have heard of Abridge before? Okay.
Awesome. Uh well, I'm going to let you hear actually from our users to start off on a little about Abridge.
>> [music] >> Full day of 22 patients, out by 4:30 p.m., >> [music] >> notes done. That's nice.
>> When I think about Abridge, I think the thing that comes to mind is it's really a cornerstone of how I practice medicine today. Um there's just no way um I would do a clinic or see a patient
without using >> I can be present throughout my clinical encounters. I don't have to think about um oh wait, did I get that? Do I need to write that down? Because I know Abridge
has my back and has everything ready for me. >> Abridge makes me feel free cuz I can really look at a patient, really listen, and not have to be thinking about what
do I need to put in the computer. >> Full day of 22 >> [music] [music] >> Our marketing team produces really good
videos and so they always light me up. Um but the goal of this talk for me, and I know that we have a lot of engineers in the room, my my goal is to talk about
health care as a domain, at least I felt in the past that there was a lot of stigma around maybe the technical problems weren't as interesting in health care. And it is true in some
ways, there's some parts of health care that might not be as tech forward. A lot of things run on fax machines, for example. Um but I want to give exposure throughout this talk of two things. One,
a bridge's journey from clinical documentation to clinical intelligence and what that looks like. Um and then two, I want to expose you to some of the technical problems we work on um and
that have to be the that are truly frontier AI product problems that have the highest stakes. A little about me, my name is Chaitanya, you can call me Chai.
Um my career has always been in AI companies and startups. I first started as an research engineer at a company called Vicarious, uh which its goal was actually to develop AGI, but they took
very different methods. They wanted methods inspired by neuroscience and probabilistic graphical models. Um and they concretely worked on robotics.
Uh then I started working at this company called Glean cuz I faced this problem in my workplace itself. How like information scattered all over the place, context is everywhere and it's so
key to decision-making. And Glean was building basically the ChatGPT for your workplace. I was there about 6 and 1/2 years, as one of the earliest engineers as we went from 10 people to over 1,100
people and work with some of the largest companies all over the world. Um but that journey was amazing. I love that product. I love that company. I love the people there. Um
but I've actually always really been interested in health care. I remember a decade ago, cover of Nature magazine was AI to detect skin cancer. I was like, "Wow." As someone interested in AI, that
was amazing. At the same time, I went to the hospital for something, and I remember seeing Oh, coming back It was a really minor thing, but I remember coming back and looking at the bill, and
I was like, "I'm not really sure what exactly I paid for." Um and so, there's these known problems in health care, and I'll talk a little about some of them, access to
care and cost. And then, we had these amazing solutions. I was like, "Why don't we bridge these things together?" Um and remember, this is about a decade ago. Um and so, I actually started this
seminar where I invited speakers who were physicians, researchers, entrepreneurs to to talk about the space. And what I learned was while there was really, really cool
technology, very little of it made its way into the clinic at that time. And so, again, my my journey went a different way. But as I as I peeked my head out 10 years 10 years later then,
actually, our technology has gotten better than ever, as everyone knows, and the AI wave has just taken over the whole world, has also influenced health care. And as you as you saw towards the
end of that video, a bridge in the matter of 2 to 3 years got its way into 300 of the largest health systems in the United States, Kaiser, Mayo, Johns Hopkins, Sutter, and so forth. And
maybe maybe you've visited some of these hospital systems. And once you're inside the hospital systems, you realize there's so, so much more you can do. And I'll talk about that journey that we've
had. I specifically work on lead our engineering teams for clinical decision support and our agentic experiences that the technology is now
enabled and how we can bring that to health care. But first, maybe maybe some of the problems that inspire us as a company at a bridge. One, one of the things that
we've noticed or many people economists have noticed over the past few decades is actually in many other industries, you actually see the cost of a good go down. And that's because the
productivity has increased. But in health care, we actually see administrative costs have only gone up over the past a few decades
and productivity hasn't necessarily increased. And a lot of our problems in health care we solve with labor, but even that we cannot actually keep up. So there's this
bit of this like productivity paradox you might have heard of like Baumol's cost disease. And it's part of partially because technology I think hasn't fully touched health care as much
as it's touched other industries to increase that productivity. >> [clears throat] >> A few other problems, you know, hospitals are shutting down and margins
are actually razor thin for many health systems. Of course, some patients have massive medical debt and then we actually and the problem that's another problem that's
very near and dear to our heart is that we hear all the time that doctors are burnt out and they actually often don't recommend it as a profession to to their children.
So what we started as as a company was working on clinical documentation. So the idea here, if you're not familiar with it, is at the end of every patient visit, the the clinician must create a
note. Our typical format is the SOAP note that has a couple different formats like what's the chief complaint of the patient and a few other sections and then what's the assessment plan? What do
we do with this patient? You have to do it right this after every single visit and there's some different variations on this depending on specialty. Typically clinicians end up often doing
it takes like 2 hours a day to write just write these notes and you often do it with what's known as pajama time after work itself. And that's a common source of clinician
burnout spending all this time outside of work and it's not the most fun part of the job. However, these documents are actually extremely high stakes because these clinical notes are often used as a
basis of billing, but also which is of course financial things are high stakes, but also have clinical impact and the reason for this is because these prior these notes are used for for
next clinician or as you switch health systems. They use it they provide context to the clinician of the patient's longitudinal medical record. So, it's actually really high stakes to
get this right. Um we started here because we it's a known a pain point that's existed for many, many years. But finally, the
technology's caught up to do really, really high-quality medical notes that's actually personalized to the clinician. This was an amazing wedge into health care industry, which has typically been
technology reticent because it really it led to they actually care a lot about getting these notes high quality and right. It led to higher doctor satis- satis-
provider satisfaction, and such they could actually see more patients. Um and it can actually help create a higher record um that helps prevent uh as it relates to billing, auditing, and
other reasons. So, we started there. Just this product alone scaled to 300 hospital systems. But I want to show you a little about where we're going next. And um and the
core thesis of the company is that every year in health care is everything is around the conversation. So, we started over here with the conversation to clinical note.
Everything else is downstream of that. Whether you it relates to billing, whether it relates to things like clinical trial matching, or whether it relates to clinical decision support.
It's all about the conversation, that sacred doctor and patient conversation, and we've just built all this administrative machinery around that. But how can we
bring it back to that conversation and actually automate some of that uh administrative machinery? So, to give you a tactical example of what this looks like, um and I'll play
this video. Of where where we're going from here. >> We've been building a solution that allows the physician to interact with Abridge directly by using their voice.
Hey Abridge, is Neaten eligible for any clinical trials? >> He may be eligible for the Abridge HF study. He remains symptomatic despite
maximal therapy, and most screening criteria are already met. But, an updated echocardiogram is needed to confirm his ejection fraction and complete eligibility assessment.
>> All right. Please order that echo for him. >> Done. Confirmatory echo ordered. >> And then, when I'm done for the day, I can just ask Abridge, "Hey, Abridge, can
you prepare my charts for tomorrow?" And Abridge is working for me. >> Pause right there. One of the things that you'll you'll notice is that we are thinking about how to revolutionize the
entire visit for a clinician. From pre-visit, how earlier in the we have suggested discussion topics. Here's things that you can talk about with your patient,
whether they're clinical or more billing related. We have after the visit, we actually create everything for you. The patient visit summary, the actual clinical note,
and we actually penned orders as you might have seen. We're able to use We are able to do this all by reading all this context. We have access to all of the EHR context, so we know
everything about the patient, the prior labs, the prior notes. We have access to the live conversation between the doctor and the patient. That's where the debugging happens in health care, where
you learn about what the patient is facing now. And then we have access to world's medical literature that we can ground and clinical guidelines that we can
ground all of our work in. So, I want to I want to switch now, given the context of where we're going as a as a product, I want to switch into some of
the key technical and engineering problems we face and inspire you on some of the what I think are very much frontier AI challenges. So, if you've ever worked on a gentic
product before, the regardless of vertical, um, three KPIs that tend to matter are quality and then latency and cost. In healthcare, I feel that we're
actually playing on hard mode for all of these three KPIs. This is a high-stakes scenario, especially when you're doing something like clinical decision support. You have
to be right because the downside is extremely high when you're wrong. When I used to work at Glynn, you know, while I loved that product, I could be wrong and it would have been fine. Maybe we
answered a question incorrectly. But in healthcare, if we answer something incorrectly, there's actually consequences and we entirely lose our trust. So, quality needs to be
absolutely high and I'll talk a little about how we keep that bar high. And then latency and cost also really matter for us when you're live in the conversation. You can't, with latency,
you can't act on information too late and you have to act on the also at the right time for it to be useful. And then finally, cost at the scale where we're doing this at.
Um, and and as an interesting aside, it actually relates to our we have a motto inside the company that our goal is to save lives, save time, save money for uh, for the hospital system and for the
healthcare industry as a whole. And it actually, I think it's funny that it really maps to the three KPIs that you care about in an A Gentic product. So, talking a little about quality, how
do we keep that bar high? For us, we really treat evals as the operating system, the life's blood of the of the company. This starts from internal benchmarks and
offline evaluation. Before we develop any product, you know, whether we're talking about clinical trial matching, clinical note, clinical decision support, coding, we start with a robust
set of internal benchmarks. This is pre-deployment and then we test that against, you know, things that we've actually seen in the wild. Then we all have a staged rollout. We
know that we need to make contact with the reality reality. Not everything offline will perfectly represent what happens in practice. And so, we slowly roll it out. Maybe it starts with the
alpha group of clinicians that we trust and under they understand the stakes. We rely to beta, maybe there's AB testing at scale. And then even after it's fully rolled out, you always need continual a
monitor. Again, the stakes are really high and you cannot get away with just being like a prototype that you ship out there and be like, "Yeah, I mean I tested on a few cases and it works."
How we do this is we always have expert calibrated LM judges. So, we have clinicians embedded throughout the entire company. The clinicians are are domain experts, but not all of us are
clinicians. I'm not a clinician. So, how can we, the rest of the company, still move fast is by encoding that clinician judgement into LM judges. You know, I think a really great evaluation system
has a property that it reflects the behaviors that you want in your product. At the end of the day, we are making a product for clinicians and so who best other than our clinicians to actually
create our judges that represent what they want. And those judges, once you have that, create a feedback loop so that anyone, whether you're a clinician or not, can actually
uh hill climb and learn from that. We also have a lot of online signals, whether how you're editing the clinical note and your typical thumbs up, thumbs down, uh and star ratings and other
free-form text. So, this is a general framework we use for our for all our products. I want to deep dive into the product that I work on, which is clinical decision support.
So, to give you an example of a clinical decision support looks like, and specifically, we're building something novel, which is contextual clinical decision support. Maybe a provider asks
a question like, "Hey, does this patient meet the criteria for febrile neutropenia?" Um and so, what we have to do here is actually a lot of context is under-specified in this question. Where
So, the first thing we're going to do is we're actually going to pull from the EHR data, previous previous lab values. Then, using that context, we're going to uh use the
we're also going to use the live conversation and we're going to use clinical guidelines and medical journals to come use as the reasoning sources for combining all this context to actually
answer the provider's question. Now, but I want to focus on evaluation. Again, the stakes are really high here. We we really can't get this wrong. So, how do you tell whether or not an answer
is correct? And sometimes I I And this is a case where the generator and the verifier gap is really small. What I mean by this is in some problems in AI, such as like Sudoku, it's really
really hard to generate a solution to Sudoku, but it's extremely easy to verify once you do have the solution. And that makes that makes it much easier to hill
climb against and build evaluation for. But in a case like this, the generator and verifier gap is really small. If I had a really really good generator verifier, then that would just be my
generator itself. So, how do I create a reference that isn't just a language model itself and ground itself to I have trust? So, what we do is we tackle this by
having many many different signals. We have a clinical quality judge, which I'm going to dive deep into, and then we tackle from we have many signals from a boundary and adversarial judge. We have
a clinical safety judge. And we also have judges that represent product aspects like tone and style, and that matters a lot as well for AI products. So, all of these are different signals
that try to get a piece of this like really really hard to measure problem and guarantee it in the way that we want the product to be. So, diving into the clinical quality
judge. So, again, I said the generator verifier gap is really small here. So, what we need is we actually need human references to tell are we generating the right thing. But you can't just create a
human golden response because there is a lot of variability in the potential responses. So, what we did is we took a lot of real clinical cases. We had independent physicians create a rubric.
So, this rubric said elements of what we wanted in the response. So, it's not here's the exact response because again, there's many infinite possible responses, but a good rubric elements
that what a good response would look like. And then we had a separate physician that actually adjudicated it, brought these two independent rubrics together, created a final rubric, and we
actually had a fourth clinician do QA on these rubrics. Once you have these rubrics, and here here's a sample rubric what it looks like. Here's actually a question and
then there's more context in the case itself. You have a rubric of what are the elements that a response should look like. Now, we can actually have an LM judge that compares our agent's
responses to these rubric elements and does some basic semantic match to tell, "Hey, is our model performing well?" As we continue to hill climb, whether it's our agent architecture, our models, or
search ranking algorithms. I want to talk a little now about cost and latency, two other really really hard problems for us. So, we do this as we said in that intro video, we do this
on the live in the conversation, and we do on the run rate of 100 million medical conversations a year. How do we do this in a way that doesn't really break the bank for us?
So, one of one place that this problem comes up is is actually in generating the clinical note. So, when you're generating the clinical note, there's
many different sections to it. There's a history of present illness, past medical history, and there's the assessment and plan. So, one of the core insights for us is rather than say using a foundation
model to generate all this, is we can actually break decompose this problem into simpler, smaller workflows. Health care is actually many specific workflows. You don't need, you know,
Fable 5 to actually solve all of your clinical notes. We we don't need frontier level intelligence for every problem. So, we actually post train a lot of smaller models for different
problems, such as different actually even to the granularity of different sections in the clinical note. And that lets us use much smaller models because it's a more specific problem and at at
much cheaper cost and latency. And we have this data flywheel that we have this unique data set of a hundred million medical conversations a year. And as far as we know, no one else has
such a large data set. So, our key insight is having a right to win in training models. There problems where the quality is already maxed out. And so, you should
train models then to reduce quality and latency. But there are other problems where the quality isn't maxed out and people say, "Oh, the frontier model would just steamroll you." Our key
insight is we can actually potentially beat the rate of change on the frontier model if we have the right to win by having the right data that they may not have and the focus on a problem that
they may not be focusing on. And that lets us still maximize quality. Another problem that I'll quickly touch on is in-visit orders. So, doctors really
aren't a big fans of pending orders, but often they'll mention orders during the visit itself, medication or non-medication orders. So, what we have is while we're listening in the visit,
as the clinician says order, we actually queue it up in the background and let them actually sign it off in the EHR. But you can imagine if we did this in a very naive way, like every few
seconds are just listening for orders, that would really break the bank. And so, a lot of our tricks are like, how do we find the right events in the
conversation to actually trigger heavier models that will actually do the order matching because you need to match the order, not just is the order said, but does it match and references orders that
are approved by the system and are relevant to the conversation. So, we have a number of different gates that are cheaper and faster that let us trigger actually larger models and hand
off to them for actually doing the end-to-end work. But the last message, this was of course a very quick talk, but the message I want to leave you with is healthcare is
a domain that needs frontier AI and actually puts it to the test at high stakes. In the past, I as an engineer myself was worried about working in healthcare. Does
like does healthcare technology actually work? Well, Elbridge has proven this at scale for for And I as hopefully I gave you a taste of some of the frontier problems that we
work on. So, thank you so much. My name is Chaitanya again. You can and feel free to connect with me at Twitter or LinkedIn. Thank you.
>> [applause] [music]