How to build an AI-Native Health Company — Dan Feng, Maven Clinic

summarized

TLDR

Maven Clinic's transition to an AI-native company required rethinking development processes: they abandoned quarterly planning for 2-4 week sprints, allowed engineers to self-merge simple PRs, and limited PRs to 500 lines to avoid rubber-stamping. The company also changed hiring and performance reviews to reward AI leverage, and implemented reliability measures like multi-model consensus for critical tasks. The talk provides a practical playbook for organizations moving from traditional to AI-native operations.

Key points

  • Maven Clinic started its AI journey two years ago and built Maven Intelligence, an orchestration layer across all products to enable AI for the company and clients.
  • The company defines becoming AI-native through three parts: internal AI tools, product AI integration, and cultural/process changes.
  • For internal AI adoption, Maven focuses on the majority of users by building shared infrastructure and easy-to-use tools, while listening to feedback and supporting multiple tools (e.g., Cursor and Cloud Codes).
  • Hiring now prioritizes candidates genuinely interested in AI, who can handle ambiguous problems, and who blur the boundaries between PM and engineer roles.
  • Performance reviews now ask 'what have you done for AI side?' and reward employees who leverage AI to multiply their impact.
  • Development process shifted from quarterly planning to 2-4 week sprints, with short 1-2 page documents for communication; mid-term (3-6 month) planning is avoided because AI capabilities change too fast.
  • Code review practices changed: engineers can self-merge simple PRs, PRs are limited to 500 lines, stacked PRs are used, and rubber-stamping is discouraged; AI code review tools are still being developed.
  • For reliability, Maven identifies acceptable vs unacceptable failures, uses multi-model consensus for critical tasks (e.g., reimbursement claims), runs integration tests multiple times with a 90% pass rate, and performs auto-evaluation with rubrics plus human spot-checking (20% for new features).

Tools mentioned

Techniques

  • multi-model consensus for reliability
  • integration tests run multiple times with high pass rate
  • auto-evaluation with rubrics
  • stacked PRs
  • self-merge for simple PRs
  • 2-4 week sprints
  • short PRDs/TDDs (1-2 pages)
Transcript (captions)

0:01 [music] >> It's time. We can get it started. I'm Dan. I'm from Maven Clinic. Today we'll share the experience how we transition from a traditional technology company to

0:23 AI native company. Before I started, I would like to do a little exercise. Raise your hand if you think you are already an AI native company. Okay, we saw a few. Raise your hand if

0:35 you thought about it, but haven't started the journey yet. Okay, we saw a few. That means most of us is in between. Hopefully this talk can help you with that one.

0:47 So, Maven Clinic is the largest digital health platform. We are focused on women and their families. So, we specialize in like maternity, fertility, parenting, menopause. We started our AI journey

1:03 just 2 years back. And this moment we built something called a Maven Intelligence. It's a orchestration layer across all our product to enable AI for everybody in this company and for our

1:14 clients. So, AI is here and improving every day. I think adopting it is not an optional. Even you choose not to, your competitors

1:25 will do. This is a quote I heard like a couple years back. I would like to share here again. Like a tractors aren't to replace farmers, but the farmers who can operate the tractor will replace the

1:37 ones who cannot. Hopefully everybody here will become farmers who can operate your tractors. That's the goal. So, first of all, I don't think there's a one single definition what it means by

1:51 AI native. And more importantly, there's no predefined playbook you can just follow and bingo, you become AI native. For us, it is really come down to three parts. One is internally with our AI

2:05 tools, whenever it's possible. It can be as simple as like generating your daily summary, managing your meeting, create a Jira task, anything you need to do today manually, you should think about to say

2:18 can use AI to do it. Whenever you want to ask other people to do something for you, you should be saying with yourself I can use AI to do it. A lot of leaders in

2:27 fact today at the Maven include our sales sale, they use AI tool to solve those task by themselves now instead delegate to other people. That's for internally. Externally, we want to build

2:39 AIs into our product which achieve two goals there. One is really focused on improve our user experience. Second maybe help us reduce our operational cost. Like a like AI

2:52 based like chatbot is really good example. It's 24/7, always available, can help our address issues, help our customer instantly. It's way better and cheaper compared to human agents.

3:07 Thirdly, I think is more important, we need to think about our culture, process, the way we work, how we can change it so we can maximize what AI offers for us.

3:18 I will touch it more on the following slides. So, when we come to adopting new technologies, there's always like screw three groups of users. One there's some

3:30 early adopters, right? For them, we don't need to do too much. Only thing we need to do is enable the tools for them and encourage them to share what they learn with the company. And what we need

3:44 to really focus on is the one in the middle. That's a majority. We should build a shared AI infrastructure for them, build easy to use tools for them. Just make the adoption as seamlessly as

3:56 possible. More importantly, we should really listen to them, get feedback, consistently inputting. For example, last year, most of folks and Maven they are using cursor. This year, a lot of

4:09 them switch to cloud codes. For us, we need to support the both. We need to meet where they are. Just make feel they comfortable to use it. And for an older places, you always have a few slow

4:21 adopters. They always have concerns, worries for the new technologies. For them, we just should have meet where they are. Understand what's their concern is. But

4:31 more important, we should be crystal clear with them where the company is heading to. So, AI is really good and execution if we know what we want to do. So, this is

4:46 change how we should hire new people and how should we reward them. We used to the way we used to work is we have a senior engineer who sense the problem, come up with solution, and

4:58 delegate to other engineers for implementation. So, we can work on it in parallel and be faster. But these days, we found the NASA and NASA engineers wouldn't like to delegate the

5:10 implementation work to other people because they they already figure out how to solve the problem. They just use AI to solve it instantly. Delegating to other people means more overheads and

5:23 less efficient. And also means like when you have a new people, you want to make sure they can solve the problem independently. They pretty much has to work on the traditional technical lead

5:34 memo. We can We cannot afford other people delegate implementation task for them. Also, when we hire a new people, we should think about what we are looking for. We definitely want to look

5:45 for somebody genuinely interested in AI. The domain is moving so fast. We want them to keep learning. Also, help the team to stay on track. Secondly, is um with AI, engineers can do way more than

6:00 they used to do. The boundaries between PM and engineers is getting blur blurry. We found like a engineers who really understand the product. In fact, they can have more way more contribution than

6:13 a traditional engineer who only focus on software side. And this is what we are looking for. And those deep understanding of the system, the ability to can handle complicated ambiguous

6:25 problem is also very valuable. This is where AI land off. When we hire new people, this is also the people we are interested in bring on board. For people

6:36 we bring in, we want to reward them in the proper way. Even in our performance with review, we start ask, "Okay, what do you have done for AI side?" We definitely want to reward the people who

6:47 leverage AI to multiple their impact. Although this is impactful everybody in the company. So, now we have the right tools. We get the right talent in the place.

7:00 And we need to change how we work to maximize the benefit of AI. The we used the way we used to work is say, "Okay, we spend the weeks, sometimes it's the months to flash out the business

7:14 requirements, finalize the design, and then do the implementation." Because implementation can be really expensive. If we didn't get the other part right in

7:25 the beginning, it's it can be very costly to change it later. But in fact, we never get the sense and right in the beginning anyway for any of

7:35 big projects. With AI, building is super fast. It's probably couple minutes you can get it done. Argument is really expensive one. So, we should really think about what's the best we can work,

7:49 how can we deliver fast. It's still okay, you can think about what you want to deliver in one year. You can assume AI models can do anything you want in one year. Based on that one, really

8:02 dream big to think what you can do in one year, but it should only serve as inspiration, inspiring, and directional. What we really need to focus on is what

8:13 we want to deliver in the next two to four weeks, right? What we want to get to the PMs and designers is say, "Okay, tell me what I need to deliver in this sprint." And the engineer will focus on

8:25 it and get it released if at the end of the sprint, if not sooner. Meanwhile, and the PMs, they have time to flesh out the next bunch of the requirements. If at the end of this sprint, they say,

8:37 "No, what we decided two weeks ago is wrong." It's totally okay. We can switch the gear, get it fixed quickly. That also means like we prefer people not write pages or pages of PRD or TDD

8:51 anymore. We prefer them to write just a short one or two pages. That one is really serve as communication, so we can iterate on it. The really awkward part is mid-term

9:04 goals. Those like a three months, six months. It's very hard to plan these days. The reason is I don't know what AI models will be capable in three months. There may be multiple releases already.

9:16 So, we prefer not focus on this one. But this one can maybe make it not very easy for most of folks who has been in this domain for a long time because traditionally we get used to have a

9:29 quarterly planning or we plan it for six months. But it's our job to get used to the new AI era and learn how to work it efficiently. So, I want to talk about the coding and

9:44 software development a little bit more here. AI coding tools is probably the most successful AI application. And it's really good and implementation. So,

9:57 you probably heard a lot of people say, "Okay, I have this AI tools. Now I can even use my phone to implement software and automate every stage." If they feel comfortable do that, it's totally okay.

10:09 But you don't have to. What I'm trying to say here is And maybe this is our journey, how we adopt those AI tools. We started with the lowest risk task, like starting with

10:20 writing unit tests, documentation. Those things are very easy to verify and the risk is super low. By doing that, one we build the confidence and we start to construct our own rules, skills, and

10:34 build our barriers. And then we push to the whole engineering team say, "Now you should use this AI coding tools for all the tasks." When they choose not to do, it's the time we really want to learn

10:45 say, "Why you don't do it?" And at this moment, we pretty much use the AI coding tools to do all our implementation. Engineers really focus on reviewing, architecturing, and evaluation.

10:59 >> [snorts] >> So, and with AI coding tools, we are writing so much code these days. Code review becomes really challenging. So, for good engineer, used to they probably

11:11 write hundreds of lines code every day. These days, they can easily write like thousands. If we keep do the code review as we used to do, we won't be able to keep up. We also tried

11:23 the multiple like AI coding review review tools. It helps a little bit, but we don't feel comfortable 100% rely on them yet. We still found those feedbacks from our engineers are very, very

11:35 valuable. And that means we need to really change the way we are doing code review to meet where we are now. And couple things we have done. One is we allow engineers to self-identify whether

11:49 they still need code review. If they think this PR is simple enough, I feel very confident, I don't need anybody to take a look. And we are fine with that one. We let them merge, but we still

12:00 hold them accountable. And if they do want code review, we want them stay with the best practices. For example, each PR shouldn't have more than 500 lines of code because nobody

12:11 can do a meaningful code review with the ones has like thousands of lines code. And we also enable the like stack the PR. What it means is that for big

12:22 feature, and the engineers can bring break it into multiple PRs where people review the PRs and they can keep working on it. One thing we really want to avoid is a rubber stamp, we call it. Means

12:36 like people submit code review, you cannot really do anything to it. You just say blindly approve it. This is the worst case, we should really avoid because that's just give us false

12:46 confidence. We think we reviewed it, it's good, and we release it. Meanwhile, we should keep working on our AI code review tools because we are thinking that's the future.

12:58 So, at this moment, we use AI tools pretty much assist each step of our software development. Our goal is it will be automate the whole life cycle from end to end, from designing,

13:12 implementation, and here is a fully release it. More important, we want the AI tools be able to monitor the live traffic and be able to catch the issue early and automatically fix it. That's

13:25 what we are still working on, and we're not there yet. The last thing I want to touch a little bit for this presentation is about reliability. So, what it means is like

13:37 for the traditional software, it does what we implement there, no more, no less. But for the Genex solutions, hallucination is there. We cannot ignore

13:49 it. And there completely eliminating them is can be very costly. Sometimes is not necessary, either. So, the way we should do is really have a holistic solution even

14:02 from the get beginning. For example, we can start with identify which failures is acceptable, which ones are not acceptable. For our AI system, for example, we have

14:14 the functionality to help our customers to schedule appointment. If we fail well to 1,000, probably it's okay. I'm not saying it's a good experience, but the users really can just click the button

14:26 again, we will reschedule for them. Probably it's okay. But if we help user to like submit their reimbursement claim, we cannot tolerate a failure because if people ask of $200, we issue

14:39 them 50 or they ask 50, we give them 200. Each case will cause a escalation right away. For those cases, we have to put in extra stamps. For example, when we receive their receipt, we will use

14:53 different models to review the same receipt. We only move forward if the results from different models agree with each other. If we really have trouble to figure it out which one is right, it's

15:05 it's easy it's okay to tell the customer, say, "Hey, we have trouble to process your stuff. Do you want us to get you connect to a human agent?" We will move from there. That's acceptable

15:15 solutions. And also we have should have a rigorous process to release our software. For us, we have like hundreds of integration tests, for which pretty much covered all

15:27 the use cases we know, and we are keep adding to the integration test the suite. And the when we run the integration test, not only pass once is not good enough anymore, because the LLM

15:39 can do different things. So, for each test case, we run it to many times. We consistently requires the high pass rate, like for example, 90% for all the time.

15:50 And the more important, [snorts] and the after we launch the software, we have our auto evolve system evaluate carefully evaluate each conversation. We have predefined a lot of rubrics, what

16:03 we think is good, what is bad. And then we will generate results, we will review the score. Besides this one, we also have dedicated a group, their job is mainly review

16:16 those conversations. We will spot a check our conversations, that helps us to say whether we need to come back to improve our systems, or our rubrics is too strict or too loose, and we need

16:29 consistently improve it. When we launch new features, then the time we say not only spot a check probably not enough, we really want to review like say 20%, and we can do it. This whole process

16:42 make [clears throat] sure we we are really confident whenever we release something, although we know hallucination is there. And that's pretty much what I have for

16:51 today, and I can stay here to take up questions, and if you have other things, you can reach out to me. >> [applause]

Frontier News · by Hyperjump Technology