Notion's Token Town — Sarah Sachs, Notion

summarized

TLDR

Notion's AI lead Sarah Sachs presents a playbook for sustainable AI product development, arguing that avoiding vendor lock-in and managing token costs are critical to building viable software factories. She emphasizes that most companies fail to scale AI because they lack a durable system of record and treat model providers as partners rather than competitors, urging teams to focus on product value, model agnosticism, and cost-per-capability-per-second tradeoffs.

Key points

  • Cost is a structural barrier to entry for AI products, and most companies get stuck at AI as an assistant because they lack a durable system of record.
  • Model providers are ultimately competitors; reselling tokens from a first-party model at a markup is a losing strategy.
  • Companies should not tie themselves to a single model provider; optionality is leverage and the discount is rarely worth the loss of flexibility.
  • Notion's auto model handles 75% of traffic by switching between models based on task complexity, cost, and latency.
  • Open-weight models (e.g., Kimi 2, GLM 2) are now strong enough for many tasks and provide negotiation leverage against frontier labs.
  • Many tasks do not require a GPU; using CPUs for deterministic operations (e.g., CSV to PDF, SQL queries) avoids token bloat.
  • Governance, visibility, and sandboxing are essential for safe multi-agent orchestration, especially as agents become more autonomous.
  • Notion's software factory approach uses a collaborative document as the system of record, orchestrating agents like Claude, Decagon, and Codex to automate tasks and save minutes per task.

Tools mentioned

Techniques

  • Model agnostic playbook (build for multimodal, switch fast, evaluate on entire trajectories)
  • Cost per capability per second analysis
  • Using open-weight models as credible alternatives
  • CPUs over GPUs for deterministic tasks
  • Governance, visibility, and sandboxing for agent safety
  • Multi-agent orchestration with a durable system of record
Transcript (captions)
[music] Okay. Hello. Okay. Before I get started, you guys, this is a huge keynote room. Can everyone like come forward because I'm talking to like four empty rows and disperse people. Do me a favor. I'm spending 30 minutes telling you all of our secrets. I can see you still. Thank you. Thank you. Thank you. Thank you. We're just going to chat. It's a giant room and there's 500 of us. This room is way larger than that. Thank you. Honestly, I knew you guys had it in you. It's really not so hard. Thank you. I also sit in the back. I also work during talks. I get it. I totally get it. I did all day, but not for me. Okay, I'm gonna start, but I'm gonna still point at you if you're in the back like you. Okay, I'm Sarah. I'm lead our engineering teams for AI at Notion. Um, welcome to my talk. Um, it's about token town. How do you go from not go from AI pled to AI poor? Okay. Um, I know that today is all about software factories. We're going to talk about that, but we're going to talk about how to do it sustainably. This is me. This is on my first day at Notion in a very sweaty subway. Um, like I said, I lead our AI teams at Notion. Um, and I negotiate AI contracts for a living. My team jokes I act like Anna Winter. So, this is a nice a nice image of me with AI Anna Winter hair after a press article referred to me that externally. Um, and that's kind of the idea, right? Uh, how do you think about negotiating between different vendors? um making sure that you maintain taste for your company. I don't do it alone. Um this is launch day of one of our recent launches. This is just a subset. Any good engineering manager points out that we have a whole company of people building this. I'm just the one that gets to come talk to you about it. So, we've been building a lot. Um this is um an example of our AI usage um just in 2026. Um and we've been really proud of how we've been able to grow that usage. and I'm going to talk to you about how you can build an AI native product and an AI native company. Um, but this is just to give me some credit that that we're doing it kind of well. Okay, so for those of you that don't know, notion's always been that durable system of record. It's always been the place where you can collaborate with your peers. Um, but today that point of collaboration is a little bit different. It's not just humans. Notion's always been the place for collaboration. And today that collaboration happens between humans and agents. Humans and humans, agents and agents. And we like to think about AI transformations going through this journey. And I'm sure some of you are looking at the slide and wondering where you are. AI as a thought partner is when we all started tinkering. We all started just going to the very first version of chat on Thanksgiving when it came out three years ago, four years ago. And we started saying like how can I send this email to my landlord to say that I shouldn't pay for repainting right then we'd copy paste it enter it into our email eventually we started getting to a place where we could use AI like an assistant AI was able to maybe execute individual tasks that's how notion AI really took off in the beginning um and it was able to save employee time but functionally was limited in its capabilities based on what humans asked it to do AI as teammates is what we were really excited to launch almost a year ago now. Um, but this is true in many products where we can do repetitive work and think about a process and have AI do that process. What I think is really interesting is when AI actually becomes that critical workflow where processes are interfacing with each other and you have entire systems running. How many of you guys feel like you have AI as a system down? Aren't you sad you came up now? Great. None of you. Exactly. We have found that no one has figured out how to do this well. 88% of people can't even get past AI as an assistant. And why is that? We have a thesis at notion. It's because there's too much siloed data and not a sturable system of record for that point of collaboration. And we believe that for your software factory to work, for your company to work, and for your systems to work, you need that durable system of record. And that is notion's mission. So doing that is expensive. Um you see a lot of companies that try and commit themselves to this vision and these are just a series of headlines all within a week of how that's painful. So you can put all of your money into a process to try and make a system and you end up feeling like this, right? You end up using a blowtorrch to light what is actually a large cigar. But you kind of get the idea. Cost is a structural barrier to entry. It makes it hard for you to serve products. It makes it hard for you to build factories. And it is ultimately, I would posit, one of the largest reasons why things do not happen at scale successfully today. And I would argue for anyone working in an applied AI company, it's something for them to be really familiar with to understand the trade-offs that they're making to build durable and exciting and enlightening product for their customers. But that's not really how the market is today, right? I'm not going to name names here, but you guys have search engines. You can figure it out. Exhibit A. A reasoning model gets upgraded. Amazing. The per token pricing is the same. What's not to love, you try it out, it uses three times as many output tokens, right? Exhibit B. A model gets upgraded, but it has an entire new digit, right? Whatever demarcation system that model family likes, it's brand new. It's 40% more than its predecessor, which is being deprecated in the next four months. These are real scenarios that we face at notion. All of you are nodding because these are common pretty much monthly now. But here's the problem. Are you growing 40% in that time period? Are you making 30 3x more revenue? No. So, how do you navigate the system? If you just auto upgrade your model and everything that you're doing, you're you're giving someone a bad deal. Either your customers or your investors, depending on how you charge and where you get your money, neither are good. Fortune 5 million companies have the capability to navigate this. They can hire large consulting teams, have durable teams on their own, and build expertise on how to navigate these trade-offs. Um, most people don't. everyone else has no ability to negotiate with leverage and they're stuck in these scenarios, right? Part of my job as that Anna Winter joke is to think about advocating for the Fortune 5 million, the non-fortune 500 companies that don't have the mass to have leverage and negotiate, but need to think about how. And I'm going to share some of the lessons that I've learned when I have kind of large amounts of traffic behind me that I think scale to those who don't. Um, this is probably less of a secret now than it was when I started giving talks like this u maybe four months ago. Um, your supplier is your competitor. I know very few people who have convinced me that that's not true. Um, you will always be getting a bad deal on tokens with someone who builds them natively, right? Sometimes the cost of goods served is extremely different. you're basically they're serving a first-party product and then you're buying those tokens at a huge search charge and then selling them again at another search charge. Um, that's not really value you can defend. You're getting a really bad deal. And if you tie yourself to one provider, you have no exit. If you build an AI product that you're selling with this structure, you are crossing your fingers and hoping that you are a viable business. I do not encourage that. This is really interesting. Dylan in some analysis posted this. I think it's it says eight hours ago. It wasn't at this point. It was probably a month ago. Um they purchased a subscription plan and they just highlighted right how different what Frontier Labs charge customers for first-party products are versus what they sell. It's a bad deal. Don't play this game or try and let me know how you win. I don't recommend. Think about everyone else. Think about what that structure means and where you have expertise. I don't think that that's winning on the token economics. I think it's about product. It's about building data flywheels and understanding your customers better than anyone else. Understanding when you need capability, when you need low price, when you need latency improvements. I promise you, you don't always need what is usually the slowest but the most capable model out there. and then build compelling UI and orchestration. And I'll show you some examples of that to justify the cost on the bad deal tokens that you do resell. The job is not to train. I mean, some of you might be training the best model and I'd love to serve it and come talk to me afterwards, but most of you are not doing that. Stop trying to win that game and think about the best product that uses many models. Help your customers, help your team. Bet on the frontier, not on the lab. and we'll talk about what it looks like to do that. This cost per capability per second trade-off is actually really intense. Um, Citadel came out with this memo um a while ago, maybe two weeks ago. I loved it. The idea is that for the economy at large, simpler models might be the most cost-effective productivity augmenting pathway. They talk about this bifurcation on frontier versus everyday usage. I really believe that. And for every product, the definition of frontier versus everyday, the definition of saturated capabilities or model capability overhangs depends on your expertise on your product. No one can replace that. And not all traffic is equal. It is a huge miss to send all of these to the latest opus model. Some of these absolutely large scale data analysis when you do it on notion will recommend Opus, right? When you triage an email inbox, if we're charging you to do that on Opus, we're ripping you off and ourselves. Think about where your traffic patterns are. And then think about how Frontier Lab model providers are structured today. I mean, it's functionally an oligopoly, right? And that's fine because they're racing to the top. And I think the top is really hard and really important. This is not to say that products don't have a place for frontier difficult tasks. I want everyone to nod and understand that's not what this talk is about. Understand when you need those tasks and it's not everything. The problem with those tasks are is keep in mind how pricing is incentivized. You can figure out who these players are. Either you are the best model. Everything above what AI can't do today is your market. You can basically price it as high as you kind of want. If you're slightly behind that best model, all you need to be is like a dollar per million tokens cheaper and you have the rest of the market. You know that economic theory about gas stations where the best gas stations are the ones that are right next to each other because they cover east and west the most. Yeah. It's the same with model pricing, which means that price does not correlate with capability growth. So for this complex task, understand what capabilities you need, but be the expert on what complexity is. And keep in mind that who handles complexity changes. Um, oftentimes you'll see applied AI companies really be super outspoken on marketing with a specific lab. That's always kind of a red flag for me when they're not model agnostic because if you look at this graph, it basically shows that they're behind every month, right? the new model and the new model provider of the best Frontier capabilities change and if you hit your ride with one particular provider in exchange for for instance a larger discount um you're doing a disservice to your customers like half of the time right so really think about if that discount is worth not actually having a Frontier product and remember that that optionality is your leverage if you don't have the capability to walk at any point you are stuck. And again, I think that's probably the most expensive decision you'll make regardless of what discount you get or the engineering work to have model interoperability. One option to navigate this is stay model agnostic. Have different models and capabilities in your system so that at any point if pricing seems unfair or untenable, you are not out of business. Notion's auto model does this really well. We have state-of-the-art models available always. Um, but we also have an auto model there at the top that handles about 75% of our traffic. Right? We have the ability to switch between models in our product and we also offer it to our customers so that they have access to these models without vendor lockin. That's part of our AI Switzerland approach. You guys love taking photos of slides. This is the slide. Okay. Model agnostic playbook. This is how you do it. Build for multimodal. It is hard to kill the cache and switch models mid-transcript. I understand that we invest in that technology. It doesn't even have to be per thread. Just think about your harness as model interoperability. Think about the cost per capability per second, not just the tokens. Here's a great example. We posted this review when we um announced our partnership with Parallel as our web search provider. If you were to look at just latency of a single call or just cost, parallel might not be the cheapest. But if you have expertise in entire web search trajectories, you'll see how it differs. The granularity of this eval is what lets us make the best decisions for our customers because we understand all of the trade-offs on entire trajectories, not just single calls. Switch fast and often. I think we talked about that. And give them something back. That expertise on use cases is also very valuable to Frontier Labs. We find that our eval program partnerships actually help us a lot with Frontier Labs and is something that we can exchange instead of extraordinarily large commits and I don't think the discount is ever worth the loss in optionality. That's a perspective you can choose to keep or not. The second option is moderate tasks understanding open weights place there. um openweight models are really strong enough to handle these tasks and the possibility to RL on top of them has also kind of expanded the upmarket growth that they can cover. I view openw weight models as basically lowering the barrier to entry on cost for our customers and they also give you negotiation leverage. So it's kind of a credible alternative that's putting that downward pressure on pricing that if there's an igopoly of two or three providers at the top is unavailable right now otherwise. I think Kimmy 26 was probably the first time that we really saw a model that outperformed 52, GPT 52, GLM 52 now is another 52 bombshell in the villa that also probably does best here. But it's no longer the case where openway models are good for just SFT on small tasks. Really think about without RL if they're capable enough for what you need. Um, and again, don't just think about external benchmarks. Be able to have expertise on your system. What are your tool errors? What's the actual latency that you need? Right? Here's an example of a benchmark that we posted. It's a little bit stale on purpose, right? But you get the idea. Philip at BE 10 this showed this slide once and I've stolen it ever since. Um, thank you. Are you here, buddy? Okay. Well, chat, hi. Um, well, he could come up and say it better, but the idea is that you don't have to be at the top, right? I'm not trying to make a case that open weight is the best model out there. Um, the case being made, however, is that um the gap gets covered eventually. So, if the tasks that you're having today are good enough, then in six months, they're probably covered by open weight. So, be prepared now. And the last thing is CPUs over GPUs. Um, we've we've recently launched something at notion called workers. I don't think that the GPU is necessary for every job. A lot of the jobs that we have are actually serving um discrete pieces of code. Like you don't need an LLM to turn a CSV into a PDF. You don't need an LLM to talk to notion tool calls if we have a CLI. You definitely don't need an LLM to do deterministic SQL queries. This is where people become token poor very quick. And I think the last option here besides openw weight CPUs and optionality is actually governance. Um there's a lot of AI governance. Um one is visibility. Um understanding who's using the data, understanding its maintainability and control. When you have model optionality, you can offer a lot more to your customers. Um here's an example of how that governance works in notion. So final tips again, think about architecture, think about open weight, and build value that transcends tokens. So we're going to depart Token Town. I know I said welcome to Token Town. We're going to spend the next 10 minutes really thinking about what to do next. So I think the challenge of the next six months doesn't have to do with capabilities. I think it has to do with security. Let's start there. There's this concept called the lethal trifecta. Simon Wilson, I think, crafted this. If you have access to private data, exposure to untrusted content, whether it be through ingestion, MCP, email, right? And the ability to ex to communicate externally, and that can include like payloads in a web search. The second you have that system, you're exposing risk. And in fact, the more autonomous your system is, the more unsupervised this risk is. I think that this is what builds valuable product, not just capability. Same with sandboxes and computers. We talked about this, but it really is something that builds better determinism in your product and also better token economics for your customers in multi-agent orchestration. Understanding what agents see and do and what persists. I think persistence of enterprise knowledge is something that's actually really not discussed enough. It's starting to be with some recent launches, you know. Oh, there is audio. So don't have your workflows look like this. And I think this is where most software factories are today, right? It's like actually your entire engineering time just spends time babysitting the factory, right? I mean, I get it. Ours started off like this. Agent orchestration is one of the most difficult tasks of making factories work. So okay, this is me telling Tay to tell people to buy notion AI. And the reason I included this slide is I am going to sell notion for a second. It's my job. Always be closing. Always be selling. Always be hiring. come find me. But I'm going to talk for a second about how notion does this today. We already have the ability to inspect tasks. And you can imagine any task that you look at um in a notion document. You can have Claude actually go ahead and scope out what you need. We've launched this manage agent capability today. So if I go ahead to the top of this task, I can actually ask Claude agent to scope out the task. Right? Ideally, it's working. Um, and you'll see it'll actually populate um an entire spec of what needs to be done. In this example, it's not ready. It's going to ask me a question. Keep in mind, this isn't a markdown file. This is an active document. Um, let's say I don't actually know the question and I go ahead and I ask my team um what to do. Imagine that you can kind of tag in your team into these systems. MJ's our PM. So in this example, she doesn't know. Usually she does, but multi-agent orchestration is important. Maybe Claude Code isn't the best at customer voice, but Decagon is, right? You can ask Decagon agents, we're proud partners of them as well, to collect the right data that you need. Okay, in this example, we think we know enough. We're going to go ahead and actually um iterate through some of this flow. I'm going to skip ahead a little bit. We asked our TL what we needed. He replied again. It's a collaborative file, not just a markdown. And we can have Claude actually go ahead and spin up the PR. Hopefully, this is looking a little familiar now. This is kind of the vision of software factories. It's what we're trying to host. Okay, Claude put up a PR. Maybe that's not enough. Um maybe I want to go ahead and ask Codeex what it thinks. Great. Found two issues. You can think about this scaling in an actual factory. So today in notion, you're actually able to orchestrate these agents together and you're not committing to a lab. You're committing to the concept that AI is augmenting and automating what you do. This is real. I asked Rejieve if I could post this. This is how it works today internally at notion. Almost all of our polish and large feedback like this is actually coordinated um through our software factories both in terms of writing to the right teams and also having coding agents take the first step. Verscell does this as well from staging to shipping to closing. And we see massive ROI gains from our customers. That's over three minutes saved on a given task. Imagine that at scale. So I think we're trying our hardest to think about the factory lens. We cannot do this without optionality and we cannot do this without conviction that we understand what models are required for which tasks. It's really wild out there you guys. I get it. The market is really young. It's exceptionally opaque. It's moving fast. I'm super grateful for communities like AI engineer to bring us together and like talk openly about these things and how we navigate it. Um I think we owe it to all of our customers to get it right and to be critical thinkers about how we navigate this together. I'm chronically online fortunately. Um you can always DM me on Twitter, you can email me, you can find me after this. Um but thank you for yapping with me and thinking about this problem and have a good day. >> [music] [music] [music]

Frontier News · by Hyperjump Technology