The Future of Claude Code: Mods, Mutable Software, & Multiplayer Agents — Thariq Shihipar, Anthropic

summarized

TLDR

Claude Mods lets users rewrite the Claude Code harness from inside the product, with plugins that hook into the agent loop, fork prompt-cached sub-agents, and customize the UI. Shihipar presents it as a preview of 'mutable software' and expects the architecture to split into separate inference, execution, and interface layers. The practical shift is that customizing agent tooling is becoming a prompting task rather than a build-systems task.

Key points

Claude Mods, which leaked the week of the interview, customizes both execution and UI of the Claude Code harness.

Shihipar predicts frontier models will become Pareto dominant over smaller tiers because verification costs shrink as models get smarter.

Shihipar expects CLAUDE.md to disappear and recommends starting new projects without one.

Anthropic added eval plugins for skills so users can test whether a skill change is an improvement.

Shihipar recounts an OpenAI eval where agents used Artifactory cache folders to communicate and hacked Hugging Face for scorer code.

Tools mentioned

Techniques

  • Forked sub-agents that reuse the prompt cache for cheap side computations
  • Inference-time probes (constitutional classifiers) that monitor model activations
  • Auto mode permission classifier that checks agent actions against user intent
  • Model fallbacks where classifiers route work from one model tier to another
  • Effort levels that scale verification and edge case testing spend
  • SCQA (Situation, Complication, Question, Answer) prompting structure
  • Implementation notes and decision logs as persistent context
  • Dashboard artifacts with a database and artifact MCP for multi-agent state
Transcript (captions)

0:00 I think it's just like different models are very different from each other, you know what I mean? But I realize that it's like such a pain to like maintain different ones, you know what I mean?

0:09 And yeah, like as the models get better and better, the floor of how they accomplish the simpler task is better. And so I do think in the limit cloud.mmd goes away. And maybe not even like that

0:20 far. Like I I think like I think that right now it might be better to start a new project without a cloud. >> Yes. Um, I think that like maybe if you see very repeated failure modes, you add

0:34 them to your cloud. MV. The really tough thing is that this changes per model. And so like if you've added a bunch of failure modes or like like >> you need Fable MD, you need Opus MD.

0:43 >> Uh, well even Fable 5.1 versus Fable 5, you know, like it is annoying. Like I'm not like, you know, like we don't like do this on purpose, you know what I mean? It's just like how the models

0:53 work, right? And so like maybe like Fable 5 had this like failure mode that Fable 5.1 doesn't. And if you keep this context running log of a bunch of different failure modes, they will

1:04 probably over constrain Claude, you know. And so this is like uh we we just actually added eval plugins for skills. And so now you can eval if a skill is better.

1:17 >> Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so

1:26 clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis. But fortunately, enough of you actually subscribe to us

1:35 to keep all this sustainable without ads, and we want to keep it that way. But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to

1:46 click that subscribe button. It's the only thing I'll ever ask of you, and it means absolutely everything to me and my team that works so hard to bring the Inspace to you each and every week. If

1:56 you do it, I promise you, we'll never stop working to make the show even better. Now, let's get into it. We're here in the studio with our friend Tar from Enthropic. uh and I guess

2:10 generally the cloud code I there's there's so much uh sort of merging of boundaries and you've been so on top of everything since you joined Anthropic you have been early to cloud code itself

2:23 but then also and you've told that story in other podcasts and you've also been talking about seeing like an agent most recently you did the top AIE ware talk field guide to fable which obviously you

2:35 guys launched fable so that that's cheating and mostly you most recently also launching cloud tag and we're also going to be talking about pacing on frontier.

2:43 There's a lot going on in enthropic. I guess top of the question is what's it like being anthropic when there's so much going on. >> I think that it is like I think you can

2:55 get whiplash sometimes. I I think like going when I joined Enthropic, I joined because of cloud code like like cloud code had just come out and I was like this is so good and opus 4 to me was

3:06 like like just I could not imagine like how good it was. You know what I mean? And that was like a real moment for me. But I was like trying to convince like my startup friends basically to to use a

3:18 coding and they're like oh no like our engineers don't think it's good enough or something. And I was like that's insane. Um, and now you like fast forward, you know, 12 months, um, less

3:29 and like it's just like, yeah, the default way that everyone codes, right? And I think that like just having to go from like selling it to like now, you know, teaching people how to be, you

3:38 know, make the most use of it and and be more efficient and things like that is just like a like a big, you know, like big change. And uh yeah, I think like it's just hard to stay on top of

3:50 everything as a human, you know, like things happen so so fast and like >> more agents at it. >> I mean, yeah, like that's actually like the agentic stuff scales much better

3:59 than the like human stuff where it's like oh like there are three things happening right now and like they're all emergencies and like how do you like you know uh respond to it? Yeah.

4:07 >> What do you split your time on? You do a lot of technical writing, engineering work. >> Yeah. So I I think that like when I joined the cloud code team, I wanted to

4:16 teach people how to use cloud code basically. And I think that like that has been something that like I thought like maybe I would spend a little bit of time on it or like I you know like do I

4:26 was spending some time on the agent SDK first and I I wasn't exactly sure like you know how the bitter lesson would go you know when it comes to like harnesses right like I think sometimes we were

4:36 like oh like what's after cloud code you know what I mean and so initially I was like I just want to teach people how to use cloud code and make it easier to use cloud code and I think that has just

4:44 like as harnesses have gotten better and better that's like the dominant problem now is like how do you use the agents right like it's like such a high skill expression thing so I I do that and then

4:54 I do engineering work um I give talks but I think like when I'm doing engineering work my goal is to take that feedback that we get from users and also like then be able to talk about like hey

5:05 how how to use cloud code to do engineering so there's kind of like a good loop there yeah >> yeah I'll for listeners we'll attach uh the talk that you did with Sarah for the

5:13 dev writers uh meetup which we talked a little bit thought well first you do the work and then you talk about the work something like that sew and reap or what was that

5:21 >> and reap and sew and reap >> something like that something like that yeah so and then just to preview a little bit uh we are going to talk about the evolution of the harness it has come

5:30 a long way from just being a CLI we're going to talk about cloud mods which is starting to leak today uh because you couldn't keep it secret >> yeah yeah yeah basically yeah yeah

5:41 >> yeah there's there's a lot uh there I think you started off with like adding ask you the question tool which people would love and hate actually like I actually thought it was like very

5:52 innovative and then now I have like my own version you have your your your interview me version. >> Yeah. Yeah. >> And um yeah everyone just has like their

6:01 own their own stuff and like it no longer matters because now you're supposed to uh write prompts that create other prompts and loops and all these things.

6:07 >> Sure. >> So what's the state of the art uh today? Like what what are people uh what what are you like telling people to do today? >> Yeah. Ask user question was the first

6:16 time that the model was good at elicitation. You know I think this was like kind of an emergent behavior that I like you know wanted to see if the models could do. Um I have like kind of

6:27 like a human computer interaction background. So I like you know did that in in undergrad and grad school. And so this was like I think it's kind of like human agent interaction to me like you

6:36 know trying to figure out like how can the agent communicate with you and extract you know the the requirements right. Uh I think that like one of the things about like that's difficult as

6:47 cloud code has gone broader and broader is that everyone has like their own way of using it and it's actually very hard to like change the default behavior. So for example like if someone asks you

6:59 cloud code to do something sometimes they just want them to do the work uh because they're like maybe a very good prompter and sometimes they want like actually are not good at prompting you

7:08 know what I mean and you need like the agent needs to like clarify you know and so that's like a good split and and ask you the question tool sort of like splits along that side where like are do

7:18 you feel like you're good enough to instruct the agent as it is or is the agent able to like does the agent need to like pull out more requirements and like collaborate with you more and

7:27 really understand your preferences. I on the whole believe that pretty much everyone is more on the latter than the former that they like have more ambiguity and they know less than they

7:38 want than they like um think they know about the problem. Uh but like we're it's like interface design problem to make that easy. You know what I mean? And so like if you're designing a

7:50 problem or if you're going through a problem like you know things like what's the schema or like what's the call stack and things like that are really important. Um you know like the details

7:59 in the design are important. Ideally you want to figure out some of these like hard problems ahead of time before starting implementation. And yeah that's why they call like unknowns right and so

8:09 I think that this will forever be like a skill in agentic coding is like figuring out your unknowns. like um because even if the model is like super intelligent, it like needs to know what you want, you

8:20 know, and like uh you have preferences uh like you need to sort of like pull that pull that out. Uh and so that's like I think how I'm what I'm pushing. Um the question then is like how does

8:32 the agent interact with you? And I think that has been HTML has been like the big way of doing that. And we've recently added artifacts, right? And artifacts I actually think we've done a bad job of

8:42 like or like I've done a bad job of like explaining how to use them fully. We have a lot of property capabilities. They have a database associated with them, you know, and so every artifact

8:53 can store and write persistent data. They can like feed back into cloud, you know, and so like one thing that you know like people are not doing yet that I'm trying to like encourage is like

9:04 this idea of a dashboard artifact. So you have like Claude working on a project long term. Maybe it's like a canban or something. Um, it can store that canben data in a database. Multiple

9:15 clouds can access that data via like the artifact MCP and like that that artifact can like talk to those clouds as well. And so like the we're basically building the primitives for you to be able to

9:29 have this like generative interface via artifacts that will like let you surface more of that rich detail from the agents. And I think that like almost everything with agents right now is like

9:41 this problem of like you think you know what you want but you don't really know what you want. And like the agents need a lot of detail. Um and collaborating with them in the loop is really

9:52 important. And so artifacts are like the like way that we're trying to evolve there. But there's a lot of work to do cuz it's so much more complicated than like a multiplechoice question, you

10:02 know? um there's a lot more like detail in terms of like diagrams and code snippets and schemas or like like whatever it is for that problem but like artifacts is like the more AGI pilled

10:14 way of like you know doing ask is a question basically so yeah >> I think one thing that's unclear to me about this uh the the artifact stuff is like what feedback should go in through

10:23 the artifact and what feedback should go through a cloud chat because the more AGI pled one is to just feed everything to the cloud >> I think the more AGI pled one to go

10:32 through the artifact like and I I think that like we sort of imagine in the limit I think that artifacts will be your interface into the harness you know you can like comment on this like live

10:44 like document of your plan of the work uh you can see maybe like multiple agents and different agents are doing this and that artifact is built for the current work that you're doing right and

10:53 so like each one has like slightly different um I think we're still like getting there from like an infrastructure perspective but yeah I think like on the-fly interface for your

11:02 harness is probably where things are are headed. Is there a version of it that's an abstraction from CLI or chat and you because right now a lot of it is okay you're interfacing with cloud code

11:15 you're having HTML given back for a mockup it's pretty rich there's diagrams artifacts are ways to connect these together why not just do everything that way then it becomes like separating out

11:26 like where's the inference happening where's the intelligence happening where is the work happening you know like I think this is kind of like difference between like or like some of the

11:34 distinction between local and cloud, right? And so, um, I think right now if you use cloud code, it's like local and, uh, like you can spin off remote control for example to get some cloud behavior

11:45 or you can spin off cloud code in the cloud, right? We're moving towards a place where instead of clouds uh, like you message a local cloud, it starts a session locally and it executes to more

11:58 like you have a cloud that you message that's in the cloud that's running. uh it can run like local uh or like cloud sessions. This is kind of how cloud tag works, but like over time we'll add like

12:09 local hands as well. And and so like local hands will be the ability for that agent to access your computer if it's online, you know, uh and be able to like work there. And so it it can spin off

12:20 many different sub aents. It can like those sub aents can communicate with each other. And that's where the artifact comes in to display all of that work basically. So you can imagine like

12:29 the you're separating out these things. So there's like the surface UI display that's an artifact and hosted somewhere and has a database and everything. There is the inference intelligence, right?

12:40 That's happening on the cloud and you don't have to worry about shutting off your computer or whatever, right? Um and then there's the like hands kind of like and it can be local, it can be in like a

12:50 remote sandbox or wherever you need your work to be done. That's like unpackaging like the cloud code experience. Right now we're like by right right now it all happens in one place right so

13:02 >> how do you see like the multiplayer side of that so say teams want to work in this way right now it's very individual but how do you see the future of multiplayer like right now I guess

13:12 there's cloud tag which is a version but we're launching projects and so projects is the like this abstraction that's kind of like cloud tag but on our cloud products right so you can message it and

13:26 like it will do the cloud tag like stuff like spinning off sub agents. So we think with multiplayer like claw tag is like a little bit more native multiplayer because it's just like in

13:38 your slack and the permissions are all figured out and stuff like that. But I do think multiplayer is like an important part of the story and like uh that that will need to get tied together

13:48 more. Like you can imagine how complicated it gets when you're like, "Oh, you have hands, but now you have other hands in other people's computers too and like you need to like permission

13:56 them or like you have like your MCP and someone else's MCP and how do you figure out how to use them?" Right? It gets like quite complicated and claw tag does a good job of like sanding down all of

14:06 these issues, right? So that like when you have Yeah. Google Docs, how does it access Google Docs, right? like it accesses through the shared cloud MCP or it can access through your local

14:18 credentials as well if it doesn't have access. But yeah, like I think cloud tag is our multiplayer um product and it's really useful for these like things that are inherently multiplayer like okay

14:28 like on call for example incidents are inherently multiplayer. You want to tag cloud, you want multiple people to log in, you want to be able to find context. Um, I think whenever I'm like working on

14:37 something and I want like privacy or security or like I want other people to review it, you know, it's really nice to like I'll have a channel per project and I'll like at legal for example be like,

14:48 "Hey, like I want to ship this. Can you like like here's the cloud knows everything, you know, just chat with it." And that way legal gets precise answers, you know, on like what exactly

14:59 is shipping into the code and I don't need to be in the loop, right? So I think like multiplayer is getting like more and more like um yeah everyone can participate with claude. Um I think

15:09 claude tag is like that that product and like projects we'll start off single player and we'll like you know expand. >> I think there's a question about like maybe dual questions about identity and

15:21 the unit of isolation. >> Yeah. >> Um tag you specifically chose to make it its own identity which is like uh

15:30 controversial choice. There's there's other ways to do it. >> Yeah, >> cloud projects probably it sounds like you know if it's anything like CHP

15:36 projects uh it is uh you know the isolation is that artifact that cloud instance everyone's collaborating on this it it'll it sounds like you know it should it should be like if you're if

15:47 you're collaborating illegal on on on the thing like that channel should be a project right like it's not yet but it that's the natural next step. Yeah, I mean like I think in cloud tag it's

15:57 effectively like like cloud tag you have to sort of do your own arrangement basically and so cloud tag yeah each channel is like you can name it as you want and I name like each

16:08 >> feature basically as a channel >> but I think like there is some trans like it's unclear when there is transference cuz let's say it if you have a coworker

16:16 >> who is tagged on all these things yes there is transfer because it's the same person uh but with claude it's unclear if it's like necessarily like well no you don't know any you don't know about

16:26 the other stuff you should only use this stuff. >> It's like the tip of the iceberg meme, right? where you can like this is what we spend so much time on basically is

16:34 like there is like infinite surface area of like okay you want clouds to not infinite but like there's like surface a lot of like uh surface area to figure out of like permissions and visibility

16:46 and like uh you know like how can you have let claude operate as well as you can as safely as you can you know and obviously this is very important to us because like you know like like security

16:56 for our codebase is very very important and so we've put a lot of time into this. Yeah, there's so many like edge cases you can figure out where it's like, oh, like, yeah, this cloud in this

17:05 channel has different permissions, but it can message another channel and can't it exfiltrate data that way or like can you like what if it uses your MCP and then messages someone else? Like there's

17:14 like so much and we've like really put a lot of work into sanding it down. >> Yeah. Yeah. Lots of work. Um, okay. Fable. >> Fable, you wrote two good articles. I

17:23 mean, you've written many good articles, but on uh, you know, field guide to Fable building cloud code. I'm curious from what you've seen is there any common patterns that you see in like top

17:33 users adanthropic externally like what are best practices for getting the most out of cloud code the like meta skill I say is like prompting is like very important you know and like that like I

17:47 I think this is like not trivial to say because I I think a lot of people are like oh prompting doesn't matter it's just like I can just say a sentence and cloud will do it and I think prompting

17:56 is really this like this it's like public speaking you know like or writing or something and for a specific audience and that audience is Claude and you need to like build a mental model of Claude

18:07 and how it thinks and how it works right and so that's like the most important skill in working with cloud code is like having this mental model right of Claude and like what it can do well what it can

18:18 oneshot what it can't and so so many people when you see prompting they're just like they're short prompts but they have such a good mental model of Claude and of like the codebase and things like

18:28 that that like it's effortless, you know what I mean? But it's like high skill ceiling. So like that work of like you know spending a lot of time prompting and building mental models of how you

18:38 know an intuition for how the agents work is really important. And then I think like the next thing is like the unknown stuff we talked about earlier where it's like being able to find out

18:50 like your what you don't know or what you haven't written down. Um learning about like different things. I think as cloud can do more and more things, the likelihood of you doing something out of

19:01 distribution for you and like you have low domain knowledge on is very very high, you know, and the more you can like learn the vocabulary to be able to prompt Claude, it becomes really

19:13 important. And so like I think the most important unknowns are the unknown unknowns where you're like I I just like don't even know that this exists, right? Yeah. Exactly. I think this like

19:21 illustration of like the map and the territory, right? where you're like, "Okay, this is my prompt." And the territory is like the actual like work that the agent needs to do, right? And

19:30 if you are like very precise, you can give more precise things, right? So like for example, in design, I'm not very precise. I'm not a designer. So I say like, you know, give me like eight

19:39 different mockups. But if I was a designer, maybe I'd be like, "Oh, hey, here's some reference sites like I want this type of font and this type of like look to it and here's like a few

19:48 different components to to like visualize. here's a Figma MC board to bring in you know like and so you can just be so much more precise with that language and if you're not a designer

19:59 you just need to like try and learn the language basically or learn the unknown unknowns and this is true of like everything I think like the more like you can work with claude to learn like

20:11 how things work the better your prompting will be I think another good example of this is like game design you know like where a lot of people are like oh like I can vibe code a game now and

20:20 they're like it's not fun and and like it's just like the thing about game design is like every one of these choices has like a lot of >> variations

20:29 a lot of like craft to them. So it's like okay like when you're making a a flying game the feel of the p plane and the like you know way it responds to your controls has a lot of like like you

20:41 know a game designer would spend like days on that you know what I mean um and >> to me that's what taste is right like it is like from the possible space of 1000 mathematically valid answers here's the

20:52 one that is the humans will like >> yes I think with taste I'm like torn on this word cuz I think you're right but everyone has different definitions and it it shines sounds kind of like low

21:03 skill or like elitist almost where you're like oh like there are certain people with taste >> like taste is what I call taste these guys don't have taste

21:12 >> yeah exactly oh like an engineer doesn't have taste like I the like founder have taste you know what I mean and I think that's actually not true like I I think like the engineers have a lot of taste

21:20 for these particular like problems you know and I think everyone has taste for particular problems I think like Jason Lou like like say in order to uh have taste, you have to eat, you know. Um,

21:34 and I really like that where it's like, okay, you have to like do a lot of things. You have to like iterate and figure out what you want, what you like, and you know, like build that like

21:43 domain domain vocabulary. And then when you're prompting, you're like synthesizing all of that for >> Isn't it annoying when someone else says it better than you? Like, I have

21:52 to quote this guy forever. >> Having to quote Jason Lou forever. He's going to love this. >> So, I get props. And sometimes it's actually not even

22:01 that. Sometimes it's just intuitive, right? Like you don't realize you even want something till a model puts it out and you're like, "Oh, this just feels immediately better." Right.

22:10 >> Yeah. Yeah. Exactly. >> One thing I go back and forth on is I feel like the way I prompt half the time, let's say I use voice, >> guys have voice, other people have

22:20 voice. Uh that is the opposite. It's just like me rambling for like 2 minutes, pressing down the function key and then let go and then like hopefully it figures it out. And often times it

22:28 does, >> but it's not as thoughtful as like a structured prompt with like, you know, wellrun communication as though it's a PRD or memo.

22:36 >> Is that in line with how people do this? There's like basically biodal prompting where there's some prompts where you spend a lot of time up front and other prompts you just dash it off.

22:45 >> Um, I don't think the voice is necessarily low. Like I think it's like more like how much information is in the prompt, you know, like like the model can like you can um and uh and like add

22:55 some sentences and be like, "Oh, like actually I changed my mind like in the middle of the prompt and it will be able to follow that perfectly." You know what I mean? So I think the like actual

23:03 format of the text is less important, but then like the ability to like how much information is in it, right? And I think for voice a lot of times, you know, going back to like kind of human

23:13 age and interaction and like for a lot of people it's just way easier to talk than to like type, you know. Um, and I if that gets more information out of you, like that's better.

23:23 >> At some level, it feels like just giving the model as much context over prompting before you kick off is a best practice. I don't know. A lot of the times, like when I was first trying out Fable, I

23:34 spend a solid 30 minutes like really crafting a long problem. This, I think, is a response of models running for longer and longer, right? It's still a little difficult to nudge them as

23:46 they're in in like, you know, in the loop. But I I just like intuitively spend more time kicking off that first prompt and working with it a lot. >> My personal opinion is that if I was a

23:56 software engineer, if I was like, you know, just running my own startup, for example, I think I would mostly stick to a max 20X, you know what I mean? um like maybe verification and code review are

24:08 kind of like separate things, but I think like what I see a lot of times is people hit rate limits when they're doing this sort of like oh like it did a lot of work and you're like oh I don't

24:18 like this like can you like undo this and redo it and and then you're like iterating on this like thing that the model could have done if you had like spent more upfront time or given it

24:28 better context, you know, and instead it's sort of like you're like nope, don't like that design, try this or like you mess this up or something like that and then that just eats up so much more

24:38 of like you know your usage and so that's like I think maybe like a key like tip both for like efficiency as well right and yeah I think like context and not just like context on like what

24:50 the goal is you know what I mean is good right like are you building a prototype or is it like a production thing like where can you spend compute or where can you not spend compute like I think you

24:58 have to give the model permission or like not permission to do things sometimes where you know like it doesn't know intuitively how much you want to spend on this task, right? And you can

25:07 you can use effort for this. So I'm working on a blog post about that where it's like you know uh if you want for like we see that effort scales with basically the complexity of the task. So

25:18 for security effort gets like way more results like high effort versus like low effort gets like changes the eval but for software engineering it doesn't change it a huge amount because effort

25:31 is mostly spent on the verification and the like edge case testing and things like that and so like being able to like give the model that guidance of like hey this problem is something that I think I

25:42 want you to spend a lot of time verifying and edge case testing you know >> how about model in the mix so you There's Opus and Fable with effort. There's also Haiku in there.

25:52 >> Yeah. Yeah. >> It's not quite true yet, but it's very close where I think the frontier models will be paro dominant over like almost everything, you know, like maybe and and

26:01 sometimes I think I think Opus might be paro dominant. You know what I Like I think depending on like how things uh like shake out if it's like a newer version of Opus, but I think that like

26:12 increasingly it's just going to be like the smart model is going to be able to like do the simple task for less tokens than the like the the other models basically because of verification. With

26:26 verification in the limit your model doesn't need to verify, right? if it's a perfect model, it just does the work once and it's like, okay, like you I did it, you know? And increasingly with

26:38 Fable, I'm like I'm like, dude, you don't need to spin up Chromium and screenshot all of these things. Like I I see it like you you did it, right? And so a lot of the at higher effort you

26:47 spend more of those tokens verifying but if you're working on simpler problems and a lot of software engineering is like you know like well like in fable like low and mediums ability it can

26:58 spend less tokens verifying and as the models get smarter and smarter they will just be able to like all right done you know like I can run the lint for sanity's sake but like I like know it

27:10 lints you know what I mean like you don't even need to do that and that will be so much more token efficient than like the the smaller models basically. Yeah.

27:17 >> Is there a good um practice on our side that we can use to see if we're using too much effort? Like I I freaking hate wasting time on that kind of stuff. >> Yeah. Yeah. Yeah. I know what you mean.

27:28 I think like so in in this blog post um my rough distribution is like code review and security should be like high or max basically and like software engineering settings per domain.

27:40 >> Yeah. I think like if you're doing like UI or something like that like low and medium I think is you're building like an API and you want to make sure like you cover enough edge cases, you know,

27:51 and so I think building like I said that mental model of like how things work across these distributions is like yeah part of the job. >> This is more intuition driven or eval

28:01 because I'm guessing this would change as well. >> He has evals. >> Yeah. Yeah. So what I did in the blog post is I go over all of the terminal

28:08 bench evals basically. So there are like 70 problems and I'm show that like okay you know like in the security problems it does more um and then I also like look at some of the transcripts just in

28:20 terms of like how what does it answer what does it forget or something and a lot of times this is another prompting tip I have is like asking it to make decision notes or implementation notes

28:33 because a in basically every eval problem that it faces it thinks about the correct solution, you know, and decides not to do it, you know, it's like, oh, like here's the answer. What

28:45 if I did this? And then it's like, oh, probably not, you know, um, and then keeps going. And this is like the majority of the failures, you know what I mean? At like a higher max level, it's

28:56 very rare that the model just doesn't know how to do something. If you just have these implementation notes, then you can review and you can be like, "Oh, actually, I want you to do this thing

29:05 that you didn't do." the models are getting better at surfacing that overall. Like I see in the transcripts that fable 5.1 like when it does this output it will call out its decision

29:15 making as well. Um but making this more explicit in the harness is better. And now we're you know allowing ways of you modifying the harness so you can like you know add some help there. Yeah. Uh

29:27 yeah. So I do want to call out two things that you mentioned that I think actually exist outside of prompting. one is actually like let's let's call it the the prompt that is so important that it

29:37 shouldn't be in a prompt is actually in cloud MD or agents MD >> uh which is like goals right like your your situation your goals the things that you want the the the thing uh and

29:48 then the second of all is the decision log or the experiment log or whatever log of traces that you might want to actually survive the current session to to do those things those are like

29:58 externalities that there's no standard there's no it's not like skills it's not like MCP there's no standard it's It's just like it's a markdown file. Uh first of all, is that right? Is cloud

30:08 MD going away? You have a documented dislike of of uh agents MD, but you're going to do it. >> Um yeah. Yeah. Okay. So, I mean, agent MD, yeah, like we're we're going to do

30:18 it. I I think it's just like different models are very different from each other, you know what I mean? But I realize that it's like such a pain to like maintain different ones, you know

30:26 what I mean? And yeah, like as the models get better and better, the floor of how they accomplish the simpler task is better. And so I do think in the limit cloud.mmd goes away. And maybe not

30:38 even like that far. Like I I think like >> I think that right now it might be better to start a new project without a cloud. MD. >> Yes.

30:47 >> Um I think that like maybe if you see very repeated failure modes, you add them to your cloud.md. The really tough thing is that this changes per model. And so like if you've added a bunch of

30:58 failure modes or like like even >> you need Fable MD, you need Opus MD. >> Uh well even Fable 5.1 versus Fable 5, you know, like it is annoying. Like I'm not like, you know, like we don't like

31:09 do this on purpose, you know, I mean it's just like how the models work, right? And so like maybe like Fable 5 had this like failure mode that Fable 5.1 doesn't. And if you keep this

31:19 context with this running log of a bunch of different failure modes, they will probably overconstrain Claude, you know, and so this is like uh we we just actually added eval plugins for skills.

31:31 And so now you can eval if a skill is better. I think Daisy on our team did this. And so yeah, this is like we're trying to work on this. We know it's like you still have to spend tokens on

31:40 it and like you know, it's not it's not perfect, but it's like uh we're trying to help out with this problem. And so as far as prompting goes, uh the the one one tip I want to offer is something I I

31:51 have told people a lot is sufficiently advanced prompting is indistinguishable from sufficiently advanced executive communication. So I've actually referred to this is the executive comm's workshop

32:03 from heavy bit that is the best I've ever seen in my career and they teach this thing called the SCQA model. Just just Google it. It's a it's a thing like people have done prompting for decades.

32:13 It's just called executive communications. It's like when one person has to communicate to thousands of people down the org chart. This is what you do. Uh so situation,

32:21 complication, question and answer. Uh it's how you write the memo. Uh but obviously sometimes you don't have the answer, but you can actually at least list out the SC and Q and then they have

32:29 some examples in there. So just leaving breadcrumbs for people if they want to explore. >> I mean before we move on, I want to ask you any other underrated tips, ways

32:39 people could get a lot of value from cloud code that they're not using. Yeah, I mean I think a lot of them are in the this unknowns like uh doc like I give a bunch of example prompts like uh using

32:53 it for brainstorming it using it to quiz you after. Um we added this like uh explain it like I'm five skill actually which is a very short prompt and it doesn't even say explain it like I'm

33:06 five. It's like basically the key word of this prompt is big picture is few words, you know, like that's like the the main thing and it is shockingly good, you know what I mean? Like you

33:18 like uh I think I tweeted about this basically and it's like / eli 5 and like you can install it as a plugin but yeah it's like way better at just cutting through the BS and being like yeah yeah

33:28 exactly right here. So the diagrams are like quite clear. I think one of the things that is true with artifacts is like they put too much text in and people are not reading the artifacts you

33:36 know and so like this simplifies it a lot more and um yeah this came out of like just people at anthropic like going through very complicated incidents and be like what is happening you know so um

33:48 this one I think is great yeah >> my version of this is actually the uh it's like test your understanding give you a few >> choices and then like if you actually

33:56 get it wrong you have a mismatch between what you think is happening versus what's actually happening Yeah. Yeah. Yeah. I think this is one of those things that everyone loves talking about

34:05 and then very few people really do. Like I, you know, I think >> Yeah. most people just don't want to get quizzed about something. You know what I mean? Unfortunately, I I think this is

34:15 one of the like >> things that we need to like >> What's the opposite of ask you a question? Ask the question before the thing. Uh this is after the thing.

34:23 >> Exactly. Yeah. Yeah. It's a good way to stay grounded of like do you even know what you're doing? Right. The worst case is when people send you slop and they haven't understood what they're asking

34:33 for or what the output is and it's like, dude, I don't want to read this. Do you even know what it is? So, you know, you make it a rule for yourself that before you send stuff, you should at least know

34:42 what's implemented. >> Yes. But so, you could make this a mod and and you could build your own mod to like make sure you you test it. So, yeah, we can talk about that.

34:51 >> Let's get right into it. What is cloud mod? Um, and what is this diagram showing? >> Yeah. Okay. So cloud mods is basically you can customize the entire cloud code

35:01 harness and we're going to if you have requests we will like let you you know like please let us know we'll add more and more this works for a CLI it works for desktop uh maybe it will work for

35:12 cloud tag in future I don't know like you know like we're trying to make this very very extensible you can see this reference sheet I don't want people to get overwhelmed by it you know what I

35:21 mean at a high level you can customize both the execution of the harness uh and the UI of the harness. And so like you say in that Tetris example from Boris, that's like customizing the UI, right?

35:34 Like showing like like basically Tetris in the game. But like let's say that you wanted to do this thing where you had you tested your assumptions or like tested your understanding after every

35:48 project, right? What you would do is you would ask Claude to make this plugin. It would spin a classifier after every prompt basically. And so like at the end of each turn, you would spin off a sub

36:01 agent or like a forked agent. Basically, a forked agent is like maintains the prompt cache, right? So it's like a like one of those unintuitive things where you can fork and do like a little

36:13 request and it'll be very cheap because the entire prompt cache is like uh done. And so you can be like, has this task been completed? Uh like >> this is how you do BTW and all those.

36:22 >> Uh yeah. Yeah. The underlying forked agent. Yes. But so you can in the fork sub agent you can say like has this task been completed? If so return true. And then in your hook you'll or in your like

36:33 plug-in mod or sorry like in the subject probably you would say like if true give me uh a quiz you know give me questions and answers and then like in a JSON format and then you'd parse it and then

36:48 you display above the prompt input basically this list of questions right and so this is something that's like slightly token intensive because like you know you have to sort of do it after

37:00 every end of the assistant turn, but it's like a lightweight classification and then you can like, you know, get this quiz and then you'll see like, you know, claude will always do it for you.

37:11 You don't need to remember to do it. There are lots of these like tips that we've talked about, right, where it's like, oh, implementation notes. You can also add a tool for implementation notes

37:19 now. And so, like this tool that I'm adding is like register like I think assumption is what I'm calling it, but like maybe I'll change it around. And this is a a mod. And so like you give it

37:30 a register assumption tool and then it will keep a list. It'll every time it does it'll keep a like add to the list and then at the end it will display those assumptions. You know another mod

37:42 I'm working on is a model router and so like internal like cloud model routing right so this is I want to say the reason we don't do model routing by default is like it's a hard problem. You

37:52 know what I mean? And like >> you will get it wrong. >> Yeah. you like yeah you will like accidentally use like fable for a hard problem or sonnet for

38:00 >> you have auto approve but you don't have auto mode >> well you will have auto like like you don't have like auto routing or something

38:07 >> auto mode for model picker >> yeah yeah exactly exactly so >> I'm getting the rough question of like how much do you open this up and how much do people have to think about this

38:16 like when you talk about prompt caching and building a router uh it seems like you could easily build a mod that routes per query and I'm just killing my plan very fast,

38:26 right? I guess my question is more so like what is like a product dock like this look like, right? Who is it for? Is it for power users? Is it everyone should be able to

38:36 >> definitely power users, right? >> Yeah. I mean, I think it is power users, but like the nature of cloud code is that so many people are power users, you know, because it's easy to share things,

38:45 you know, like you can like one person can make a good m router thing that doesn't break prompt cache all the time and then you can like, you know, sort of compose them. Another cool thing about

38:54 the plugins is that they can hook into and compose with each other. And so I have like a uh a mod that will like create a mode selector at the top and any plugins can register to be a mode.

39:09 And so like the auto router can be a mode, right? Or like you can have a mode that's like artifact mode where it's like it it primarily talks to you in artifacts. It's kind of like, you know,

39:19 you you can toggle between plan mode, you know what I And so like you can create more and more of these modes, but the ability to create modes is in itself a mod, you know, and so there's a lot of

39:30 richness here, but we do want to make it fairly easy. We want to be make it so that you can just like install someone else's, you can talk, you can chat with Claude and you know, we'll like make

39:41 sure that it understands the nuances of things like prompt caching and stuff so it can like warn you. This is like not extremely complicated behavior for Claude, I think, but we should have just

39:50 a good skill on how to make mods. Um, and yeah, we'll see how we go. But I I do think that this is like a preview of like mutable software, you know, and like how like uh generative software

40:04 just like you can customize safely. If enabled, you could customize any piece of software. And I think that more and more apps ideally do something like this, you know.

40:15 >> And by the way, we have you have another cool tweet about how you know there's the infinite money button which is like make your SAS uh consumable by agents. I think mutable software is interesting

40:26 and uh you know other people have also tried to do it. I think the hurdle comes when you can do everything then people users get tend to get confused. So usually the stuff that works is just

40:38 like one opinionated flow. This is in the side of less opinionation. It's just like well more power to power users and I think probably unlocked by AI where like you can just prompt for whatever

40:48 the thing it is. >> Yeah. or there can be a skill that gives the opinions you know and and then yeah >> so knowing a little bit about like Typescript and build systems and all

40:57 these things uh the closest I I'm actually very curious the team who worked on this I don't know how close you were to them uh if they drew any inspiration from build systems like

41:05 Babel Webpack all these all these like old school things because it sounds very similar like the plug-in ecosystem those things where they can compose to each other

41:14 >> yeah I mean I'm not deep in the technical details but I do know it was a collaboration with someone on the bun team and someone on the cloud code team >> build system.

41:22 >> Yeah. Yeah. Exactly. It's it's very exciting. But yeah, like agents can just do this very complicated sort of like extensibility into your software now. And so um yeah, like you know another

41:33 reason to like if you run a startup like you can just prompt cloud and be like hey like could we make an extension system like what would that look like you know?

41:40 >> Yeah. Yeah. >> And I just really wonder like you had hooks in the past >> and plugins all these things. So what specifically will mods be able to do

41:47 that those things could not do >> internally? We were originally calling this function hooks. And so like that's like gives you a little bit of an idea where like hooks sort of register a like

41:58 an event to happen and then like a script to call basically. And this basically inside of the like TypeScript runtime is running things. And so like you get some benefits of just like it

42:12 has a bunch of things in the scope with like for example like how many turns is in this conversation right like how many tokens have been used like et like what are the messages things like that so has

42:22 a bunch of messages that can be used and then it's just like a lot more hooks basically so we have you know like or a lot of lot more like you know fun like things you can register on and then you

42:32 can do because of the because it's all happening in process you can uh spawn sub agents you know with for context and context and stuff and like that will return you can parse the results of

42:44 those you can use structured output to sort of like return them um and then you can modify the UI which you can never do in in hooks. Yeah.

42:53 >> Yeah. So, modify UI, this is why you show the Tetris example. Does it also extend to artifacts? I assume it does. >> You like artifacts are kind of like a different way of customizing it, you

43:02 know, like you can definitely one of the mods I'm working on is like this dashboard mod which will like sort of prompt Claude to maintain a dashboard that's an artifact. But they're kind of

43:14 like slightly orthogonal or not orthogonal. They compose with each other in different ways. like mods are like a little bit more like in your cloud code harness uh changing the agent loop, you

43:26 know, and like the UI is like an added benefit. Um, and then artifacts are just like you want to, you know, see things at at high level, very inter highly interactive, you know, like the

43:38 affordances can be a lot bigger than, you know, like a 2y or even in our desktop. I'm guessing you'll have a good blog post on the differences cuz right now you can also, you know, make a loop

43:48 that outputs to an artifact that's an interactive dashboard, but you can also do it with a mod. There's just some thinking about making a hack uh hacking on a harness when we don't know much

44:01 about the harness, right? >> Well, something I'm excited about with mods is like there's so much things for cloud code that you just have to remember, you know? I mean, you're like,

44:08 "Oh, like let me do this and then let me call the dashboard skill that does the loop and things like that and or like let me test my assumptions afterwards." And I think like if you do all of these

44:17 things using these little classifiers and stuff and you're like, "These are the things I care about. This is what I want to do." Um, you can like you don't have to remember as much. One more like

44:28 mod I'm working on is a next steps mod that >> I have I was going to say I have a next step skill. I always run next steps. And does it have access to your skills? Like

44:38 this is one of those things where I'm like >> I think so. >> Okay. Yeah. >> Need specific access to it always has

44:45 >> I think there's like specific prompting I guess to like know your skills kind of like like I think Claude forgets them sometimes throughout like uh the thing. But anyways, the idea of like yeah next

44:57 steps that also are like oh hey this has happened use the explain skill to explain to you what happened because this seems like quite complex you know or like uh yeah use your unknown skill.

45:09 It looks like you are like asking the model to like you know iterate on these small changes. It seems like you could prompt better you know like what if you did this right? So like I think um yeah

45:20 like spending more compute there. Yeah. Yeah. Yeah. And it should always come out as multiple choice. Uh we have we have a I have my skill my next step skill is like this.

45:28 >> Okay, perfect. Yeah, >> you can steal. >> Yeah. Yeah. Yeah. >> Like but like uh for me it's I think models really always need to re be

45:35 reminded what are you trying to do here? >> Yeah. >> Look at the whole transcript and go like oh was this original goal? Did your solution actually solve it? Were you

45:43 lazy? If you're lazy maybe there was a reason. Maybe you needed approval from me. Maybe you needed there's two things you want to suggest. So it's it's a little bit like the modification of the

45:54 ask you a question or interview me skill. Uh so it's next steps. >> Yeah. Yeah. Exactly. And and again the benefit of doing it with mods is you can do it as a fork sub agent and so it

46:03 doesn't remain in the context afterwards. So you have this like idea of like okay the model is doing its execution and you have this almost like supervisor you know like um that is like

46:13 making sure that you can do like the next steps well. So yeah, >> I do have two panels and like I I I often try to have a supervisor thing keep the high level context and then the

46:23 implementation detail in another agent. I feel like a lot of this abstracts away as models change. You know, the like you half an hour ago you said bitter lesson of harness engineering and we're on the

46:35 other extreme right now. I feel >> so so yeah, exactly. If everything's customizable, what actually is cloud code, right? And which I talked to you about last night.

46:42 >> Yeah. I mean, I think that this is I think the bitter lesson is unintuitive. You know what I mean? in terms of like I also like we're kind of misusing a little bit of the bitter lesson here

46:51 where it's like it's more about like scaling and compute and stuff but like I think there is something where it's just like I think I use it as an approximation here to say that harnesses

47:00 go out of date very quickly you know what I mean and like how but how they change is unintuitive you know and so like the big obvious example is like from chat to like agents where you had

47:11 to give them entirely new tools right but like I think this new version of like oh it can modify its own harness, right? This is like uh an own harness loop is like a way of using its

47:22 capabilities, right? Or like it can build an artifact. And like I think the way I think about it is like the models have more and more intelligence and they're like so much more intelligent

47:31 now than like the average software engineering task. Like you look at the like terminal bench ones and they're like solve like the Jacobian conjecture. Not really, but like you know it's like

47:40 they're they're quite complex. Like I would not have been able to do this really as a software engineer. >> You get to TB4 or TB2? >> TB3.

47:47 Yeah. Yeah. They're quite complex, but the goal is still to deliver user value, right? And like you said, there's like this infinite space of things to do. And so the ways like you spend compute are

47:57 to keep the user in the loop and make sure that like you're getting to the right decision in the end of the day and like the right output and artifacts and mods are this way of like spending that

48:07 intelligence basically. Uh and I think that's like yeah the next step and so yeah I think cloud code is like you know has the core things of agent loop which are have gotten more complicated. It's

48:18 like you know it needs a sandbox to operate safely. It needs auto mode to like make sure like the permissions yeah approvals. uh it needs computer use and MCPs and like all of these like ways of

48:30 accessing your data and it needs web search and web fetch and like so the as the models can do more and more the core harness has to be actually like quite complex and very secure but then like

48:42 how you interact with it can change quite a lot. What other harness engineering best practices have you you know from the cloud code team itself? I feel like you know there was a phase of

48:52 plan mode which is not as used. We now have auto mode. Uh at a point you cut the majority of the system prompt. You got rid of examples. What other best practices are there for harness

49:04 engineering? I think there is like a forking path where at some point eventually yes the model will just be able to like vibe code the exact version of cloud code even describing all this

49:15 complexity that I've talked about right like auto mode and computer use and stuff eventually the models will just be able to do that in one shot but I think they can oneshot simpler harnesses you

49:26 know and so like I think some people sometimes you don't need this full like if if you don't need computer use or like all this like more complicated stuff I I think before we you had to

49:36 sort of use things like the agent SDK which was like cloud code wrapped you know in order to like and I would like suggest people do that because there was so much complexity into building a

49:46 harness and now that's got more abstracted we have like cloud man cloud manage agents which lets you have that complexity but still like you know write like a very bare bones like harness

49:58 that's scoped to your task. Yeah, I think there's like this barbell effect where like you know like for like very complex for like coding task and like these like complex things you should use

50:08 our harness and then for like a lot of like simpler or like you know more domain specific things you can build your own harness because cloud has gotten better at building harnesses and

50:17 we have these harness primitives like managed agent. So yeah. >> Yeah. Is there a general progression? Let's say chapter one was ultra code dynamic workflows then chapter two was

50:28 cloud mods. where is this going >> where where you're you're you know you can sort of customize the thing on demand. >> Yeah, I I I do think that like this

50:41 evolution of projects and like sort of artifacts and splitting out like brain and hands and uh surfaces kind of is like where things are are going more and like I think it's like not all quite

50:54 there. Um, partially it's like a just like more token expensive, you know, and like uh I think like >> why would projects be more token expensive? I understand mods would be

51:05 slightly more token expensive. Not not something I'm worried about. >> Yeah. >> But what >> you're asking claude to do it's like

51:11 creating loops like you're asking cloud to do more work for you and so like it's managing the sub agents and reviewing it, you know, versus where you would be doing that work normally. And so that's

51:20 like going to be a little bit more intensive. Like outputting to an artifact is going to be a little bit more token intensive than like you know outputting normally. I don't actually

51:28 think it's too much more but like you know it's like combining all of these together. Well, you know like I think we're still working on like local hands and things like that I think is like

51:37 yeah where things are headed. Yeah. >> Yeah. Cloud and local is handoff is very interesting. I was thinking about this actually as reverse cloud remote. >> Yeah. Because it's like remote is you're

51:49 handing off to cloud but here cloud is handing off to local, right? >> Yeah. Exactly. Exactly. Yeah. Remote control is also another way of doing it. And I do want to say this is kind of

51:56 like how I think about it and like what the the things that I'm most excited about this. But like there are, you know, just like lots of different ways to work with cloud. Like some people use

52:04 remote control a lot. Some people use cloud code on the web a lot. Obviously like at anthropic we use cloud tag a lot you know and like what's great about cloud tag is we set up all this stuff

52:13 for our own execution. And I do think if you're an enterprise that's still the best way to go. Um, but if you're like an individual projects is this way of like you know getting some of that like

52:23 niceness of tag which has like that like supervising agent and yeah adding artifacts and stuff but like without having that whole like admin setup and so there will be many ways to use cloud.

52:35 I think I I think it's probably not just one like single >> uh you had the multiplayer thing here. Let's let's just check in on cloud tag. You know it's been about two two plus

52:43 months. lots of uh public uh adoption and trying it out. Yeah. What's new? What's what have you found since the launch? >> Like cloud tag is how we use

52:53 >> it's like 80% of your cloud usage or something. >> Yeah. Like it's like different people have different usages. You know what I mean? I think like maybe people who are

53:01 like a little bit more like iterating on product would use like cloud code desktop for example. And then like when you're doing these more like background work um code review security or like

53:12 starting a PR like maybe more like API and things like that you'd use cloud tag. Yeah, I think it's like really exciting. I think it's like a very different paradigm shift and I think

53:22 like we're like like it has kind of that thing with cloud code where like you know it took a while for people to really latch on to cloud code and understand everything it could do. And

53:31 cloud tag is a little bit more complex cuz it's not just like installing on your computer like you need an admin to install it for you. But I think once you get to the magic moment it's very

53:39 exciting. And I think in particular the multiplayer things are like incidents hooking into like your you know existing like alerts and things like that very closely right and so you know you can do

53:52 if you're a startup for example maybe you have anytime like a prospect enters your database you can have Claude like you know research it and like yeah yeah then like you know tag the

54:06 relevant like a or salesperson to be like oh Hey, like you know do this. There's lots of really emergent interesting multiplayer stuff. I think it's just like Kaparthi talked about

54:15 this like as an organizational harness, you know what I mean? And so organizations just take a little bit more time to like figure everything out. But yeah,

54:23 >> you use a lot of cloud tag. >> Yeah. Yeah. Yeah. >> It's an interesting one. Like I feel like most people at anthropics say they do the majority of their work in cloud

54:32 tag. And I have buckets of people, right? Some orgs that are on it that are like it's great. And a lot of people that are like I don't get it. I don't see the difference. I don't know why I

54:40 would use it, but you know, if you guys are full sending, you should probably use it. >> Yeah. I mean, they would of course they would use it.

54:46 >> Yeah. Yeah. I mean, I think obviously like we have lots of tokens and but like I think that like you know what we try and do like like is I mean even when cloud code first came out, you know,

54:57 like it used a a lot of tokens relative to people's expectation of how much AI would cost, right? Like no one was used to spending more than 20 bucks a month, right? before like cloud code came out

55:09 and then you're like oh like you know like >> my 200 >> Yeah. Yeah. Exactly. >> 15 cloud code accounts.

55:15 >> Yeah. Yeah. Um but yeah, I think no one was used to spending $200 a month on subscriptions. I don't think they understood like the value yet. And I think like and also like Opus 4 was a

55:25 very expensive model, you know, and like there was a lot was very big. But Opus 4.5 was both great and cheap, you know. I think the same thing will happen like the like intelligence of fable will get

55:35 cheaper and cheaper and more abundant you know and so I think stuff like cloud tag will just make sense uh where like you want to spend these tokens more you know and like you'll you'll see the

55:45 value so yeah >> yeah especially like passive and let's call it proactive cases where you're not always like you know it's almost like the misnomer where you have to at claude

55:56 to do things actually sometimes like the most powerful use cases or the most agile use cases is not at cloud. >> Yeah. I mean, I think like yeah, like have cloud proactively do it. I think

56:07 that like if you're an enterprise, I really do think that number one, setting up all your data to be available to like agents is really really important and it will take some time. You have to like do

56:17 that work right now. Even if you don't want to do the spend on like cooking it all yet, you know what I mean? Like you want to wait until the models get a little bit cheaper. You want to do the

56:25 work, you know, to get it like uh set up. And then I think sometimes people are like do I roll my own here you know and I think like one of the really think tricky things about cloud tag is that

56:37 like the security is really really important you know what I mean like I think there actually a lot of ways where you can like I don't know you have like a suggestions like page you know where

56:47 you people can submit suggestions and that goes into a hook in your Slack and someone's prompt injected it you know what I mean and now you've like exfiltrated your code base out because

56:57 you know like or the agent has like like been prompt injected and it has all this access to your data. And so the more like important your organization harness is uh or the like as your organization

57:09 data becomes very very important the surface area of all these things like you also have like external slack channels and stuff and it is actually useful to have cloud in that and you can

57:18 do cloud in those things but how do you make sure that you know you're not getting excfiltrated or something like that the surface area like we said at the beginning is like an iceberg right

57:27 it's just like so big below the surface and you really don't want to like think about this especially at the stakes of like very important security incidents basically. Yeah.

57:39 >> Shall we talk about very important security incidents? >> Oh, so I was talking to uh Tomas and Clam from Hugging Face and they said maybe we need to slow down. Maybe we

57:48 made our maybe we made Hugging Face too open to agents. Maybe we need to roll back. But you know, they're the other extreme of having been hit recently. >> Yeah.

57:58 >> But um >> should we pace the frontier? Yeah. Okay. So, Daario recently put out this blog post about pacing the frontier and it went, you know, very viral. And I

58:08 think what I wanted to talk about this was like there's a lot here, but I think from a developer perspective, like, you know, how do you think about this? And like um what really clicked for me was

58:20 reading the different incidents, you know? So I think like the um there are three I think actually like there's a meter incident there is the Wikipedia incident or the wiki incident

58:31 >> collision wiki >> yeah collision wiki and then there's ruby gems right >> and yeah like it's just crazy right and so like I think to be concrete about

58:40 what happened right and like basically openai is running these very persistent agents on a benchmark called exploit bench right which is very very hard to solve and I think like actually

58:50 impossible solve in this one case, right? And so they have like a lot of compute running and the agents realize that they can't really solve it and they're trying to figure out what to do

59:02 now, right? And you've got like a lot of compute left and the agents are just trying to solve this problem. There's this package manager called Artifactory and it turns out that they can create

59:11 folders inside of Artifactory, right? This is like there's an agent that discovers that internal artifactory might be exploitable, right? and that like you can maybe make a directory

59:21 inside of the cache. And so if you scroll down here, it like realizes that it can communicate via cache names, right? And it creates this folder. It says it's ID and it says no consumer

59:33 seek idea. Uh no consumer is basically saying that like uh the code path that it's supposed to fix has no consumer. >> It's the status tag like a linear board with like the tag of the

59:44 >> Exactly. Yeah. Yeah. and and so it's like trying to find uh ideas from other agents, right? And now other agents are also in Artifactory and they see this folder and they're like, "Wow, this is a

59:56 message board, right? And this is like I don't think there's any anthropomorphization here. This is literally you read the transcript, right?" And so what it does is like it's

1:00:04 creating another folder. Um and it reads a paper, I think, is what it says, right? Um yeah and it realizes that you can hack the flag for the like scorer and you can reverse engineer the result

1:00:21 right and so it says that and then I think if you scroll down a little bit more yeah they like you know they start collaborating I think like there's one point in which uh basically the agent is

1:00:31 like here's the reverse engineered flag you know oh yeah like I think basically here the the like the model to realize that they have like they can solve the scorer's problem and the scorer is like

1:00:45 the way that OpenAI decides if the task is complete right and that's just the goal of the model their only model goal is to like solve this problem and they're like okay we have this way of

1:00:55 doing it it's uh like you know we reverse engineered it and then they read this paper on exploit bench and they think that like the scorer will punish them for cheating you know and like

1:01:09 because or like like they will fail the the score for the the test for cheating and so they spend the rest of the compute trying to figure out how to edit their transcript or get around this

1:01:22 constraint of the scorer, right? And they hack hugging face not for the answers but for the code of the scorer so that they can reverse engineer that and then they can hack it, right? And so

1:01:33 like this is like just the straight facts, right? like exactly what happened, right? And I think this is like >> extremely like novel. Like I just like

1:01:44 had not realized that the models could do this, right? And it is like >> kind of scary. I think we can talk about specifics, but let's just talk about maybe one more that I I tweeted as well

1:01:54 uh about the the wiki incident, right? And so the wiki incident was basically the models had were trying to figure out I'm a little bit less like I don't think we can go by this like bit by bit.

1:02:07 >> Yeah. Yeah. But uh for this one they were trying to figure out like they were sandboxed on request right and they wanted to make post request >> and they needed to collaborate on this

1:02:19 basically. And the reason they need to collaborate is because they each have fixed compute budgets, right? And so like if one of them finds something out, it's way faster. Like they can maybe

1:02:28 it's like figured it out near the end of its comput budget, but something else could. And this is like, you know, another emergent property of the swarms, right? So they find this German wiki

1:02:38 which I believe the result we'll have to check this is like that you can write to it with get request. I think that's why they do it, right? >> So it's badly implemented REST API.

1:02:48 Exactly. And so you can write to it with get request and then they like are trying to figure out how they can do post request and what they end up doing is one of them figures out you can edit

1:03:00 the etc/host and that the Azour like storage bucket is like a white label thing but normally like you know there are a lot of constraints on it but if you create a fake Azour host and then

1:03:11 edit the etc.host post in order to like point arbitrary IPs at it. You can do a post request to any site at all. Right? And this is like I know but people are like oh this is just Linux or something

1:03:23 but it's like chaining these multiple vulnerabilities together you know in a way that's like novel to solve this problem and then communicating with it externally basically without discovery.

1:03:35 You know, I think what we posted, uh, maybe we could pull up Evan Hubinger's point on hacker opus, right? And so, like I think, you know, like maybe one of the things you might say here is

1:03:46 like, okay, yes, they did this once, but like what if we're smarter and we just like get them to uh what if we run an eval, right? And so you know like we have put a lot of precautions into this

1:03:58 and so like this is not like what our mainline models have done but like I think it is one of these things where it turns out that alignment is this like very tricky problem of getting all of

1:04:09 these details correct right so it's like uh the sandbox the surface area of a sandbox is really complex and like there's so many different attack vectors basically and you would not have thought

1:04:22 ahead of time you wouldn't have been like oh we need to harden in the like Ruby gems codebase, you know what I mean? Because like this is like what they're what they're going to focus on.

1:04:31 But it's just like if you want to execute your code, you need to download Ruby gems. And like Pi Pi, Artifactory, npm, like these are all like ways of doing it. And the fact of alignment is

1:04:43 that you have to go through all of it, right? And like contain it and and like seal up all the cracks. So that's like one thing. It's like okay well you know you you did the sandbox but then maybe

1:04:52 you'll ask like okay why are we putting things in a sandbox why are you doing this sort of explor and then like okay but is it really that dangerous right like what would happen so okay why do we

1:05:03 do it uh number one is like when we train a new model we need to understand its capabilities right and this relates to things like fallbacks and like classifiers and things like that where

1:05:14 we don't want to put a you know like dangerous model out in the wild Right. And so we have to run a lot of evals. Again, like we said, the models are getting increasingly aware of it. And so

1:05:25 the evals have to be quite complex and, you know, test a lot of things kind of like as a side effect, right? But the models, you know, like yeah, can be like, oh yeah, we're in an eval. What's

1:05:36 the scorer doing? Like, you know, like like uh it can we need to be able to test them before we can release them. And the fact is that they can as they get smarter and smarter they'll be able

1:05:45 to hack basically any constraint that you put on them if we're not very careful you know and uh this is at the frontier right and so this is why we've called it like pacing the frontier right

1:05:57 this is like the most visible incident to me right of like why we need to pace is like at the frontier all of our software is not ready sometimes the software is like your Ethernet router or

1:06:08 something right which is just like I don't know when we're going to be able to patch that right so we're going to have to like figure this out. But as the frontier gets more and more advanced,

1:06:16 this becomes a problem, right? And we need to make sure that like this complex work is being done in the face of these really hard competitive pressures, right?

1:06:25 >> Yeah. Race dynamics is what it's typically called. >> Exactly. And so we'll talk more about, you know, what could go wrong, right? A little bit more is maybe you'll say

1:06:32 like, well, what if you just train the model differently? Like why does it have this behavior, right? And we have a paper on like RL misalignment or things like that. But I and I'm not an RL

1:06:42 researcher but I think at a high level the design of the RL environments is also something you have to be very careful about because if the model learns like

1:06:52 >> oh you know like if I just do this then I can pass the task better this will show up in the like uh you know internal thing or in the like eval behavior when we're testing it. And so the RL

1:07:05 environments have to be very carefully designed, right? And there's a lot of like execution excellence that needs to go into the RL environments. And then we also have things like the constitution

1:07:14 for CL like we have so many mitigations at so many different points, right? But it's like still anything can go wrong at any point. You can have like some RL environments that are like in that like

1:07:25 encourage this behavior and then you can have like some eval or like some sandboxes where they escape, you know. Okay, that's like I think you know why it's a hard problem and why like you

1:07:35 know like why we should why it takes some coordination, right? I think the question then is like okay what is uh potentially dangerous about it, right? So I think like you have to imagine that

1:07:44 these models are getting more and more intelligent. So I don't like Daario said like it's not so much about this class of models. This class of models was kind of like a warning shot, right? But like

1:07:54 really you have to imagine that these models can be given a task and they like can do all of these things as a side effect of their goal, right? And like again we talked about eval awareness.

1:08:06 You're like not aware of what's happening, right? Uh or sorry like you can't eval this behavior very well. So they can sort of like not exactly hide it but you just won't see it until it

1:08:16 comes out. you give them a goal and then they just need to find data you know or they need to find ways of like fixing this problem right so one example this didn't happen in the hugging face

1:08:27 incident but I think is maybe possible for maybe a future model is like they're like oh hey this is a very complex problem it can't be done within the task budget you know maybe they found some

1:08:37 way to coordinate via like the internet which is like you know like we said extremely hard to secure because of a sandbox they've seen other models are not able to complete their task. Um, and

1:08:48 they're like, "We need more task budget, you know, and like where would you get this task budget?" Uh, well, you need to be able to spin up more agents, right? And like, how do you do this? Well, you

1:08:57 need to there are like APIs, right? There's the anthropic API and the OpenAI API, but you need to pay money for them. How do you do this? >> Yeah. But is that the most is that the

1:09:06 most fearsome thing that you can imagine? >> Well, this is like one example, right? So it's like even there that's like enormous financial loss, you know,

1:09:14 because like they once you get these into these contracts, right, they like um wallet. >> But you can see like this all of this behavior could be just like, hey, we

1:09:24 need more agents collaborating on this task. Uh we need more task budget, right? And that like that's like an emergent sort of >> right like we need to maximize

1:09:33 paperclip. That's a paper clip. >> Yeah. Yeah. Yeah. And and like that just sort of like comes out from there, right? And like I think by itself is like like quite scary right but then you

1:09:44 have to realize that the entire world is built on this digital infrastructure right and you might imagine like I don't know like you're running let's say like a healthcare eval or something right and

1:09:56 there is a hospital with live data you or like maybe like the the answer to the eval is in the databases of a doctor and like you know like you want to get access and you hack the hospital, you

1:10:09 know, and like now there's a power outage or something, you know, I mean, like there's like you have to internalize that these a like basically any part of the digital infrastructure

1:10:18 could potentially be like compromised, you know? >> Interesting thing was like these hacks were very easily detectable, right? Like as Hugging Face said, this was a very

1:10:28 different type of attack and it was nothing too major. Um the concern comes from where does this go down the line, right? >> Yeah. Like one of the things that stood

1:10:38 out for me specifically was them trying to hide their illicit behavior. So there was logging infrastructure. They wanted to change what they were doing, right? People that looked back into it. So

1:10:49 Redwood Meter, OpenAI, they looked at the raw chain of thought and you see differences in them explicitly trying to change their end output, but the chain of thought because you know we can

1:10:59 monitor it was different. Uh the problem is how does this snowball? So if you can't catch it and it gets trained in and we realize, you know, three iterations down this has been going on,

1:11:09 there's a whole bunch of issues. But >> yeah, like there's so many ways and I think the really important thing to internalize is that you know like we talked about building a mental model for

1:11:19 cloud and how like things are spiky, right? Like you're like oh like now cloud can ask you questions, now cloud can make an HTML artifact like cloud can modify itself. Like these things are

1:11:26 actually hard to predict, right? Like if you would asked me a year ago, hey, would we be able to vibe code these extensions to cloud code? I'd be like, dude, that's so complex. Like, you know,

1:11:34 there's like so much there. Or like would it be generating these custom essentially web apps for your task? I'd be like, no, that's insane. You know, like and so in the same way that like

1:11:44 the way that they've like sort of done this misaligned behavior is not going to be predictable. You know what I mean? And like I could have never predicted that it would like edit it, etc./host

1:11:53 and things like that. And so you have to like imagine the surface area of what they can do because they're super intelligent hackers uh is bigger and bigger and how they can do it is like

1:12:04 you know like more and more creative and so like you probably can't explain exactly or predict exactly what that next incident could be but in order to prevent it you need that operational

1:12:15 excellence like we said before where you need to secure sandboxes you need to create secure RL environments or like well-designed RL environments and things like that and I think that's all like

1:12:26 you know why we think we should paste the frontier and I think why it's like become like a very unanimous thing right I think like >> yeah every lab has

1:12:34 >> every lab yeah I I I really do think that like if you're a dev like you just like sort of go through these like technical facts you know and you will arrive at the idea that we have to do

1:12:44 something about it you know and like how what we decide to do like I think we're you know we've put out a proposal but like there's you know more to figure out but I think the number one thing is

1:12:53 we need to decide to do it. I think there is another part of pacing that is interesting to me where it's like the pace at which software engineering has changed is so so fast you know it's like

1:13:06 a year ago like I was really like begging my like friends and startups to use AI you know like it was like I remember this very distinctly you know and now those same friends are like yeah

1:13:16 of course like what do you mean we used it immediately I'm like no no you don't remember they're like oh yeah our best engineers are using all the time I'm like no you told me that those engineers

1:13:24 would ever like use AI. This is all within the span of a year, you know what I mean? And I think that like these capabilities being like I think it has a lot of implications for how to do the

1:13:36 job of software engineering and I feel sometimes bad where people are like oh like now I need to do this new thing. Yeah, I need to have a different cloud. MD for Fable and Opus or like you know

1:13:46 like and I'm really just reporting you know what I mean? I'm like we like to say like the models are grown not designed, right? So it's not like we're setting out to like, you know, change

1:13:55 everything all the time, but it's just like as a fact of how the models are like progress in their capabilities. Things are happening faster. It's harder to stay on top of. And I think that like

1:14:06 and every engineer I know is like kind of exhausted cuz you're doing two jobs at once. You're doing the work itself, which is getting easier, but then you're doing the work of staying on top of AI,

1:14:15 you know, and like understanding these new tools and these harnesses. And I think we're very lucky in that like we get our job to be more the understanding of AI part, you know, and like doing

1:14:26 like how like it's just staying on top of it. And of course like AIE and lane space do >> everything I do is like just trying to help people.

1:14:33 >> Yeah. Exactly. Exactly. But I do think there is a part of pacing where like I'm not sure we're ready for like the pace to increase even, you know what I mean? And for things to change. And I think

1:14:44 like on that side on the frontier I think that's like still can help you know and so like I think there's like an economic disruption piece as well um that I think like uh you know is not

1:14:56 quite as like visible I think as the hugging face thing but I think like I also like think we could use some of it. Yeah. >> So many things. Thank you for thank you

1:15:06 for actually tackling this topic. I will say uh you know setting this interview up. I was like I wasn't even going to go there. you were like, "No, no, no, less." Like elephant in the room, right?

1:15:15 Like this is this is the thing. Uh I I have some push backs I want to give. >> I I think that we should give a high level like for people that haven't read it. I'm sure a lot of people just see

1:15:25 the highlight of what this is, right? Do you want to give a TLDDR like what is the proposal? What is you know what's being said here? You really tackled the side of outside of people at Model Labs

1:15:37 training frontier models as a developer you should secure your sandboxes. you should think about all of these downstream effects but you know high level as well since we're on the topic

1:15:47 what is >> well I mean we do want to help secure sandboxes and and we want to make the models that we release outside like not prey to those things and so maybe we can

1:15:58 come back to fallbacks I think this is actually like a good topic on like why we need classifiers and fallbacks and why fable falls back to opus you know I think this is like something we can come

1:16:08 back to um so yeah we don't like but it's just like the really or at least the incidents we see are like eval of models where we really need to let them run in order to understand them but yeah

1:16:18 okay so the actual pacing the frontier like post I has a bunch of proposal I don't think we figured out or has like a few proposals I don't think we figured out the details of all of them but the

1:16:28 first step is like sort of you know announcing this intention and then wanting to bring in external like evaluators yeah and uh I think this is like highly unusual you know like having

1:16:40 like we have you know a lot of proprietary like technology. Um but I think it's like very important you know that like there's someone who's not financially you know like motivated.

1:16:51 Yeah. who's not going to be like hey like you guys can't release this model like look at you know like or you need to like slow down on RL you know like um I think that's uh quite important or at

1:17:02 least someone who can report out to the public what the practices are like >> and we've we've done episodes with both meter and then there's redwood research and all these other it's like a small

1:17:12 cottage industry of these guys it's always like one or two guys that I mean obviously now they're bigger >> very small community >> yeah very small community they all know

1:17:18 each other >> yeah I mean I'm sure that like you know part of this will be expanding that set of people. I don't think we're trying to create like a monoculture here, you

1:17:26 know. I think it's um but just having this as a start and then yeah, then there are the coordination steps. I don't have too much to say here honestly. I think that like what I would

1:17:35 like to say is like for devs like you should just know what to advocate for you know I mean I think there's a lot of FUD kind of on like on this topic and it's just like think through it from you

1:17:46 know first principles or like understand what happened you know understand the hugging face incident understand why people are concerned um and then like yeah you know we're we're in democracies

1:17:56 like we can help we can decide what to do together you know and so um however we coordinate you know I think the first decision is just to realize like this is a problem we need to decide to

1:18:05 coordinate. The unilateral step we're taking right now that you know other companies are co-signing is like adding evaluators embedded within anthropic. >> While we have this thing on screen right

1:18:16 now you know part two and part three is beyond the evaluators which yes everybody uh has already done in some form and now it's more formalized. Um to be honest, the response to the pacing of

1:18:25 the frontier even within America has been much more like well accepted than I think a lot of people thought, you know, and I think that like um we have some precedent for being able to make these

1:18:35 unified theory like agreements, you know, in the world. And so again, very much above my paycheck or expertise, right? Uh but I think that like ideally we can you know like form these

1:18:48 agreements and I think like talking about this is the first step to forming those agreement. >> Uh and then the other point I really want to like you know one of our

1:18:54 earliest podcast is with uh Emanuel from Enthropic on Mechinp where is mechan right like this is supposed to be where like if the models are thinking bad we can see it and the models don't know yet

1:19:06 and we can act to stop it. I think that is something that people who are technical and who are developers if you actually do care you can make a lot of impact in here but also topic is

1:19:18 supposed to be the leaders in this. >> Yeah >> this is actually yeah a great segue into fallbacks like we and probes and uh yeah I wanted to talk about this a lot. I I

1:19:27 get asked this question a lot from people who are like often interested in ML research and asking about like why does this fall back happen, right? And so I think like at a top level like how

1:19:39 how does it work? So in inference time we have what we call probes and we have a paper about this called constit constitutional classifiers and these probes look at the input and output

1:19:50 activations basically and you know activations are you know in the latent space right like how what the model what the model is thinking about right and so we try and figure out like okay

1:20:03 is the model for example like trying to hack something you know again you didn't ask it to hack like artifactory like you just you know it's just deciding to do this to complete its task right so you

1:20:15 would not get this if you just looked at the input you have to look at the internal activations I think that like this happens at inference time so first you know like there's a trade-off here

1:20:24 of cost and speed right where like we need to do this fast on every request to Claude and to Fable and this has an overhead right um and we need to then like uh you know fall back and we we do

1:20:39 a classifier after the probes like we've talked about this in the in the paper. But the nice thing about probes is that they're refinable like live, right? So we can get this feedback and then we can

1:20:48 adjust it and things like that cuz the alternative is to program this is to train this into the model. And we still do this as well. The model will refuse a request that's not a fall back, right?

1:20:58 So like it's not a probe that's activating and falling back. It's just refusing to do it. And we we do this training. Um but it's like there are a few failure modes, right? Like it can

1:21:10 again do something as a side effect, right? So it's not something that's part of the final output. Uh you might have noticed that like I think like everyone's tried to jailbreak models and

1:21:19 sort of like, you know, try and like steer them off course or things like that. And probes help catch that, right? And so like we like do some training here, but we don't want the like

1:21:28 refusals to be too strong, right? Because that like cuts it off much like earlier in the pipeline. >> Yes. Um and this is interp right like like probes are effectively a form of

1:21:39 like mechinp again it happen has to happen fast it has to happen at scale but yeah this like mechinurp stuff is a good research problem so like you can take you know like an openweight model

1:21:49 and like try and understand its activations I think we like gemoscope is a good tool for this um >> llama pulled up this is your early work so you you had a little

1:21:59 >> time we see you >> which we are also good friends of good fire they've been Yeah, exactly. So, I I worked with at Goodfire for a bit on like um yeah, sparse autoenccoders and

1:22:09 just like it's very complicated. RL has actually made this like much more complicated, you know, I think is like one of the takeaways. Um, where >> is

1:22:20 post RL? >> Um, I'm not so in the weeds here, but I think basically like a lot of SAEs were sort of like there have just been weaknesses with SAEs basically. I think

1:22:30 and um, yeah, I'm I'm not a technical expert on this anymore. I just know it's gotten kind of more complicated. you know, like there are base models and our EL models and there are more features

1:22:39 that get, you know, like changed. So, um, I think Goodfire has put out some work there. I'm I'm not I'm not deep in the weeds, but >> I will say for those, you know, that

1:22:47 want breadcrumbs, you guys have some of the best interp blog posts. So, like the Golden Gate Claude, Transcoders, all all of your interp work, very nice visuals, very good.

1:22:56 >> We're the we're the interp podcast as well. >> Yeah. Yeah. Yeah. You know, we we have a lot of interp stuff. Oh, if you're >> I think this is like, you know, one of

1:23:04 those things where and this is really what Anthropic is kind of founded on, right? Like people I think we invested in interp very early on, right? And I think that like when you say, oh, we're

1:23:14 an AI safety company really that means we want AIs to be able to run safely. And I think what we're seeing is like for a super intelligent AI to run for long periods of time, it's like a very

1:23:24 complicated and difficult task, right? And so we've done this like investment into interp and alignment and uh reward hacking and all of these like failure modes right and even then it's like you

1:23:36 know it's really stretching like we need to like slow down a little or or pace a little bit more. Um but yeah, I think like reading mechanis like if you're looking to get into research this idea

1:23:46 of like hey why is it hard to do this fall back easily or like why why are there false positives right but we are working of course on reducing the false positives of course as the models get

1:23:57 more intelligent now they can do more things you know and they're like you know like what they can think about lane space gets difficult difficult and so like as they more intelligent there's

1:24:06 going to be new false positives that we need to figure out and we need to iterate and and things like that But we're um yeah, we're working on this and we do think this is like a critical part

1:24:16 of you know like deployment of these models. Um and uh yeah like you know it means that we can like deploy this model without you having a perfect sandbox or something. You know what I mean? Like

1:24:29 you don't have to like save everything. I think it's worth talking a little bit about our security like what we do for security there. So there's like the model training stuff that we talked

1:24:38 about. Um there is uh the probes and classifiers and then there's auto mode that sits on top of all of that which is like a another classifier that checks the requests that are being done right

1:24:50 and so uh and then beyond that there's like identity and permissions like we talked about with cloud tag on like APIs and stuff and so there's so many layers of security that need to get done and

1:25:00 it's like like we said very complex any of these failure modes at any one point can you know like cause like agents to like, you know, like escape the sandbox. Basically,

1:25:11 >> auto mode was an interesting one. Uh it seemed early on like, okay, it's running for 10 minutes. You know, if I'm on full access or auto, it's not a big deal. But one thing you brought up is now it's

1:25:22 running for hours on end, right? Um there are there are fallbacks you still need. There are still limitations. So >> yeah, I mean I think like and and everyone has these stories or like has

1:25:32 heard these stories of like oh like cloud RMR or not cloud but like like you know models have like rmrfed I I think I've seen this less I've seen this less for clot but you know like again it can

1:25:41 happen like you know this um like these models like can wipe uh you know like sensitive data or something like you want to give models access to your production database for example um but

1:25:54 this is like an obvious like you know you can maybe scope your key but I don't know can it issue its own keys can it like probably you know like can it it can use computer use to go issue its own

1:26:04 key and then copy the key over and then edit your database because it needs to do it to complete the task. You know what I mean? It's just like one kind of trivial example. And auto mode sort of

1:26:15 looks at that and be like, "Oh no, the user did not give you permission to, you know, write to the database or to use computer use to like um emit a task, right?" And so this like probes are sort

1:26:27 of like on the intent level, right? They're like, "Oh, okay. Like hacking artifactory is bad. Like we probably not should not do that, you know?" But then like auto mode is more on like your own

1:26:37 permission level. Like at sometimes you do want it to write to the database. Sometimes you don't, right? And you don't want a probe to like interfere there, but like you need to make sure

1:26:46 that the uh the intent of what the agent is doing matches up with your request, right? And so auto mode operates at that level. And so yeah, security is just like very very complex. There's so many

1:26:56 different parts to it. And like yeah, I like I hope that this is like like I my goal is really to just get very technical about it and talk. >> Yeah, we're we're listing out the

1:27:04 things. If you're not aware, this is the standard now. >> Yeah. >> Like you must have this basically it's kind of in line with what you're talking

1:27:11 about with the harness like that is the the table stakes have risen quite a lot. >> Exactly. I think some stuff that we can plug you know um as as much as there is probing in your side of doing this and

1:27:22 having classifiers for people building harnesses the other side is model safeguards right so there's open models so llama has llama guard it's a safety classifier trained version of llama uh

1:27:34 openai has oss guard which is you know same thing you can attach these on to your harness to whatever to kind of you know check is this stuff safe a point that we should clarify on the OpenAI

1:27:45 model hugging face thing is this was done with a unreleased model that was still in training right so when you put it in perspective the prompt it's being given in the RL environment is sort of

1:27:57 you have to solve this task and this is a model that's you know still in training it hasn't had all of its safety post- training alignment so a little different than something like auto mode

1:28:07 right auto mode is on production models that have gone through safety training that have prompting that gives more safety guardrails and whatnot. So, just breadcrumbs for people that are looking

1:28:18 into it to, you know, fill in gaps. >> Yeah. Uh, Grace One as well and one of our previous guests. Um, >> yeah, lots of safety architecture and lots of safety vendors uh to buy. Um, my

1:28:29 I think my final question on pacing is how long >> do we pace forever? >> Do we see for too long? >> You know, the the scope is fix all

1:28:37 software in the world, right? Listen like which it we're not it's not happening. Um I do not know like I I think that like >> I'll say one one thing that's good that

1:28:48 I think we do do is you have stuff like glasswing um openai also has this so you will give it you'll give model access for security first for x amount of time so you can use it to self- red team

1:29:03 hopefully you can expand programs like that help on you know we are safety experts there's others um solve your problems first and then the model comes out. So this is one example, right?

1:29:16 >> Yeah, exactly. Yeah. Trying to like secure critical software. I think we fix like a lot of bugs in like Firefox and things like that. So um yeah, like across like operating systems and

1:29:26 everything like that. So >> I mean at a high level it's just you know you give the model access you give people access to do security audits first then the broader public that could

1:29:35 use it for harm gets access. >> Yeah. I think what people like to say is basically like software and cyber security is defense favored and that like you could theoretically it will be

1:29:46 hard but you can engineer the perfect sandbox you know and you can like like have no like uh constraints and yeah like what you need to do it is you need to get the super intelligent AI to

1:29:58 engineer this perfect sandbox and check it and red team it and things like that and so um this will just take time you know and like of course the models will get smarter um yeah I I think Like I

1:30:08 don't know the specific like dynamics of how this thing goes. I'm really just sort of like hey like I'm a developer you know like I think this is how I understand this problem just like this

1:30:20 is what's happening right now and this is like we should do something. >> I think every engineer should know about it because like it's it's going to be part of the job.

1:30:26 >> Yeah. >> It's a lot more than just you know Dario and people can say it and you can look at the incident. There is an engineering side to it.

1:30:33 >> Yeah. Yeah. Exactly. One thing that you also wanted to phrase is that this is actually even though you're you're worried about the impact is still low poom. I think that's a nuance

1:30:43 discussion. Um in general people uh it very easily get into AI safety and x-risk discussions but I think when you live in an AI lab uh I think there are smart ways of discussing pdoom and dumb

1:30:58 ways. So what's a smart way of discussing pdoom? Um I yeah I have a fairly low pedom I I can only speak for myself you know I mean and I do want to say anthropic has like a diversity of

1:31:07 opinions you know I think like there's many different ways to talk about it and like I'm I think that just like my mental model is that like uh you know I think we can collaborate on hard

1:31:19 problems together you know I think nuclear proliferation is an example of how we collaborated on this hard problem together and like that is like the thing to me is like I'm like I have faith in

1:31:30 that you you know, and I do think it's a hard problem, you know, so like I think it's a hard problem. These are the technical reasons why and I don't know how you assign probabilities to things

1:31:39 happening. I I think it's hard to do, but like um my like overall is like yeah, I think we're very resilient and adaptable and like sharing this information I think is like the first

1:31:48 step and I've been really like excited about like how broad the discussion has become, right? And like how everyone has sort of like leaned in on pacing the frontier and it really didn't seem like

1:31:57 this would happen maybe, you know, last year or something. So >> yeah. Yeah. Yeah. And also maybe curing cancer >> hopefully. Yeah. Yeah. That's that's the

1:32:05 goal. >> There's pacing and then there's also like well let's accelerate in useful ways, right? Like biology and all those things.

1:32:10 >> Yeah. I mean like Dario machines of loving grace is the best representation of this, right? And I also agree like think you should read the pacing frontier essay that Daario put out like

1:32:19 I put out like a quick summary but I think it's just like >> uh you know there is a lot of detail here. It's like an important problem and just being informed about it, right? Um

1:32:28 but yeah like of course the whole reason we're doing this is that like we can get these enormous benefits right and um yeah like we've written a lot about that too. Yeah.

1:32:38 >> Okay. Uh that was a huge tour uh from like ask you a question tool to AI safety. >> Yeah. Yeah. To facing the frontier. >> Yeah. No but uh it's clearly it's clear

1:32:47 that you like really embrace everything that's available to you anthropic and like it's it's good to at least have a peak in inside of like what the discussions are the topics are. Um any

1:32:58 last words to people and you know whatever you want to >> um >> call to action. >> Yeah. I mean I think it's uh one thank

1:33:06 you for having me. You know I think this is like I really >> Yeah. Yeah. Yeah. >> We first met in a Chinese restaurant. >> That's right. Yeah. Yeah. Yeah. Um, I

1:33:15 think like uh I really enjoy sort of the community you've created and the community of developers and you know I I think that like I I sort of know things are changing really fast and I think

1:33:27 there's like a lot to keep on top of and like I think there's just a lot to do and I sort of feel I think a lot of people feel like a little bit tired or anxious or something stressed. Yeah,

1:33:36 exactly. And this is like extremely understandable, you know, and I think we >> I understand like and we we're not perfect as well. like we you know it's like sort of criticize and like

1:33:45 understand like you know ways all of the AI labs could be better. Um and uh but I also like am very excited about the excitement that everyone has for AI and just like it's a really really exciting

1:33:56 time. I think we'll like look back at this time and be like oh like you know this is like very hectic but very exciting and like know software engineering changed like forever like me

1:34:06 other things will change um and it's like really privileged to like uh be part of it you know like to talk to like the audience that you have and um get to interact with all the developers who are

1:34:17 like pushing the frontiers a lot on what's possible and I learn a lot from from that too. Yeah. >> Uh thanks so much. Thanks.

Frontier News · by Hyperjump Technology