Run a $10,000 AI Model at Home, Here’s How

summarized

TLDR

Open source models have narrowed the capability gap with frontier labs to 3-6 months, with recent releases like Kimi K3 and DeepSeek V4 Flash bringing agent-level capabilities previously only seen in closed models. The real leverage now lies in specializing open weights through fine-tuning, which requires good evals and high-quality data rather than massive datasets.

Key points

Open source models are now only 3-6 months behind frontier labs in benchmark capabilities.

DeepSeek R1 first democratized reasoning by publicly showing model reasoning traces.

Fine-tuning to 95-99% quality requires only hundreds of high-quality examples, not millions.

Fireworks AI started as a PyTorch platform before the ChatGPT era and pivoted to LLM inference.

The speaker recommends starting with off-the-shelf models and fine-tuning only after gaining traction and data.

Tools mentioned

Techniques

  • speculative decoding
  • mixture of experts
  • quantization
  • reinforcement learning for fine-tuning
  • evals-driven development
  • remote development environments
  • AI agents for CI automation
Transcript (captions)

0:00 All right. So, Ditro, you're the co-founder and CTO of Fireworks AI, and I want to hear your thoughts on the open source run for the last few months because it went from being quite behind

0:09 to being almost at the cutting edge. So, what's your thoughts about open source models right now? >> Open source has been gaining popularity recently. I think the general kind of

0:16 consensus that open models depending how you count on benchmarks maybe like behind three to six months based on progress compared to Frontier Labs. And I think what we've been seeing maybe

0:27 starting from GM 5.2 2 and now with Kim K3 uh also with like new Deepseek Flash which is kind of this jump in capabilities which for Frontier happened maybe like last November December where

0:38 everybody went from copy pasting code from uh from like agent window or maybe doing like only limited tasks kind of switching to really agent development sub agents running longer term stuff

0:51 doing loops doing goals etc. So this kind of level of capabilities effectively came now to open source models and that's why we seeing uh a lot of excitement and a lot of focus there.

1:02 >> So I guess uh I would say like the first you know really open source moment was obviously deepseek uh R1 then V3 then uh I mean Kim K3 feels like the second really hit and then literally like two

1:14 weeks later we got the Deep Seek V4. Which of these first of all let me ask you which of these like do you think is a bigger moment Kim K3 or the V4 flash? I I think probably K3 uh I I would

1:24 probably group like K3, DM 5.2 and uh flash kind of in the same category which is like much better capability of doing kind of a stuff like similar to you know jump when was when like Opus 4.6

1:37 whatever was really was released last year. I think that's similar capability jump and yeah you brought up kind of deeply 3 R1 really a push which democratized reasoning remember those

1:47 days like open source models thought I would probably be like llama which like simple chat bots the whole like 01 series from open AAI was like very cryptic and secret without showing you

1:56 the uh the trace and everyone came out and like I think the most fascinating back then was one and a half years ago the fascinating was that a lot of customers were like just look even the

2:07 consumers were just like fascinated by seen the reasoning traces for the first time, right? Because you wouldn't see it from like yeah because because being open source because being open you can

2:16 see it from R1 and it was just like fascinating to see how the model thinks. So that was kind of one huge uh huge jump in capabilities and yeah current jump coming from both scaling as a base

2:27 for general reasoning and just like much better RL for longer horizon tasks means that yeah you can you you can hand over if talking about coding or like some kind of coork stuff you can hand over

2:38 much much bigger tasks and not babysit the models model so much uh while they come both a fraction of the cost and being open allows you to customize it for particular application. I wanted a

2:49 great local agent that runs on my machine and is capable of doing actual work. So when Apoex open- sourced Frontier Agent, I had to give it a real shot. Frontier Agent is a framework for

3:00 running agents and it runs on their open weights 35B model which easily fits on my MacBook. Frontier Agent plans the task, writes the code, runs it, and gives you a finished result. And you can

3:11 do all of that 100% offline with complete privacy. But if you need more compute, you can run the same system in the Appodex web version. Now, I wanted to see what it's really made of. So, I

3:22 gave it the following task. How fast does a new open model get a quantized version that people can run at home? Chart the typical delay over the last 2 years. Then, a group of agents went out

3:33 and found the data on hugging face. And after that was done, another group of agents cross-cheed everything to make sure it's accurate. It then created the chart, reviewed it thoroughly, and gave

3:43 me the result. Everything was complete with sources and verifications for every single claim. So if you're serious about AI, give Frontier Agent a star on GitHub and grab the 35p model on HuggingFace.

3:54 And if you don't have a powerful computer, go to appex.ai. New users get free credits on sign up. It's the first link below the video. Thank you Apple for sponsoring this video. So what are

4:04 the implications like on the market on the AI scene as a whole? You think there's going to be like large shifts towards these open source models? You think there's going to be more

4:12 self-hosting, more fine-tuning? How do you see the whole AI field? >> Several factors from here. I mean once there is some price pressure uh generally and we see this with some

4:21 frontier labs also dropping dropping pricing recently which is I think is great for everyone to me like the most important part about about open weights models is uh ability to specialize and

4:32 customize and that's what we built fireworks around right uh I mean we we do a lot of inference but I really position as a special like specialized intelligent company right and what I

4:41 mean what I mean by that is that uh it's like DNA of Frontier Labs is build one single model which is going to be the best for all tasks. try to approach AGI ASI whatever you whatever you call it

4:53 from both uh kind of leverage unique data and insights for every business and use case and just driving kind of cost quality trade-off down uh we strongly believe in model kind of specialization

5:02 basically fine-tuning resin learning per use case and this area has been advancing dramatically over the past I would say like half a year or so we've seen a lot of folks who like uh start to

5:15 getting their product their product getting traction they get data they they start training models for for their use case and they're getting both quality improvements for their use case and uh

5:24 dramatic cost savings. And to me the most exciting part of kind of open weights ecosystem besides just driving usage of those box is this this ability to specialize and uh having having the

5:35 future where it's not maybe one or two frontier models controlling everything but really having like thousands or millions of specialized models for every niche and for every use case and of

5:44 course uh I mean as you mentioned I think there is kind of another angle of this about control and uh kind of reliability in some sense I think there's been a lot back and forth with

5:56 some frontier labs around both like login data and refusals and kind of changing models behind the back. So >> controlling not your weights not your model uh kind of mindset. Uh so I think

6:09 just from like I mean we we also provide primarily API right but because it's open ways derived models we you can download it you can run somewhere else you you as a as a business as a customer

6:19 are in control of what's uh what's been run I think that's also like very very important factor and finally there are local on like on device local use cases on your laptop on on your phone

6:29 >> so when people go to fine tune their own models what do you see like as the primary motivation is it the quality is it optimizing costs what's the most common reason usually usually quality

6:38 but quality can come from two two sides right it can come from side of okay I want to I have my my let's say I have my startup right I started as wrapping GP uh rapping in GPTO cloud right I got I

6:51 got user traction I got a lot of usage I actually have my unique insights of how users behaves how like make make experience better I probably squeezed out a bunch of quality from doing like

7:03 harness optimization and prompting and all of the all of this stuff And now uh now I want maybe I got to like you know 90% of what my ideal user experience would look like. Now I have the data I

7:14 have actual way to measure what is what is good what is not. Uh and fine tuning allows you to go from know those maybe like 80 90% over to like 95 99%. Because now I can not just optimize prompting or

7:27 maybe harness engineering. I can optimize entirety of the stack with my harness with my data with actual usage patterns and insights which I have uh bake some of that into model into model

7:38 weights and I can build the frontier model for the specific use case and in the process uh drive the cost savings which are always important usually with combination of that some people come

7:47 from okay my actually my out of the box model doesn't doesn't get me where I need to be I want to drive first quality and costless is a second factor some folks come from like okay I'm paying too

7:59 much for Frontier because I probably don't need uh general model from uh to to to drive my product I try something open source so maybe I try GLM or something uh it's good enough but

8:10 slightly worse let's say out of the box through prompting harness engineering I store some fine tuning it becomes better and suddenly I get like 10x cost host savings so it's really I would say it's

8:19 always like quality and cost ordering what which is first which is second kind of varies by the customers >> so like what gave you guys insight to start fireworks you know like a month

8:28 before CH GBT release because that was like after that it kind of became obvious that this is the you know going to be the next big thing and uh it's going to be a huge market and a huge

8:38 economical incentives to startups but you guys started like roughly a month before so what was the key insight >> yeah that's actually an I mean an an interesting story origin wise um I mean

8:50 myself and many my co-ounders were working previously on pytor which is as I mean you have attentional need on on your wall, right? This is like framework in which a lot of I pretty much all of

9:04 like deep learning models these days implemented or big chunk of that. Uh so we've been working on this from you know origin like 2016 or so to like 2022. So we went through this journey of kind of

9:17 adopting research patented research deep learning etc. that uh that was period of time where uh there was no foundational models right so it was prejai pre foundational models so you would I mean

9:28 maybe you would have some pre-trained transformers like bird but all but all of them had to be like heavily customized for for use case or people would just train models from scratch for

9:37 example for computer vision for self driving or use cases like that we saw the power of building AI for different use cases right we saw the power of deep learning and like scaling up data and

9:46 scaling up infrastructure allows you to train much much better models and build magical products So the initial motivation of fireworks was actually okay we there's a lot more businesses

9:56 which don't use deep learning don't use AI we're going to build a platform which going to help them and and uh so the kind of the motto of like model specialization building specialized

10:04 intelligent was there but because it was pre kind of genai large scale models uh foundational models era it was focused more on just like traditional deep learning and building kind of platform

10:16 around pytorch so that was like original pitch that's how we that's how we got started we had actually pay some fake customers. So it was going well getting traction but then as CHGPT got dropped

10:26 like maybe two months after we got started and uh after that like some of the like open source based model uh based models started appearing again it was all very nancy it was you know

10:37 before llama one early models which were I mean not not very smart but still impressive by those by those standards. So it was becoming clear that okay uh this kind of the segment of starting

10:48 from foundational models and fine tuning it a little bit for particular vertical or just using foundational model directly is going to be have like so much more applicability because it's

10:58 simpler and you can build on existing intelligence much more. It became clear that okay this segment actually is going to be much bigger than just you know deep learning training from scratch and

11:06 that's why we kind of soft pivoted. We want to help every business and every developer build specialized intelligence stay the same. But because of kind of technological shifts, the way we do this

11:16 uh changed. So that's why we focus pretty much exclusively on LLM. >> So right now, what do you think is preventing cuz I I think like everybody would like to have their own fine tuned

11:24 models, right? Like whether it's just cool to have your own model or you know that you have some data, you want to have a more efficient model for that specific use case. What is preventing

11:33 most people or businesses from doing fine tuning efficiently? because I still would say it's not a solved problem. >> Yeah. So I mean it's definitely harder than prompting and I want to be clear

11:41 that like even as all the advancements makes advancments that's not something which you should start as a as a as a developer prototyping your startup like certainly start with of the shelf model

11:51 if you get traction and if you get data and you kind of scale uh that starting to make to make sense and I think the turning point is basically evals right basically having ability to know what is

12:02 good and what is bad at scale for your product >> and that's what we see I think a lot of great startups do very well is basically focusing on evals because even if you're

12:12 not doing fine tuning you want to know which of the printing models you want to use you want to know how to optimize the hardness you want to know all this you know to know all the details uh so

12:21 starting starting from evol being like very rigorous and like do you evaluate correctly performance in your product do you know what users actually react to it do you have ability to kind of simulate

12:31 this on some like offline metrics and data sets and kind of are you constantly improving this as intelligence advances So you're able to capture interesting failures and cases where it's where it

12:42 doesn't perform well. I think that that's a crucial thing and that's honestly like what majority of like AI native company work goes to in the kind of evoling quality and once you have

12:52 that sure you can apply that to optimizing your harness and prompting but you can also take this and apply and apply it to fine tuning right you can take the same throw into and data sets

13:02 you can throw them to RL etc. So I think yeah the I mean the pretty much the binary criteria whether it it's going to work or not is like did you actually collect the data for your use case and

13:13 do you know what's good and bad looks like and being able to uh write a walls. I think what's exciting about like our kind of transition to our own kind of more post for training is that compared

13:24 to like we talked about earlier about traditional deep learning is that data set size like quality over quantity became much more important know if you're trying to train pre-train from

13:34 scratch you need to have like you know millions billions trillions uh tokens or images right I mean if you're trying to train computer vision model you have to have like you know millions of images or

13:43 something right uh you can get meaningful you know improvement with very high quality, you know, data sets of car environments in the order of like hundreds, right? Or even or even less.

13:55 Uh, and it's really important because there's already enough base intelligence uh in the model. So you you effectively try to like teach it something specific for this use case. Uh so it's much

14:06 easier to know scale and and quality of having I'd rather have like hundred of good environments and good wells than having like you know 10,000 but of questionable quality. By the way, if you

14:17 want to talk to other people who are on the cutting edge of AI, make sure to join my new Discord server that I just launched. It's completely free and it's the place where you can discuss with

14:25 other people who care about AI, who are serious about AI, and where you can hear about new projects and new startups I'm building first before everybody else. So again, the link to the Discord server

14:35 will be below the video. Make sure to join and send a message to the general channel saying hi. So is this like one of the main if you had to think about a business in the current stage of you

14:44 know the world is this one of the ways you would build a mode is like collecting high quality data that other people don't have and fine-tuning your own models or like what are some other

14:52 ways to think about modes because you know with software being so easy to build it it really shifted from what was considered having a mode you know 5 years ago or even like a couple years

15:01 ago. >> Yeah exactly. So I I think like the actual implementation software is cheap now. uh I think still building high quality products still has a has a

15:10 premium right I mean people call it taste or uh whatnot but it's basic uh but it's basically ability to understand users and iterate quickly yeah your iteration cycle became faster because

15:22 coding front end is not no longer a bottleneck but uh but I like but it's just a small like coding is just small part of iteration loop so I think still like building the great products became

15:31 faster but not infinitely faster and still requires a lot of insight and taste uh but uh yeah the big I mean the biggest leverage as implementation goes

15:40 down is unique insights about your area right so whether it's you know if it's if it's consumer product do you understand consumers and do you have data about product usage and what good

15:49 looks like if it's some uh you know application for particular vertical like do you understand that vertical really well and are you able to leverage unique data uh or leverage like unique insights

16:01 about about that area and uh And in those cases, yeah, implementation became cheaper. But maybe implementation was small small part of the loop. uh but abil but the ability to have insights

16:13 and have data and have evolves uh comes to the premium because it allows you to first of all build a good product prototype of v1 to begin with and after that compound on this kind of

16:26 self-improve self-improving data loop and you see this with a lot of a native startups right uh I mean coding is probably the most uh obvious segment right and cognition had you know great

16:39 success and kind of grown exponentially usage and uh they started with building great UX around printer models and kind of increasingly go going customization specialization there and uh since like

16:52 last year both of them are training customized models based on open weights because they have user data they have uh actual patterns and the data is the data is their mode they have better coding

17:05 data than print labs uh at that point so that allows them to build you know better models and hit higher quality in some cases or and definitely hit like quality versus cost

17:15 trade-off because you can specialize specifically for uh for like more narrow coding tasks and you and you see this you see this with kind of many customers in like other verticals some of the

17:27 people about publicly right like uh of folks like Perplexity or Gem Spark or uh like other consumer applications where like Harvey probably from kind of more specialized specialized verticals

17:43 for legal for example. All of them are you know transition from being just a rapper to kind of collecting data insights and usage and uh building the mod around that by specializing like

17:55 their entire stack starting from starting from harness and kind of product features all the way to weights. But for example, you mentioned Crystal, right? They I think they were like at

18:05 one point I don't know if this is accurate info, but they were running like over >> they were basically over 50% of your revenue. Is that accurate?

18:11 >> Uh that was uh I mean I think was that was the case for like pretty much all inference pro like open basic inference providers for us and our competitors last year. Uh that's no longer the case

18:21 because actually the market diversified quite a bit. So yeah, we had uh different concentration uh maybe last year. That's that's not the case anymore specifically because I because I mean

18:32 that's normal in a market where you know where you have frontier and breakaway capabilities. Whoever is kind of is is on frontier usually grows exponentially but then there is more uh kind of other

18:42 similar verticals and other similar products in different areas which kind of scale scale the same. >> Yeah. My question was about to be a bit different. It was like aren't you like

18:52 worried that you know once because it is a it is good that you know company get some data can fine-tune it you know can can run fine tune models but like if they have enough data like cursor they

19:03 can pre-train their own models right so like what happens like if some of your biggest customers switch from you know just running fine-tuned models for inference to like actually pre-training

19:14 from scratch >> first of all we can you know help them as part of the pre pre-training and scaling uh there as well I think that transition the whole pre-trading is

19:22 probably going to be very very small number of people very small number of people it makes sense to go there right because effectively you're signing up to be like becoming effectively in your lab

19:32 right and >> extremely >> it's extremely difficult it's it's extremely costly so I think it comes kind of on degrees on degrees of

19:40 customization right uh you maybe you start with life post training then you scale up post training then you start scaling up mid- training kind of trying to inject more of kind base knowledge

19:50 for your domain, not just like to tune the quality and like eventually you say like okay now I'm gonna raise a billion dollar and like going to invest it in getting a bunch of comput

20:00 and maybe like try my own architecture the last step is actually very very costly and if if anything like amount of people who do this probably went down over time right because if you remember

20:11 like I mean if you remember like original language like LLM wave we talked about you know 2022 right like I think the big theme of like 2023 was pre-training like see like mosaic ML was

20:24 like bought by data bricks it was like very very hot with many cu with many customers trying to do LM pre-training from scratch uh I think the you know representative there was like there was

20:35 Bloomberg GPT right where there was like financial financial models on all the Bloomberg data etc which ended up being beaten very quickly by by just generic pre-trained models uh precisely Because

20:49 yes like a lot of a lot about pre- training like compression general human knowledge and that's you know it's it's hard to do it's very costly and it benefits a lot from very aggressive

20:59 scaling as you know so which means that number of people who can actually do this decreases over time if they want to stay near friend here uh what's exciting about open weights because it allows you

21:08 to amvertise this uh amortize the cost right you can yeah >> can pre pre-train and compress it once but then specialize to so many different verticals propose training and I think

21:19 that's that's much more stable kind of uh equilibrium. Of course, there would be people who kind of maybe graduate to retrain their their own model but I think it's going to be very very small

21:31 number and uh there's going to be like 10x 100x more people who enter the funnel and uh and focus on building like for their niche but they might not never might never reach a scale where it makes

21:44 sense to pre-train. Why did you make the decision to not go into hardware and to focus on the like software optimization layer? >> It's a good question. I think um there's

21:55 a lot of I mean we look at it from the perspective okay how we can acceler can accelerate the mission of helping everybody to own their intelligence and specialize for their use case right I

22:06 think that's primary focus of the company and we see like a lot of value in kind of building the platform building you know building features around inference and training making

22:15 sure that that's kind of co-designed and coupled and coupled together well and that we are actually serving and customers uh at customers problems and I think that's where a lot of kind of

22:26 value in any gap uh is uh on comput side I think the I mean there's a lot of folks who focus on just you know scaling uh like sca scaling compute and scaling there are a lot of new clouds right so

22:39 there's a lot of value in being able to you know aggregate and abstract out that compute because it's so diverse and so heterogeneous and that's what and that's what we're doing so we call it like

22:49 virtual cloud effectively we run on I know several dozens of all this new new clouds and and what traditional hyperscalers you can abstract out so you don't need to worry

23:00 about any any of this uh whereas the actual GPUs or other accelerators are uh so from customer perspective you don't need to worry about compute procurement and all this uh all this pain but then

23:11 like does it make sense to actually know go on the lower levels of like managing data centers and energy and etc I think at this point for us it doesn't because even kind of computer orchestration and

23:23 developer platform and kind of focus on use cases is already a lot and I think as a startup it's very important to focus on on key differentiation and what you're doing well uh without uh without

23:35 being distracted by uh by other uh by other things. So that's that's kind of short answer. I mean again I don't I wouldn't preclude uh something like this happening in the future but the again

23:47 the key focus is still like comput orchestration and platform and we have working with many wonderful providers who focus on you know energy and scaling and managing data centers which it's

23:58 also very a very hard job. So you mentioned that you know being focused and knowing what your core bet is as a startup is one of the keys of building a successful startup like what what are

24:08 some other things that if somebody you know wants to build a new AI company in the second half of 2026 what advice would you give them what type of ideas would you entertain like how would you

24:18 approach this how would the first 30 days look like? Yeah. Um I mean coming back to like the idea that like implementation is the the cheap like implementation is cheap right cost goes

24:28 to to zero but I think insights are become come at a premium. Uh again we are we are operating more like B2B business right so it's it's it's kind of it's a different book kind of focusing

24:40 on focusing customers and scaling etc. I think there a lot I mean there's a lot of interesting companies to be had in uh in areas where like applying AI to particular verticals. So if you I think

24:53 we'll we are seeing this and we'll see a lot more of like folks you know from maybe some like manufacturing manufacturing niche or transportation whatever like all the different uh

25:03 verticals which are kind of more traditional uh kind of forms of business. I think there's ability to like bring AI there and if you can if you understand really deeply particular

25:13 domain you can pick up some amount of how you can do it differently with AI and build a gener be have the first mo advantage and be a generational company in in that area. I think that's one uh

25:26 kind of factor of that on consumer side it's just hard to say because yeah I think like consumer success is usually very hard to predict. Uh I'm sure we'll see a lot more like exciting application

25:38 just because cost of cost of implementation goes down. Uh but it's you know really hard to bet on like what's what's going to what's going to be there.

25:46 >> So I'm wondering like you know you guys are scaling super fast. I think uh 4x in about 7 months you know from the 4 billion valuation to like almost 18 billion. How does that look like in the

25:56 company? Like what are the things that are breaking? What are the things that are changing? Walk us through what it's like to go through this insane growth. Yeah. Uh I mean yeah I mean the founder

26:06 is kind of as a company grows you always need to be changing and adjust uh and adjusting processes. So it's it's pretty it's pretty exciting. Um I think we you know recently crossed

26:18 that bar number which is which means like you can't over 150 people so you can't really remember everybody in the face which is which is kind of exciting but a little uh scary milestone. I think

26:29 uh you know when companies scale and you know I've seen like I've been before at Meta for like 10 years so I've seen this kind of scale maybe not from not really from startup stage but pretty close to

26:39 that to kind of this gigantic organization and yeah like a lot of kind of communication and scale kind of communication breaks down as you as you scale up and that's you know that's why

26:50 big companies usually operate worse than startups right because there's less less internal coherence uh so several things we can kind uh have to to counter that. uh first of

27:02 all like organizing internal processes with AI helps a lot right because huge huge amount of uh kind of effort and kind of human work in traditional organization is just like aggregating

27:14 and passing information right like oh I have this team and I'm I'm going to have somebody whose job is literally like oh summarize goals and summarize progress and pass it on to maybe another person

27:25 on another team who whose job is similar and like align gaps and then have discussion etc like a lot of that information passing can be automated much more. So we obviously use a lot of

27:36 internal agents, right? And if before you know maybe had like Slack overload and you had to uh organize some human-driven processes to structure this, today you can throw agents on that

27:48 and uh and build upon. So can you give like one example like that maybe you know maybe don't give your best one that you guys you know have but like one example that you think companies should

27:58 implement like some specific AI agent that's like saving you a ton of time internally. Yeah, I I mean again probably predictable sense but just uh making like knowledge making knowledge

28:10 base kind of like making sure like information of your company like how it actually works is accessible to agents, right? Which might be as simple as dumping

28:18 important uh information in uh in like markdown files and putting them in the repo, right? and uh making sure that you have like uh if you want to add in some information there, you can ask an agent

28:32 to add to to an extra context there. uh you know some some things which uh which a lot of people I think also doing is like recording all the meetings and being able to kind of process and

28:43 summarize that's again uh very very basic things I think part which kind of exciting for us like from scaling perspectives there's you can do a lot more like kind of more closer product

28:55 technical you can do a lot more like automation in terms of what they're doing right like I mean for example we run a lot of specialized uh kind of specialized deployment for different use

29:05 cases where maybe even if model is not fine- tuned uh or if it is fine tuned but just like deployment configuration is specific to particular use case. So you end up with uh like different

29:16 configurations if you want to meet different latency targets or you're trying to optimize for cost basically throughut versus speeds stuff like that and a lot of that previously would

29:26 require you know human per engineer performance engineers who would focus on individual use case. You can scale that up dramatically with agents and you can focus you can f focus humans on setting

29:37 up you know the right benchmark and the right environments the right way to uh kind of automate and measure things. uh but then you can pour more compute and uh be able to scale to like more use

29:48 cases or more models or kind of like more hardware types because like your hardness and your like what you what you're measuring coming back to evolves right like in this case like let's say

29:58 performance and what you're measuring is the most important that's where humans need to focus uh but then like actual like implementation and kind of more like horizontal scaling from use case to

30:07 use case is more scalable with compute than humans and that allows you to know fundamentally teams more li and uh being able to you know sc like don't have all those like

30:20 off and square overheads which you have is like scaling organizations humans I think that's today probably easier to do on more kind of engineering technical sides obviously I think the moment like

30:31 on kind of GTM side there is yeah for a lot of force multip multiplier from again propagating information aggregate information getting insights of the business but still kind

30:42 human interaction side uh side of the story takes bigger part of people time which mean that yeah kind of G GDM side of the uh company needs to like scale more

30:53 uh scale more proportionally to the number of customers than engineering technical side so that's kind of how I would say but yeah advice on uh automated stuff I mean yeah if you are

31:02 even simple sense right like I mean if you don't have your agents in slack and everybody in the company is not using them yeah that's you're already behind And just giving

31:12 those agents the right uh the right tools is important, right? So you you should definitely have like a few folks in the in the company who like legit thinking about kind of a infra

31:25 for for your business. Uh and that can be that can be using other building something from scratch. That's actually I think less less important. What's important is are they actually thinking

31:36 okay for the process which happen in my company do I do I have the like can can the AI do some part of it does it have the right inputs and do I have like right way to monitoring what goes well

31:47 and what what doesn't go well and we've seen this you know once you kind of set it up uh like ex like people love to use it and people kind of pick it up both on technical and uh and nontechnical sides

32:00 because yeah it makes you more productive and people want to focus on you know creative insights and kind of interesting kind of scaling problem of process design not on just being

32:12 information faster from from one S channel to another to to Google doc to to a slide deck that's I think that's actually exciting >> let's talk about inference for a bit

32:22 what does it take on a large model like Kim K3 to run at very high tokens per second like what's what's the things that people are missing usually I >> I think it's you know kind of core idea

32:32 like if you ask AI to list you like the I this the list of techniques for making inference fast. It probably will give you like 20 30 bullet points which will be mostly correct. Uh the hard part is

32:43 actually making sure that they all uh kind of work together in practice and kind of code and code design in in actual implementation all the way from kind of hardware orchestration and GPU

32:54 kernels to kind of how you shard and split it across multiple uh multiple GPUs or other accelerators to how you kind of orchestrate different traffic and route different requests to

33:05 different places. So all of those kind of layers need to kind of come together to provide a really reliable and fast inference. So and that's kind of what makes kind of inference inference job

33:17 fun. Uh and that's why I was talking earlier about kind of dedicating not just models but also specializ in model deployments for different case because trade-offs uh would be different

33:29 right like if you want to run Kim S3 really fast it's probably going to cost you more and it means that your servers will be configured differently from how you would serve Kim S3 for overall

33:40 throughut like for I mean for example you would probably you know invest a lot more in speculative decoding and you're going to try to kind of burn compute to uh get the answer faster. You're

33:49 probably not going to load your servers very high very high basically don't have like very high batch size. So the proiness uh process is faster. Uh the consequence of that there are like

33:59 implications of how you like shut and partition models across different GPUs like what would be your replica size etc. And it you know it becomes like very important to make sure that like

34:09 when new requests come in they don't get like blocked and queued up across others. So like kind of you start optimizing your load balance and to be on super low like TTFT like basically

34:17 time to first token. So we get you get fast fast answers on the opposite sides uh if you're trying to for example you want want to just lower cost and you're running some ground agents 20 tokers per

34:28 second or 10 tokens per second is good for you. uh to those become different right that you now want to like load up those diplomas as much as possible have huge batch size so everything is gets

34:38 like all memory reads gets amortized on GPU so you probably your deployment side is going to be huge and going to be spun spend multiple GPUs just running those gigantic mixture of experts layers with

34:50 really high batch size uh and again your routing will now be much more focused on making sure that like everything is stocked up to the to the limit and kind of chugging along and spitting out as

35:03 cheap as possible. And uh I mean those are just kind of two examples of just configuring the deployment. But as I said like even before you go into model specialization there's a lot of sense on

35:14 let's say quantization trade-offs or training customized speculative decoding models and stuff like that which you can do to specialize particular use case. Even if you don't change the model, your

35:24 domain probably has particular data data patterns and you can train a better speculator for that. Right? And that's some of the things which we we our platform does does automatically for

35:34 different use cases. >> You briefly touched on background agents. I'm wondering if you think like running agents will move more and more towards the cloud the same way like

35:44 before AWS, you know, most companies who had mainframe or their own servers then they moved to cloud like once they needed more power. Do you see the same trend happening?

35:54 >> Um it it it depends right. So I think um it really depends on your case. So when I talk about background agents, I'm mostly talking about you know application like application use case

36:04 where you like human is not in front of the agent output. So you actually don't care about latency in this case. You know the actual sandbox can run on customer side and model can run on

36:15 fireworks for example. you I mean usually for identity cases because single tool call is probably going to take few hundred milliseconds and you know syncing between pool calls is going

36:26 to take few hundred millconds like network latency is not such a huge factor so uh even if you run model in the cloud and your uh hardness locally or the opposite it doesn't make that

36:37 huge that huge difference performance wise so I was just talking about general like speed uh yeah in terms of deployment cloud I mean I think majority Majority of products do run agents in

36:49 the >> I meant like for for development >> in the cloud >> for development. Uh yeah. So I think >> they use

36:55 >> for like coding. Yes. Uh I think I I mean I think so that that makes sense and you see this with you know multiple products doing this right. I mean Corser >> like David pushing a lot pushing a lot

37:07 of clouds because >> and and we see it also internally we uh I mean when we for example develop for GPUs right? you don't you're not going to have your like we were in the

37:18 sitation from the beginning right like because you you're not going to have your GPU in like the center level GPU under your desk it's just not very practical so you

37:26 kind of even in pre-agentic era we would all like develop on remote machines you know connected VS code with SSH whatever dev containers uh so yeah it's it actually like very natural that okay now

37:40 I'm if I have my remote set up I can just throw agent in there in the same docker container and actually don't have like ID attached to this. This transition was even easier. Uh yeah, so

37:50 I think it makes sense for increasing number of applications. Uh because yeah, if you if you don't need a human sitting in front of uh the laptop, you can just I mean you can scale much more

38:01 horizontally, right? Uh I find all this uh you know there was a way of you know smaller uh small thingies for your laptop so the screen doesn't close when you when you travel around. I think

38:13 it's, you know, it's kind of a silly sign of this transition, right? Yeah, it makes sense to transition your dev environment to the cloud. And again, like from my perspective because I was

38:22 working with GPUs for development, that's been already the case for like years. you can cloud and yeah I think like all even preent all the tooling was mostly there for the for doing remote

38:33 development and it allowed you to scale better and this just totally makes sense because you can spin up you know tens hundreds thousands of parallel development so I I yeah I see like

38:45 current you know opt optimal interfaces as of today is yeah it's a it's a slack chatbot agent whatever virtual co a coworker whatever you call it right who you can ask to you know do

39:01 fix some sins and then go spin up spin up an environment sends you you know sends you a preview of what it works or like whatever report makes sense for your type of product maybe maybe you

39:12 review the code maybe you don't depending on the trade-offs for your applications but your prim your primary interactions become this kind of like asynchronous and embedded in where in

39:22 where you do development and we see a lot with software development but also with like other tasks, right? Like I think one of the early things we did was like hook up like model evaluations,

39:33 right? Like because we run a lot of benchmarks, we run a lot of evaluations for that. We just like hooked up that uh like set up this agent that like allowed him to like scale on to to more GPUs and

39:43 set up this infrastructure hooked up to slide and like pretty much everybody doing just if I want to run some eval on some model in some cases a lot of people just do use it as use this like bot

39:54 internally and it's been like very powerful. Yeah, I mean makes perfect sense and I think I'm approaching it from a like a practical standpoint where like you know

40:04 if you run couple agents locally it's fine and again I'm talking about like active level development not building apps on top of it but like if you think about it from first principles and if

40:13 you you know I think all of us believe that like we will run more and more agents there's a number which your computer cannot even handle right and even if you have the latest MacBook with

40:23 like 128 GB of RAM stuff like that you probably cannot run thousands or tens of thousands of agents on your computer. And also Greg Broken tweeted recently that like the agents will need more like

40:33 each agent need will need more resources like even you know the next generation of models whatever they will be running tons of database tests or you know computer use or or just way more

40:43 computation way more resources that can fit in a consumer you know desktop or laptop then by that logic even like working actively maybe you have like your main agent you know that's your the

40:56 one you're interacting with but like if see the future where we can productively use hundreds or thousands of agents live in coding. It has to be on the cloud. And you know the the the biggest

41:07 realization for me came from like when my cursor GUI started crashing at like 30 agents in parallel where each agent was like in a different work tree and then I was like okay so it has to happen

41:18 in the cloud because you know even having like a top backbook it was still crashing and then like okay I I I moved to you know herder CLI setup which is better because it's way more lightweight

41:27 but still I know that like soon enough that will lag as well and like soon enough like I do have to have some mini VMs or use one of these manage products from like like you said cognition or

41:38 cursor that offer these cloud agents and they're pushing them heavily. So I'm I'm wondering like how you think about this like what is your current coding AI coding setup first of all and like how

41:46 you think about the evolution of this overcoming mods. >> Yeah, I mean as I said that's that's already the case. Uh I think you know for me personally because or like for

41:57 kind of a big big chunk of our team like whenever you work with specialized hard like specialized hardware you kind of already was were in that world right you you were not running 20 agents locally

42:09 you were running like you were running like 20 different Docker containers on GPUs even even before and you just transition from having to example like cursor window attached to cursor window

42:20 not not attached. So I think that the transition was in some sense easier but we kind of mostly mostly did it and that allows you to also do a lot of other stuff right like I mean even simpler

42:32 since like like basically at this point you shouldn't be you know fixing your CI ever right like even if you're doing some code review mainly for your product because it's more complicated and agents

42:42 don't do good good job and you can't disconnect it but like simpler stuff like yeah CI shouldn't be failing right like if something fails on the PR agents should just jump in the cloud and I go

42:53 and fix it right like it's something which we have internally like been very helpful we have a I mean there are a lot of startups which kind of offer similar similar products so I think this

43:02 transition is pretty much underway or kind of closer completed for a lot of uh for a lot of companies and yeah it it makes sense because the moment you are more uh more detached the value of it's

43:16 actually running locally uh goes down but then the value of being able to horizontally scale build in the clouds goes up and and that's again like I think the highest leverage activity you

43:28 can do development wise right it make sure that kind of environment setup and CI and like whatever you are evaluating is actually set up correctly coming back to know when you ask me what what is the

43:39 mode for for startups and I told you you know kind of eval and knowing what is good looks like I mean the same is true for software development right like kind of in a more local localized scale like

43:48 you code is like code is cheap and less important but like the value of green CI and having good tests and making sure that uh you can kind of automatically

44:00 say whether you know your code base is moving in the right direction or not in the right direction is very important because you're going to dispatch all the hundreds of agents and if you don't have

44:09 like kind of the right structure and harness and ability to then to scale like what's it and be able to spin up the right environment and test the changes and kind of build evolve base

44:20 that that's up you know your your code base and kind of your development process is just going to collapse and it's not you know it's not very different from like kind of traditional

44:30 software development scaling up right uh I mean you would see that like in in kind of in big companies etc like you you often see like the most senior and experienced engineers actually working

44:42 working on like CI and uh making sure that like test and like style guides are in place I think what we are seeing now is uh since everybody got uh ability to use

44:54 agents for for coding if actually everybody transitioned from kind of being software like hands-on software engineer to being kind of TLC like TLC senior distinguish engineer running a

45:05 bigger team and that's why the same kind of mental patterns like yeah writing code is not the highest leverage activity you can do but setting up the right environment the right test and the

45:15 right like guidelines so all your AI you employees which you're ting actually move in the right direction like that's the highest leverage activity and I think that kind of mentally mapping from

45:28 from the idea that everybody transition to kind of being a TL of 100 agents uh like is is a useful way in my opinion to kind of think which you know which processes uh make sense in which areas

45:40 it makes sense to invest. So as this scale, what percentage of your day is still spent on building something and like what tools are you using when you're doing coding?

45:51 >> Yeah. Um I mean I think we companies we kind of use all of all of us. We use a lot of cursor we use uh also like codex GPT we use cloud code and uh code opus babble etc. So I think it's it's a

46:07 combination combination of all this. have some you know agents which we build ourselves as I said you know for stuff like evolves performance optimization uh so it's a kind of it's it's a mixture of

46:17 tools I would say both myself and company right like actually writing like typing out code is probably negligible at this point zero uh code review or like high level kind of high level code

46:30 review or like review with agents where you like have another human quiz some PR with agents written by an agent of another of the original author. I think that

46:40 still happens quite a bit and in my opinion it's still very important for like depending on the type of the product. Uh so for if you have something like standalone and defined and where

46:52 you can actually describe how it like how it should work in kind of PRD style you know in a document that's probably a good good example of a product where you don't need to look at the code anymore

47:06 and for us it would be all types of I know internal tools or something like yeah I want this functionality I don't care how it works it's pretty easy validate that it works uh where agents

47:17 no go break down at this point and kind of you still need to you still need to look at the code is like more complicated like integrated systems right like if I'm working on like LM

47:27 inference agent uh and I just let uh agents lose they're going to create a lot of complexity right you still kind of need they're going to write a lot of code they're going to kind of hack it

47:38 together to maybe work for this use case but uh in inherent complexity of the codebase is going to go up so I still see a lot of value in kind of human input installed get in driving this down

47:49 and re like at least high level reviewing all the PRs going in. So that's that's actually where a big chunk of the uh like development work comes in, right? So you might maybe you want

48:01 to implement some feature in complex system which interacts with a lot of things. Yeah, you're going to fire up different approaches and maybe like try this with multiple agents run and see

48:10 like what comes out more like more lean and more more well integrated and and less complex and maybe like iterate multiple multiple times to that. So that's probably how like day-to-day for

48:20 like uh for for folks uh on our team who like working on more like systems or like influence engine or like training training infrastructure looks like and

48:32 yeah as as models become smarter I think more and more of that can be uh can be captured but it's not not not there yet. >> All right so I want to end it on this note. Um thank you for your time. I

48:42 appreciate it a lot. Where should people go? >> Fireworks.ai. Uh so I mean again if you're if you're a developer if you just want to try to use you know try uh open

48:52 models you can go and try uh our serverless API that's pretty much compatible with like open AI cloud API so you can call it directly you can plug in it in your in your favorite uh in

49:04 your favorite hardness so you can try all kind of latest and hottest models usually available on day zero whether it's Ky or GLM or deepseimatron or or etc.

49:16 Uh we also have actually have new product which launched called fireworks nexus but with idea of kind of routt intelligent routing between models. So you can swap out uh kind of endpoint in

49:30 your local in your hardness whether it's open code or cloud code or whatever and uh enjoy uh kind of enjoy routing to open weights models for maybe simpler task and fall back to fronting models in

49:42 other cases. So that's another thing you can try and uh yeah if you're building your company and um uh definitely try open open weight models which progressed a lot as talked about maybe they they're

49:55 good for out of the box >> or or once your product hits some scale and some uh and some having some evolves and having some data go try go try fine tuning. So we have uh kind of

50:09 multi-levelled trading platform. So you can kind of try from more high level of like oh here's my reward function and a button uh we we take care of everything. But if you want to experiment and try to

50:22 implement your favorite RL algorithm or whatn not we've got you covered too. So you have kind of lower lower training SDK where you can customize loss. and customize your data sampling all the

50:34 nitty-gritty details if you want to do kind of more research stuff. >> Awesome. I'm going to link all the products and your socials below. So again, thank you Metro for your time and

50:44 have a great day. >> Yeah, thank you. Thank you so much for having me. Eat them.

Frontier News · by Hyperjump Technology