I Built a 32GB Local AI Server for $1,500

summarized

TLDR

A $1,500 32GB VRAM local AI server built from three RTX 3060s (12GB+12GB+8GB) performs surprisingly well for LLM inference alongside a 5950X and 1,000W PSU using risers on a B550 board. The triple 3060 setup achieved ~50% of the text generation speed of a single RTX 3090 on Gemma 4 26B Q4 and about 44% on the slower Qwen 3.6 27B Q4, with prompt processing much closer between the two. The builder recommends the 3060s as a cost-effective backup rig, noting they are far less inflated than 3090s, but warns dense models like Qwen will be slower and advises using adapters or batching tuning.

Key points

  • The build uses three RTX 3060s (two 12GB, one 8GB) providing ~32GB VRAM, powered by a 5950X CPU on a budget B550 motherboard with PCIe Gen 3 x1 risers.
  • Prompt processing on triple 3060s vs single 3090 was remarkably close on Gemma 4 26B Q4, with differences shrinking to margin-of-error at 128K context.
  • Text generation on triple 3060s reached ~68 tok/s on Gemma 4 26B Q4 and ~17.2–17.8 tok/s on Qwen 3.6 27B Q4, roughly 50% of 3090 speeds.
  • Watts peaked at ~580W for the triple-3060 rig, well under the 1,000W PSU, making it an efficient budget option.
  • The base system cost ~$800 for CPU/mobo/cooler/PSU/case/SSD, plus ~$700–750 for three used 3060s, totaling ~$1,500–1,550.
  • The builder recommends using 12GB 3060s for 36GB total and notes that for image/video generation a single 3090 is still preferred.
  • Llama.cpp ran fine over Gen3 x1 risers (1 GB/s), while VLLM would benefit from full PCIe lanes.
  • The server was set up on Proxmox with Hermes Agent running in an LXC container using Llama server.
  • Defaults used were temperature 1.0, top-p 0.95, top-k 64, flash attention on, and no UI.
  • The system is ideal as a backup inference rig that stays always-on while the main rig undergoes tinkering.
  • Used 3060s are far less inflated than used 3090s ($1,000+ vs ~$250 each for 12GB 3060s).
  • No special batch tuning was applied; numbers are baseline and could be improved.

Tools mentioned

Techniques

  • batch tuning
  • flash attention
  • prompt processing benchmarking
  • text generation benchmarking
  • Llama server inference on Proxmox LXC
  • power limiting GPUs to ~85%
  • using PCIe risers (Gen3 x1) for multi-GPU
  • compression ratio management for large contexts
Transcript (captions)
What is the performance of 3060s like on modern LLM? We're going to be checking that out today and we're going to be using the rack frame that we used to have the quad 3090 build on. We're going to be evaluating these 3060s though today and look at the performance of them against things like Quinn 3.6, Gemma 4, and we're also going to test out Kermit's agent. I'm going to take you through the build for this, then we're going to go through the test with the 3060s, then I'm going to pop in a 3090 and get a test against that to see head-to-head. I think that is a good kind of same price budget point comparison. So, we'll also at the end of that look at all those results. Then, we're going to do a cost breakdown and I'll give you my conclusions. Let's get building. All right, we're going to get started on this now. And I will also be setting a power limit on these somewhere around probably 85%. Yeah, it's been a long time since we did something with the 3060s also, so I thought this would be fun. We've got the two 3060s in here that we featured several times and I added in another 3060. This was from the wife's desktop. She's getting a 5060 Ti now, so congrats to her on that very nice upgrade and this is getting her old 8 GB 3060 Ti does have 448 GB per second of VRAM bandwidth. However, it has only 8 GB of total capacity. They come in usually about 220-ish, it looks like used. Instead, the $250 on the 12 GB variant makes it a little bit better of a value. That extra VRAM really goes a long distance. And there we go. Let's get this powered up. Nice. I already had the 5950X, which about $282 is what I paid for that and that is exactly what it's about right now used at the lowest price. So, that one is really holding its value quite well. The side mount here, I think is a very good upgrade. If you had seen the way that we had it laid out before with the quad 3090s, we had it mounted up there against the top. So, the side mount here is actually probably a little bit better because it gives you a lot of free area in here and there is no like giant tube that you have to like deal with as far as where you place GPUs. That's very nice. Only 16 GB of RAM in this, that's just basically what I had on that system. And the motherboard is a B550 and this is a gigabyte eagle and the Wi-Fi 6 variant with five PCIe 16 mechanical wide. However, four of those are only X1 as far as the electrical. So, those are going to negotiate at Gen 3 X1 and that is why these risers make sense in that configuration. As far as the GPUs that we've got, I really think 32 GB is a great place to try to target. 24, which the 3090 has, is also a really great starting point. At the end, I think it'll really make a very clear comparison of what the differences it look like between those two. So, to start off, we're going to be using the 3060s for our testing and we're going to be loading this up against Hermes Agent and you can see here this is Proxmox on the desktop. If you're not familiar, check out some of the videos and also the written guides on digitalspaceport.com if you want to get up and running like this, which is a very great compact way to store a lot of machines virtually on one system like this one system is the prox three node in my cluster and you can see that it is the 5950X listed here as the processor. As we look at our Hermes Agent and the resources that I've got set for that, I've got the memory set to auto expand from 1 GB to 4 GB and four sockets total, so about two cores will be on that and it's about 32 gigs of SSD space that I allocated. You really don't need that much. And for our LXC container that we're running, even though it says Docker, we're it has Docker installed in it also. We're actually running this outside of Docker, but the Llama server that we're running here, you can see the model that we're going to be using is the Gemma 4 Unsloth 26B A4BIT, and that is going to be the Guff in UD 24KXL. The port is 9876, the host is the IP address of this machine, our API key is NerdTastic, tools we have set to all, reasoning we have turned on, flash attention we have set on, temp is one, top P is 0.95, top K is 64, and we have no UI specified. As we load this and get it running over here, it will basically be providing the service that we will use from our other machine. So, you can see the three GPUs that we've got specified here, the 3060 Ti and the two 3060s. So, roughly about 32 GB of space that we've got for utilization. Pretty awesome. Toss up in V top, so you can see the GPUs getting loaded up there with the model. So, let's get Hermes started up here, and this is hitting our endpoint, which is this machine here. Just to make sure, I'm going to toss it a howdy really quick and just make sure that we see some spiky little thinky actions going on. Looks like it. It is ruminating. And you can see the context being pulled in is 256. That's set to auto detect, so that is accurate. So, the next thing that I'm going to do is I'm not going to actually go through and score all these. I just want to see it do it, and this is basically going to my website where I've got all the information about the prior version one testing outlined, and it's very nice for a LLM to be able to parse it and to be able to kind of consume it and replicate it and provide you a big batch of answers all at once really quickly instead of going single one by single one. A complete installation here that is just fully local. This is very stock, fully local. It is very akin to the guide that I put together very recently that you should definitely check out where if you are local, I give you some real good meta in that that can help you get up and running with excellent speeds, especially if you're looking at Quinn. That in particular, 3.627B, has a couple of quirks you want to make sure that you have set right so that you are having the best experience you can running that model. And let's come over here really quick and just take a quick peek. And you can see So, it did a bunch of prompt processing. Now, it's doing TG. So, it's like 50-ish tokens per second, it looks like 49 tokens per second. So, yeah, it looks like it gave us answers to all of that all at once, which is exactly what we asked it to do. So, 26,500 tokens is what we came up with there. We'll type {slash} status. Uh I'm not 100% sure that this is accurate, so this shows token activity of 147,000. I don't think that's what happened at all. But, definitely there were some tokens that happened. Performance of this is still really good. I mean, for it to be able to go to the website, read that, answer the 10 questions, and do all that is pretty freaking sweet. And so, I am impressed with the performance of my new mini monster. I've got, of course, big monster, but this is mini monster here. And mini monster is good to have as a concept, like a smaller rig that you can make sure, you know, if you're offloading the big rig or you're doing maintenance on it or you're tinkering with it. Like, I'm non-stop. It's probably a result of being a YouTuber, tinkering with that machine. So, having a small machine that's always on is substantially less disruptive, and I can set that as the Hermes backup model so that it'll fall back to that in case the big machine's down. Freaking sweet. Let's take a look really quick at what the just performance tokens per second looks like on this. So, 1,000, that was almost 2,000 tokens per second. At 4,000, it was 3,200 tokens per second. At 16, let's see what it is. 3,500 tokens per second. At 32K, it was 3,200. So, you kind of the arc peaks right around 16K tokens, it looks like on this. At 64K, it is 2,700. We're still maintaining a really respectable level of performance. And this is not with any batch tuning or anything like that, either. So, you could definitely probably improve those numbers by a decent amount, as well. And the wattage right now looks like it's hitting 526 W. And I'll also run a quick generation test for you, so we can check out 512, 1024, 4096, and 8192 on text generation to see what that performance looks like, also, here. And that looked like it was about 22 GB max that we were using on the prop processing test on the three GPUs combined. So, this is a good alignment because it'll fit in the 3090 completely, also, when we test against that. The text generation at 512 was 67.78 tokens per second. And at 1,000, it held very close at 67.16. And on text generation heavy usage, the wattage is looking like 350 W right now. At 4,000, it's hitting 65 tokens a second. And at 8,000, we've got 64.48. So, next up, we're going to test the Quinn model on this and see what kind of performance we can get. I'll take you through the same steps with the Hermes agent. So, I'll quickly run through these settings. These are a little bit different, but I wanted to make sure that this was going to fit fully in a 24 GB footprint first off, so that we would be able to compare it between the 3090 and this. Of note, you would be able to fit a larger context window like 128k without any problem if you used the GPUs the additional VRAM that you get to get to 32 GB. Unsloth 3.6 Qwen 27B U D Q4KX L. We've got tools on reasoning off. That's important in VLLM. I'm not sure if that's going to translate to be important here or not. FA we've got set to on. Temp we've got set to the recommendations from Unsloth which are temp of .7, top P of .8, and top K of 20 with a min P of zero for general-purpose tasks which is what we're pretty much going to be asking it to do here is some general-purpose tasking. And it looks like it started up there. We've got the presence penalty to 1.5, no UI, parallel set to one, fit turned off, and the context size limited to 65535. That technically should have been six, but I typed in five there. Okay, and it is up and running. So, let's get over into our terminals, and you can see it's over here running away. And yeah, if you add that up, it is like 23 GB right there. And we'll just run Hermes here. We'll give it a quick hatty, make sure it is initialized and running. And as it dense, this is going to go slower than an MoE by quite a bit because again all of the activations versus just a few of the activations is a pretty big difference. Also have it run through the Digital Spaceport question set. And and yes, I know somebody's going to be like, "But it keeps saying Gemma 4 down here the entire time." That's just because I had that set to the name, but it is definitely not Gemma 4 running. Just make sure that it can process all that. That does quite a bit of testing, and it also tests of course the tool calling. And it looks like it is finding the questions here. So, it looks like it was was to parse those out. the utilization and yeah, you can see it is definitely using quite a bit of the resources. So, we're seeing generation side about 16.25 tokens per second, which when we get to the benchmark side, I think we'll see that that's kind of going to hold. I'm not going to go as big as I did with the other and I'm only going to go up to 65k, which is I believe there's a 60k context cutoff for Hermes, so you would need to be able to set a little bit above that. We're at 65k and if you set like a 50% compression ratio, like that's going to pretty much about halfway through that, go ahead and do a compression. That's a lot of compressions if you were looking at just 24 GB. So, I know there's a Q3, you might want to consider that if you're dead set on going with the Quinn, but definitely quite a bit slower. It's writing the code now for Flippy Block Extreme. It looks like there are some problems with the answers that it gave. Whether or not the Q6 would even be acceptable is a good question. So, when we move on to our testing here, it's going to be much briefer. I'm not going to go through the same question set necessarily when we throw the 3090 in here. Aside from just seeing the speed differences of it, I think there are some interesting things that we're finding out here. It's good to run these kind of tests. So, let me shut this down. I'm going to swap in the 3090 and then we'll get that test. Nice. All right. So, you can see the 3090's now installed and let's grab just the kind of performance benchmarks that we're getting on the 26B A4B. And definitely you'd want to be tuning your batch on this also. Probably two 2,000 would be the upper bound. So, we came in at 15 183 at 1024. At 4K, we came in right at 4095 tokens per second. At 16K, we came in at 3940 tokens per second. And 32K came in at 3500 tokens per second. 64K came in at 2861. And 128K finished at 2100 tokens per second. You can see the TG a little hint of that. We'll run the TG next here, but TG 128 came in at 140 tokens per second. Next, we'll run the token generation side at 512 tokens, 1024 tokens, 4096 tokens, and 8192 tokens. And the TG 512 is checking in at 133 tokens a second. Holding almost exactly the same, just a little bit of speed up there at 1K. At 4K, we hit 131.5 tokens per second. And we hit 130 tokens a second for our 8K context. All right. I think this is a lot of good information. Let's take a look at what this looks like on some charts. So, I have to say this is kind of surprising. I did not expect the 3060s to hang that well compared to the 3090. So, there's a lot of caveats. I'm going to give you the cost breakdown, and then I'll give you some conclusion takeaways and some frequently probably asked questions that I anticipate. But yeah, I did not expect the 3 3060s to hang this good. So, looking at the prompt processing first, then we're going to move on to the text generation, and we'll look at the Gemma 4 and the Qwen 3.6 in that order for each of those. And I do have charts for you also, so you can look at those as well. But to talk about the numbers a little more concretely, let's take a look here at the kind of prompt processing that we got from 1K all the way to 128K tokens. And definitely, we did see a peak that happened somewhere around the 4K to 16K range. Probably around 8K is where it would actually have peaked for these. And the triple 3060s hung in there really, really good at 1K, 4K, 16K, 32K, 65K, and all the way up to 125K. So, the biggest gap that we saw was around 4K to 16K between the triple 3060s and the 3090. And when you're looking at this many thousands of tokens per second. Now, of course, this is all baseline, the most basic benchmark I could have done. I did not do NTP, I did not do the special tunes of these that can run a little bit faster here and there. I did not do any batch tuning or anything. I wanted this to be highly reproducible for anybody. And so, I just baseline everything, just simple. Like, we're just testing the prompt processing. We're just testing the text generation. But, yeah, seeing this and seeing 3,000, 3,500 all the way back down to 128K, 2,000 for the triple 3060s I thought was surprisingly decent compared to what we saw for the 3090. Now, the 3090 did peak above 4K tokens per second at 4K, which I definitely think is, you know, it's holding and showing what you would expect from it. All the way up to about 16K it really held that performance, but after it hit about 32K it kind of closed down and all the way to the remaining 128K. At 128K kind of probably the outside edge of where you would want to have your context window set if you're using something in Gentic and a compression point about 50% of the way through, maybe 65% of the way through. You're looking at about 2,026 tokens per second versus 2,109. So, like, margin of error kind of differences when we got all the way to 128K between the two. Now, that's on Gemma 4 26B 4B at Q4, so factor that in. Is that something you are good with is a question that really only you can answer at the end of the day. It's pretty decent. I mean, it's surprisingly decent. And if we look at that in kind of a chart format, this is what it looks like here. And you can really see it's not the differences that I was expecting. You will see bigger differences, but those will be in the text generation side of it. Moving on to the Qwen 3.6 27B Q4, from 1K all the way to 128K, we saw a very similar story pan out. Although at 4K, it was not that big of a difference that we saw between the two. Of course, a dense is going to go slower just across the board, just like an MOE is going to kind of go faster across the board. So, we did see a lot of lower token performance, but many people will make the observation, and I will second that observation, quality of tokens matters more than just going stupidly fast. So, it has to be accurate while it's going fast. There's a balance there that you want, and there's a certain amount of inaccurate tool calls that you can kind of take or sustain in a session. And especially when you're doing agentic work, missed tool calls add up, and you run into timeouts, things stop working, all of a sudden you got a bunch of problems. So, seeing a start at 1K at 624 tokens per second on the triple 3060s, all the way peaking at uh 16K at 1,055 tokens per second, and then degrading down to 128K at about 731 tokens per second, kind of a nice arc there. But, that really is what I would have expected, and it really was not that far off from what the 3090 was able to do. So, 831 was where we started at 1K on that, but we ended up very, very close to the triple 3060s at 754 tokens per second at 128K. To kind of give you the chart view of that, this one I think is even a little bit closer than what we saw with the Gemma 4. So, prompt processing is where you're going to spend a lot of agentic time and having that really good tuning your batches. You can definitely get these numbers wait, wait, wait, wait faster than this. But looking at the text generation, there's where we're going to see the biggest difference. And we'll take a look at the Gemma 4 first. So, the text generation, I only did 512 up to 8K. So, usually if you cap that around 4K, you're probably going to be a little bit happier. 50% performance was not what I expected for the 3060s. I thought they would be like 25%, but 50% is what we got. 68 tokens per second on Gemma 4 for text generation all the way down to 64 at 8K. Compare that to the 3090 with its 133 that really did hold out very well there and all the way down at 8K, it was 131. So, only two tokens per second difference there. Consistency, of course, but when you're looking at that, that is, I think, a very, very respectable difference that is not that big. Now, if you're looking at Twin 3.6 27B, I would say there's a warning sign there. And that is 17.2 tokens per second all the way up to 17.8. I mean, 17.8 is not good, but it is not as bad as I was expecting. I was actually expecting to hit like some single digits on that on the triple 3060s. Now, the single 3090 did substantially better. It was all the way from 38 up to 40 tokens per second at 8K. Check out the text generation, of course, on the Gemma 4. It was a little bit like 50% and even when you go to the Twin 3.6, it's a little bit like 50% difference. So, both of those I thought were very respectable numbers for the 3060s. Is it okay to go 17 tokens per second? I got to say, it matters which model you're going with. If you're going with like Gemma 4 and you're cool with Gemma 4, I would say that, you know, that performance was totally fine. But it's 60 tokens per second. If you're looking at 17 tokens per second in Quinn 3.6 27B, the dense specifically, you're going to want to run adapters. You're going to especially want to be very mindful of making sure you've got your batching tuned. But those things are very possible and you can absolutely increase your performance, so that should be very doable. So, my biggest takeaway was I did not expect the 3060s. So, I mean, everybody talks a lot of smack about 3060s, right? In inference. And I'm like, let's test it. And I'm glad I tested it because this is going These three 3060s are going to be my backup rig. And that's a good concept to have. And for me, that's going to allow me to be flexible with tinkering around with the main rig, which man, if you're a YouTuber, you end up tinkering around and messing with stuff all the time, pretty much. So, having a dedicated rig that you're not going to be messing with is pretty nice. And for me, I think that's really good. Also, if you're looking at paying for inflated things, I don't like paying for inflated things. So, maybe $750, $800, not that bad for a 3090. If you're looking at what they are today, I mean, $1,100, $1,000 to $1,200, $1,100 probably being around the average. Those are very, very inflated from what they were 2 years ago. Whereas the 3060s are like almost not inflated at all from where they were 2 years ago. So, there's something to consider. There's definitely something to consider there. Let's talk about wattage really quick. I know a lot of people like to talk about wattage. So, I saw about 390 while it was processing be kind of what the typical working profile for it was. It did peak uh the max I saw was 580. And that was with the three 3060s, not the 3090. This is a 1,000 W PSU. I was not even close to like 800 W. So, 600 W is what we saw it for the peak there. And that's with a 5950X. Granted, the 5950X is not running crazily at the moment when it's doing inference. A single thread is, but the entire processor's not. So, that probably could have gone up another 100 watts is my guess. Uh you also probably, if you are familiar with AM4 tuning, could get that down lower or have a different processor that's not that processor to save some wattage and cost because $282 for a 5950X may not be everybody's idea. So, let's take a look at the cost for what this looks like as built right here. So, the 5950X $282, the most expensive single component in the build and definitely something I would suggest you would want to maybe you've just got an AM4 laying around. Like, honestly, reusing what you've got is kind of another huge theme that you continually see me hitting on. 3060s, once the most popular GPU in the world, there are tons of these things out there. They were the most popular because they were probably the most produced and they were the most used and that's because they just hit that budget category so nicely. So, there's a good chance you might even have a 3060 and you could have just written it off and said, "No way, that thing will not be able to do anything." I think you should reevaluate that. But, definitely the 5950X, the most expensive component at $282, probably not something you have to go with. I would recommend going something cheaper. The Gigabyte B550 Eagle, this is a $110 motherboard. This is new. It does have a lot of capabilities. I think it is a pretty decent motherboard. I wish it had 2.5 gigabit in on it, but that's not a deal breaker necessarily. The Corsair H170i. So, there's a newer model of this out, but these are usually on sale for all the way down to $100 all the way up to about $150. And these are really cool. I do think these are great 420 mm water coolers. If you can get them, great. And they come, of course, with the iCUE controller and all that stuff. So, you can control a lot of fans and a lot of lights and stuff, and they work in Linux. I've got just 512 GB NVMe in here. It's a Gen 3. Like, 30 bucks is about what those cost. The power supply that I've got, definitely a 1,000 watts is more than enough for either of these as built. As built here, like I mentioned, 600 watts pretty much capped at the power like that I saw being used. I would imagine 700 would be kind of if everything was peaking all of a sudden at once and you had some additional add-ons maybe where you were at. Of course, if you were going for multiple 3090s, I would recommend a totally different rig. You should go and watch the build that I talked about with the Threadripper. That's really if you want to go that route with multiple 3090s or better cards, probably a much better way to go. And of course, you get full width lanes with that for doing even better things. If you want to go with VLLM, there's a really critical distinction. You will have a performance impact if you don't have full wide lanes going to your GPUs. Llama.cpp, I mean, it barely pushed 100 megabytes. That was pretty much it per second. And like even these little Gen 3 X1 risers that I've got totally able to handle that, no problem. They can go up to 1 gigabyte per second, by the way. So, that's I think, I mean, 100 bucks, 110 bucks for the power supply seems reasonable. The GPU rig frame, I do like this rig frame, Sluice, whatever name it is. Uh 65 bucks for that is about what those are going for. So, the base build is actually more than the GPUs as built here, which that I'm not saying is great, but I already had this stuff laying around. So, 800 bucks for the base build awfully. If you're looking at the 12 gigabyte 3060s, 250 each for those and the 3060 8 gigabyte, 200 bucks for that. I would of course recommend probably going with 12 gigabiters all the way across the board if you were doing 3060s though, so that you ended up with 36 gigabytes of VRAM. Also, if you are thinking of going and doing video generation or image generation, definitely the 3060s, nah, I would skip those and you really do want the biggest, best, baddest card you can get. So, a 3090 in that instance, exactly what you should be after. So, these are my thoughts on it. It actually hung surprisingly well. Well enough that back up inference rig for me, and I think it's going to perform really well. I do want to tweak, if I can, the 5950X, see if I can get that changed to where maybe it's wattage is a little bit lower. I think it's idle at 95, I can probably get that lower. I really do feel like I can get that lower. But, overall, surprised. Surprised in a good way, and I'm glad I tested it out. And so, huge hat tip to all of our channel members, everybody that buys me a coffee, everybody on Patreon. Really do appreciate everything you can do for the channel. It does allow us to operate without sponsors, and I really do appreciate that. And you get a actual opinion that you don't have to worry, is this guy full of blank? And like, it's just good to test things out and to think outside the box. I feel like I have the freedom to do that because I don't have to worry about pleasing a sponsor. So, those are my takes on why I really do have to say, surprising. And thank you, because you guys enable me to do this kind of cool, crazy, out there research, which who would have thought, right? Who would have thought? Of course, you can find out more about running and setting up your software, if you want, from the playlist up here. And if you're looking for more goodness about hardware related to local AI, check out the playlist down here.

Frontier News · by Hyperjump Technology