ORIGINAL REDDIT POST
Can I run DSv4Flash-0731 with 2x16Gb VRAM and 128Gb RAM? what about Qwen3.8 27B?
I have a rig with an RT 5080 and 64Gb DDR5. I'm considering adding a 5060ti and 64Gb more RAM (a bit on the limit for my B650PLUS mobo and 850W PSU, but I think still doable). Which quants would I be able to run for DSv4Flash-0731 and Qwen3.8 27B? And at…
I have a rig with an RT 5080 and 64Gb DDR5. I'm considering adding a 5060ti and 64Gb more RAM (a bit on the limit for my B650PLUS mobo and 850W PSU, but I think still doable). Which quants would I be able to run for DSv4Flash-0731 and Qwen3.8 27B? And at which speeds +/-?
Collected discussion
Performance won't be steller on DSv4 Flash. Stick to Qwen 3.x 27B. If the prices on the RAM and the GPUs are decent then go ahead with it. Also, consider the 5070ti instead of the 5060ti. Better specs, although you'll only get more VRAM for context, 32GB won't help with running larger models.
I can get the 5060ti for 500€< the 5070ti is 800, I could get a second 5080 for 1k instead. 64Gb DDR5 is 650€ today. Performance won't be steller on DSv4 Flash because of the quant or because of the speeds? I'm ok with slow. so having 32Gb split in two will only help with context for dense models? I was hoping I'd be able to split the model in two but I guess not. What about MoEs? Also no effect on model size, just extra context?
Should I get an R9700 instead?
Wouldn't it be beneficial for Ds4 to have more than 96VRAM? Should I get an R9700 (1300€) instead? Not sure what else to get with more than 16Gb VRAM for that price range, maybe an RTX4000PRO (1800€)?
if you're okay with slow then get the ram
Not for DSv4 flash, just for running a better quant of qwen 27B + more context.
I have an R9700 and its beautiful. Noisy, but beautiful. You can't go wrong with 32gb VRAM and RDNA 4. I had problems on Windows with Vulkan but RocM has been great, and even faster
2x3090 MSI 3 fan edition, quite quiet and temps up to around 70 degrees, on longer sessions bit more. Aggressive fan curve to be as silent as possible.
try not to mix nvidia/AMD/Intel unless you know exactly what you are doing.
unless it's unified high bandwidth RAM
You see a lot of that on B650Plus motherboards?
I would prioritize VRAM over RAM for most cases.
Get a 8GB CMP 170HX and unlock it, thank me later
You can get DS4 Flash 731 for free on openrouter.ai right now... Just sayin' before you splash too much cash...
I have the sapphire r9700 ai pro. Super noisy so I undervolt it. What variant are you running? Hows its noise?
If you have the money, skip everything desktop and get a Cascade Lake Xeon or Epyc Rome. You'll get a lot more memory bandwidth. Server DDR4 RAM is still quite a bit cheaper vs Desktop RAM. 192GB DDR4-2666 is like €350, and you get almost 128GB/s memory bandwidth. Pair that with a single 3090 and you'll have a much better experience, while running the full 160GB DS4 Flash.
.... and 64Gb more RAM .... Get additional GPU instead.
Get the R9700 32GB for a bit more. ----- > so having 32Gb split in two will only help with context for dense models? I was hoping I'd be able to split the model in two but I guess not. What about MoEs? Also no effect on model size, just extra context? That's not what I meant. You can obviously "split" the model in half per card but there aren't any models better than Qwen 3.x 27B. Next big step up requires more than 64GB of VRAM. So, adding another GPU in your case means that you can run a better quant of Qwen 27B and have enough space for a decent amount of context. On 40GB (2x 7900XT) I can fit the 27B Q5_K_XL quant + mmproj/vision + 130k unquantized KV cache (context). Unless you use Tensor parallel (which doesn't work well with mixed GPUs), you can only opt for pipeline parallel i.e. the GPUs will process the task sequentially rather than in parallel. Token generation speed would roughly be the same.
I would get a 5070 TI instead of 5060 TI. Smaller gap in bandwidth between the 2 cards.
You can split the model in two. Look for split layers param on llama.cpp docs.
With a 5090+3090 and 128gb of ram I am getting around 13tk/s (gen) and 200 tk/s (proc) with the iq3xxs / q8 kv cache and 64k context. Ram is about at 110gb used. At 32k context I get 20 t/s / 300 t/s. You might be a bit short to use that quant. If you put that kind of money a server build with more memory bandwidth is worth investigating.
I have two R9700 and can’t run deepseek-v4-flash. I’m unsure why are people suggesting you get this GPU to run this model.. On the other hand, for Qwen3.6 27B, it works great
Memory bandwidth is key. VRAM is fast. Regular RAM is, by comparison, SLOOOOOOOOOOWWWWWWW.
You can. But it will be slow.
Qwen3.x-27B lower quant is probably by far the best option for 32GB VRAM. Better upgrade to 2x3090 or 2x3080 (20GB variant) to run 3.x-27B with higher quants/context. 2x3090 unlock Q8 with full context.