REDDIT 原始帖子
Ssd stream models on strix halo
Has anybody here tried ssd streaming deepseek v4 flash or pro on a strix halo? What about kimi k3? I dont want fast speeds, just a good architect that can plan things over night and delegate to smaller models during the day. If someone can point me in this…
Has anybody here tried ssd streaming deepseek v4 flash or pro on a strix halo? What about kimi k3? I dont want fast speeds, just a good architect that can plan things over night and delegate to smaller models during the day. If someone can point me in this direction im also willing to try out and report back numbers
已收录讨论
doesn't dsv4-flash fit already in q2? On kimi k3, the answer is more likely to come to you in a dream than that thing ever finishing.
You can fit a few quants of ds4 flash yes, I'm using an antirez one and it's really impressive and capable
Do you have anything special configured for it? Ive found qwen 3.6 27b better for most agentic tasks and always default to it over flash Edit - im using the antirez quant
I think even q1 is 600gb
Yes, i want to run the full quantization and stream off ssd if possible, not happy with the quality of the lower quant versions of deepseek flash. Haha, maybe a q2 of kimi
You may have a look at this project: https://github.com/igorbarshteyn/llama-kimibri/tree/main/tools/kimibri
I put just 10% of a model on optane disks, it wasn't even 1t/s
TIL that I'm kimi k3. Also how the fuck does that happen? I'll go to sleep feeling fucking frustrated and have a dream that I resolved the problem, and bang, the morning I actually finish the problem with the dream solution
Yes, I streamed a 220GB GLM 5.2 Quant on Strix Halo. Got to ~2tps generation speed.