REDDIT 原始帖子

DeepSeek V4 Flash on DGX Spark

Does anyone have experience with running models like a V4 flash on a setup utilizing two DGX Spark? Memory would be sufficient but I have absolutely no idea if the throughput is somewhat reasonable or not. We need this in an environment without internet so…

原帖正文r/datascience

Does anyone have experience with running models like a V4 flash on a setup utilizing two DGX Spark? Memory would be sufficient but I have absolutely no idea if the throughput is somewhat reasonable or not. We need this in an environment without internet so cloud resources are not an option.

已收录讨论

3 条评论

u/Cupakov

Check out either r/LocalLlama or even better, the Nvidia DGX Spark forums: https://forums.developer.nvidia.com/c/accelerated-computing/dgx-spark-gb10/dgx-spark-gb10/721

u/thefossguy69

You can run a few quantised models on a single node too! I started my 0731 experimentation with the UD-IQ1_M (1 bit) quant from Unsloth (https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF) and I'm really shocked to see a near-frontier capability model, running locally with 250+ tokens per second.