ORIGINAL REDDIT POST
DeepSeek V4 Flash on DGX Spark
Does anyone have experience with running models like a V4 flash on a setup utilizing two DGX Spark? Memory would be sufficient but I have absolutely no idea if the throughput is somewhat reasonable or not. We need this in an environment without internet so…
Does anyone have experience with running models like a V4 flash on a setup utilizing two DGX Spark? Memory would be sufficient but I have absolutely no idea if the throughput is somewhat reasonable or not. We need this in an environment without internet so cloud resources are not an option.
Collected discussion
Check out either r/LocalLlama or even better, the Nvidia DGX Spark forums: https://forums.developer.nvidia.com/c/accelerated-computing/dgx-spark-gb10/dgx-spark-gb10/721
Thank you!
You can run a few quantised models on a single node too! I started my 0731 experimentation with the UD-IQ1_M (1 bit) quant from Unsloth (https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF) and I'm really shocked to see a near-frontier capability model, running locally with 250+ tokens per second.