Kindly Benchmark Higher Quants of DeepSeek-v4-flash Against Qwen-3.6-27B Q8!
Kindly Benchmark Higher Quants of DeepSeek-v4-flash Against Qwen-3.6-27B Q8! I am running the UD-Q2_K_M of the model locally, though I can run Qwen3.6-27B_Q8_K_XL at around 70t/s with MTP activated. The question I am constantly asking myself is: Is it worth…
Kindly Benchmark Higher Quants of DeepSeek-v4-flash Against Qwen-3.6-27B Q8! I am running the UD-Q2_K_M of the model locally, though I can run Qwen3.6-27B_Q8_K_XL at around 70t/s with MTP activated. The question I am constantly asking myself is: Is it worth running a slower higher quantized version of the Deepseek-v4-flash? I have no idea. My gut feelings tells me that Qwen3.6-27B_Q8_K_XL, coupled with online search, should be better than a highly quantized Deepseek, a model that takes up 100GB on my disk. What do you think?
已收录讨论
Can you stop spamming the same post? Run some benchmarks and share the results instead.
Its a useaful and actionable comment though.
Or instead of benchmarks - just do your tasks. Stop fantasizing.
But that could serve as a basis for comparison.
I like DS4-flash way better, for my applications.
oh, you know, wait a couple of days and benchmark against qwen3.8-27B
If you don't like my post, don't read it and don't post useless comments that does not help anyone.
If you don't like their comment, don't read it and don't post useless replies that does not help anyone.
Just too problem dependent to eval for me.
If you don't like his comment, don't read it.
it might as well be brain dead at Q2 but since you already have both, just run a few tests thru them to generate a few apps and compare.
I run them both (deepseek q2_k_xl ~90GB, and qwen 3.6 27b q8_0) and deepseek is definitely way smarter. However, qwen often "good enough" and much faster