On 24. April 2026 DeepSeek-V4-Flash(-preview) came out.
Now, on the 31. August 2026, the full version has been released.
Quote from: 192 GB+24 GB VRAM is nice on July 21, 2026, 11:06:14(the current DeekSeek-V4-Flash is a preview version and the final version is supposed so be released this month)
And, as said, the full version has been released this very month indeed, both, not only on their API, but also the weights:
Quote from: reddit.com/r/LocalLLaMA/comments/1vbidkp/deepseekv4flash_has_been_updated_the_official/DeepSeek-V4-Flash has been updated, "The official release of DeepSeek-V4-Pro will follow soon"
Quote from: reddit.com/r/LocalLLaMA/comments/1vbk5ob/new_deepseek_v4flash_achieves_50_on/New DeepSeek V4-Flash achieves 50 on ArtificalAnalysis Index, 1 point below GLM-5.2 and GPT-5.6 Luna
Quote from: artificialanalysis.ai/articles/deepseek-v4-flash-0731-scores-50-on-the-artificial-analysis-intelligence-index-10-points-above-previous-deepseek-v4-flashDeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, a 10-point jump over DeepSeek V4 Flash (released April 2026) that puts it 6 points ahead of DeepSeek V4 Pro
[..]
Weights from e.g. here: huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF (follow their guide on how to run this model locally (and therefore privately)).
On 128 GB RAM and/or VRAM systems, it's worth trying a 3-bit quant (smallest is 103 GB).