QuoteHowever, the high starting price of $2,599 initially limits their target audience.
If it was 128 GB RAM, then no, but it's 32 GB and can't even fit still SOTA (in its class size) (or any other models in similar 35B MoE class size) (and there's still no smaller MoE model that outperforms it) AI LLM model[1] huggingface.co/unsloth/Qwen3.6-35B-A3B-MTP-GGUF at a minimum good Q4_K_XL quant with proper context:
Quote from: reddit.com/r/LocalLLaMA/comments/1sq94qx/is_anyone_getting_real_coding_work_done_with.. I've come to the conclusion that (1) 32768 is the biggest context I can get away with in an adequately smart model, and (2) it just ain't enough.
It's also possible to calculate the available context[2]: 16,384 [KV cache context tokens per GB]*(32 [GB total memory] - 6 to 8 [GB for the OS] - [GB quant filesize]):
~34,400 context tokens = 16,384*(32-7[3]-22.9) (confirms the quote)
~18,600 context tokens = 16,384*(32-8-22.9) (1 GB more for the OS -- this is how much every GB matters)
Qwen4-27B has been announced[4] and hopefully it's MoE (=will run fast on 32 GB unified memory devices), then it will revive them:Recalculating context from existing 27B Q4_K_XL (17.6 GB) quants:
121,000 = 16,384*(32-7-17.6).
And if the KV cache arch is improved in Qwen4 as well, like in [5], let's say just even doubled per GB, then 121k would become 242k!
[1] artificialanalysis.ai/models/open-source/small?models=qwen3-8-27b-medium%2Cqwen3-8-27b-low%2Ck2-horizon-mova-36b-a4b%2Cg9v3-39a5b%2Ck2-horizon-7b%2Cqwen3-8-27b-non-reasoning%2Cqwen3-6-35b-a3b%2Cmuse-glimmer%2Cgemma-4-26b-a4b%2Cqwen3-6-35b-a3b-non-reasoning%2Cgemma-4-31b%2Cqwen3-8-27b#artificial-analysis-intelligence-index (K2 Horizon has expensive KV cache and may reason above average[5])
[2] reddit.com/r/Qwen_AI/comments/1vo8pjz/qwen3827b_kv_cache_works_out_to_64_kibtoken_so/
[3] techpowerup.com/353395/windows-11-26h2-update-lowers-ram-consumption[/quote]
[4] reddit.com/r/LocalLLaMA/comments/1wmxfjs/qwen_4_announced_at_apsara_conference/
[5] huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash (see "Global KV Cache Per Token (Bytes)" image)