Qwen4-27B (confirmed[1]) (unconfirmed: MoE and/or dense) will revive 32 GB system memory devices:
Based on from what we know from Qwen3.8-Flash-Next (Qwen4 preview), such a model could be roughly 20-30% smaller and still perform better. Smaller means the 30% would be streamed from the SSD with only a 5% performance penalty vs if it was streamed from the RAM[2]. Including a possible reduction in context tokens memory requirements from 16,384 to something lower (DeepSeek show the technology already used by huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash (see the "Global KV Cache Per Token (Bytes)" image)).
35B - 25% = 26.25B (confirms that a 27B MoE model may be possible at same or better performance using mentioned technology that already exists).
Let's recalculate the context tokens we'd get using from what we know from already existing 27B models:
121,000 = 16,384*(32-7-17.6).
[1] reddit.com/r/LocalLLaMA/comments/1wmxfjs/qwen_4_announced_at_apsara_conference/
[2] notebookchat.com/index.php?topic=324510.0