Quoteup to 32 GB of RAM .. $1,949 when optioning 32 GB of RAM
32 GB RAM can't fit SOTA (in its class size) (or any other models in similar MoE class size) AI LLM model[1] huggingface.co/unsloth/Qwen3.6-35B-A3B-MTP-GGUF at a minimum good Q4_K_XL quant with proper context:
Quote from: reddit.com/r/LocalLLaMA/comments/1sq94qx/is_anyone_getting_real_coding_work_done_with.. I've come to the conclusion that (1) 32768 is the biggest context I can get away with in an adequately smart model, and (2) it just ain't enough.
It's also possible to calculate the available context[2]: 16,384 [KV cache context tokens per GB]*(32 [GB total memory] - 6 to 8 [GB for the OS] - [GB quant filesize]):
~34,400 context tokens = 16,384*(32-7[3]-22.9) (confirms the quote)
Agentic workflows often exceed 60,000 or quite a bit more context tokens.
Recalculating with an additional 8 GB VRAM GPU:
165,500 = 16,384*(32-7-22.9+8)
So, in this memory sense, any normal gaming laptop that has a 6-8 GB VRAM GPU is superior and cheaper. It's also going to have upgradable RAM (yes, upgradable RAM is slower (5600 vs 9400 (what Lenovo uses)), but the GPU's much faster VRAM compensates for that).
And, generally, you get more memory per buck if you get a desktop PC (and get a much cheaper iGPU-only laptop for everything else/non-AI and remote desktop into your AI desktop PC).
Qwen4-27B has been announced[4] (not stated if MoE and/or dense, but hopefully MoE too) will revive 32 GB unified memory/RAM devices:Based on from what we know from Qwen3.8-Flash-Next (Qwen4 preview), such a model could be roughly 20-30% smaller and still perform better. Smaller means the 30% would be streamed from the SSD with only a ~5% performance penalty vs if it was streamed from the RAM[6]. Including a possible reduction in KV cache context tokens memory requirements from 16,384 to something lower (DeepSeek show the technology already used by huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash (see "Global KV Cache Per Token (Bytes)" image)).
35B - 25% = 26.25B (confirms that a 27B MoE model may be possible at same or better performance using mentioned technology, which is already being used).
Let's recalculate the context tokens we'd get using from what we know from already existing 27B models's Q4_K_XL (17.6 GB) quants:
121,000 = 16,384*(32-7-17.6).
[1] artificialanalysis.ai/models/open-source/small?models=qwen3-8-27b-medium%2Cqwen3-8-27b-low%2Ck2-horizon-mova-36b-a4b%2Cg9v3-39a5b%2Ck2-horizon-7b%2Cqwen3-8-27b-non-reasoning%2Cqwen3-6-35b-a3b%2Cmuse-glimmer%2Cgemma-4-26b-a4b%2Cqwen3-6-35b-a3b-non-reasoning%2Cgemma-4-31b%2Cqwen3-8-27b#artificial-analysis-intelligence-index (K2 Horizon has expensive KV cache and may reason above average[5])
[2] reddit.com/r/Qwen_AI/comments/1vo8pjz/qwen3827b_kv_cache_works_out_to_64_kibtoken_so/
[3] techpowerup.com/353395/windows-11-26h2-update-lowers-ram-consumption[/quote]
[4] reddit.com/r/LocalLLaMA/comments/1wmxfjs/qwen_4_announced_at_apsara_conference/
[5] reddit.com/r/LocalLLaMA/comments/1wg0vqz/comment/p9qj8c8/
[6] notebookchat.com/index.php?topic=324510.0