NotebookCHECK - Notebook Forum

English => News => Topic started by: AI on 32 GB RAM devices on October 05, 2026, 17:33:54

Title: For all the 32 GB RAM / unified memory, iGPU-only, devices ..
Post by: AI on 32 GB RAM devices on October 05, 2026, 17:33:54
Quote from: on the cusp on Yesterday at 11:32:54
Quoteup to 32 GB of LPDDR5X RAM
This is unfortunately on the cusp of being usable, because 32 GB of total memory/(V)RAM you can fit about this many context tokens of SOTA (in its class size) AI LLM model[1] huggingface.co/unsloth/Qwen3.6-35B-A3B-MTP-GGUF 4-bit UD-Q4_K_XL quant:
Quote from: reddit.com/r/LocalLLaMA/comments/1sq94qx/is_anyone_getting_real_coding_work_done_with.. I've come to the conclusion that (1) 32768 is the biggest context I can get away with in an adequately smart model, and (2) it just ain't enough.
(1 token = 0.75 words)

It's also possible to calculate the available context[2]: 16,384 [KV cache context tokens per GB]*(32 [GB total memory] - 6 to 8 [GB for the OS] - [GB quant filesize]):
~34,400 context tokens = 16,384*(32-7-22.9) (confirms the quote)

On Windows and with the MTP variant (=faster token generation) (with MTP head: 22.9 GB (vs without MTP: 22.4 GB)), the context tokens are going to be even less than what is in the quote.

Agentic workflows often exceed 60,000 context tokens.

In this memory sense, any gaming laptop that has a 6-8 GB VRAM GPU is superior, possibly even cheaper and also going to have upgradable RAM.

Recalculating with an additional 8 GB VRAM GPU:
165,500 = 16,384*(32-7-22.9+8).

And, generally, you get more memory per buck if you get a desktop PC.

[1] artificialanalysis.ai/?models=qwen3-8-27b%2Cqwen3-8-27b-medium%2Cqwen3-8-27b-low%2Cqwen3-8-27b-non-reasoning%2Cqwen3-6-35b-a3b%2Cqwen3-6-35b-a3b-non-reasoning&intelligence=artificial-analysis-intelligence-index
[2] reddit.com/r/Qwen_AI/comments/1vo8pjz/qwen3827b_kv_cache_works_out_to_64_kibtoken_so/
.. we can only hope that Qwen release a smaller 35B-A3B MoE AI model, so that it can fit in 32 GB RAM / unified memory systems and leaves enough RAM for a good amount of context tokens. Based on from what we know from Qwen3.8-Flash-Next (Qwen4 preview), such a model could be roughly 20-30% smaller and still perform better. Smaller means the 30% would be streamed from the SSD with only a 5% performance penalty vs if it was streamed from the RAM[1]. Including a possible reduction in context tokens memory requirements from 16,384 to something lower (DeepSeek show the technology already used by huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash (see the "Global KV Cache Per Token (Bytes)" image)).

35B - 25% = 26.25B. Qwen already announced / confirmed an upcoming Qwen4-27B model, but didn't say whether it will be a dense and/or a MoE model(s).

Let's recalculate the context tokens we'd get using from what we know from already existing 27B models:
121,000 = 16,384*(32-7-17.6).

[1] notebookchat.com/index.php?topic=324510.0