Since this has AI in its product name, but the total memory (RAM+VRAM) is only 32 GB RAM (and no dedicated, e.g. 8 GB VRAM, GPU) for a 3500 bucks price, this must be said:
RAM in consumer devices is always slow, so you want to use a appropriate MoE model, not a dense one and currently the best one, in its size class, is still Qwen3.6-35B(-A3B) (Qwen3.8-27B is released, but unfortunately no Qwen3.8-35B MoE so far). But if one uses the minimum recommended 4-bit quant of this Qwen3.6-35B, one will fit only about this much of context tokens:
Quote from: reddit.com/r/LocalLLaMA/comments/1sq94qx/is_anyone_getting_real_coding_work_done_with.. I've come to the conclusion that (1) 32768 is the biggest context I can get away with in an adequately smart model, and (2) it just ain't enough.
More about running AI models locally: notebookchat.com/index.php?topic=315954.0 ("Your own ChatGPT, offline: AI without the cloud on your laptop") and comments.
In this memory sense, any upgradable RAM + 6-8 GB VRAM gaming laptop (used for 700-800 bucks) is superior and much cheaper, even new.
PS: The new Qwen3.8-Flash-Next (Qwen4 architecture preview) ("Number of Parameters: 125B with 6B activated, plus 51B n-gram embedding and 4B MTP") allows to offload a big part of its parameters on a SSD and still get a good speed.
- reddit.com/r/LocalLLaMA/comments/1vyq2v4/megathread_qwen38flashnext_release_day/
- github.com/ggml-org/llama.cpp/pull/27742
- huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/discussions
- github.com/0xBakeer/qwen38-flash-next-spark/blob/main/docs/how-it-works.md
But this doesn't mean that in the future a model for 32 GB RAM devices is going to come out. If you want such a model, your best bet is to let Qwen know it.