News:

Willkommen im Notebookcheck.com Forum! Hier können Sie über alle unsere Artikel und allgemein über notebookrelevante Dinge diskutieren. Viel Spass!

Main Menu

Post reply

The message has the following error or errors that must be corrected before continuing:
Warning - while you were reading a new reply has been posted. You may wish to review your post.
Other options
Verification:
Please leave this box empty:
Shortcuts: ALT+S post or ALT+P preview

Topic summary

Posted by Pascal76
 - Today at 15:46:38
>fast LPDDR5X RAM (7,467 MT/s)

The integrated memory controller on the Core Ultra X7 358H officially supports dual-channel memory speeds up to LPDDR5X-9600 MT/s (as well as LPDDR5X-8533 MT/s).
Intel's OEM Branding Requirement: Intel established 7,467 MT/s as the strict baseline threshold for LPDDR5X RAM in laptop designs using Panther Lake processors

=> maybe it is fast but it could be much better ...
Posted by 32 GB RAM and AI
 - Today at 14:48:26
Quote32 GB of RAM
QuoteRAM not expandable
If you plan using this mini-PC for AI (extensively):
Since the total memory (RAM + VRAM) in this mini-PC does not go above 32 GB: Know that with current SOTA AI LLM model Qwen3.6-35B-A3B[1] (in its size class) (only 3B parameters get activated per generated token -> perfect for RAM-only, no dGPU, devices), and its bang for the buck, 4-bit quant, you will be restricted to about this amount of context:
Quote from: reddit.com/r/LocalLLaMA/comments/1sq94qx/is_anyone_getting_real_coding_work_done_with.. I've come to the conclusion that (1) 32768 is the biggest context I can get away with in an adequately smart model, and (2) it just ain't enough.

In this memory sense, any upgradable RAM + 6-8 GB VRAM gaming laptop (used for 700-800 bucks) is superior and much cheaper, even if new. The additional 8 GB of memory/VRAM of the GPU make all the difference in being able to fit and run huggingface.co/unsloth/Qwen3.6-35B-A3B-MTP-GGUF/blob/main/Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf.

It is possible to stream the 3B active parameters from a SSD, but you'd have to inform yourself on how many tokens per second you'd be getting (ask/search e.g. on /r/localllama, /r/localllm).

More about running AI models locally: notebookchat.com/index.php?topic=315954.0 ("Your own ChatGPT, offline: AI without the cloud on your laptop") and comments.

[1] artificialanalysis.ai/?models=nemotron-3-5-lightning%2Cmuse-glimmer%2Cqwen3-8-27b-medium%2Cqwen3-6-35b-a3b%2Cqwen3-8-27b%2Cg9v3-39a5b%2Cgemma-4-31b&intelligence=artificial-analysis-intelligence-index#intelligence (Qwen3.8-27B (dense architecture -> 27B active parameters per token) scores much higher, but it's much slower) (you can look up non-reasoning scores or other models in the table)

If you still want this 32 GB total memory mini-PC: According to

reddit.com/r/LocalLLaMA/comments/1we8tl1/3827b_has_ruined_353635bs_for_me_its_just

, Qwen3.8-27B (dense) is, unsurprisingly and as can be seen in the previously mentioned artificialanalysis evaluation, a superior model VS Qwen3.6-35B-A3B for most tasks. So, if you are (really) ok with waiting many times longer for a much better answer (for simpler tasks, Qwen3.6-35B-A3B can be good enough and is much faster), you should consider huggingface.co/unsloth/Qwen3.8-27B-GGUF (Q4_K_XL). Due to also being smaller (27B vs 35B), Qwen3 27B has the advantage of giving you many more context tokens[2] on a 32 GB RAM-only device.

The slow speed of running a 27B dense quant on a DDR5-only device is ultimately going to annoy you, so, as previously mentioned, you should really think about getting a device with an at least additional 8 GB VRAM GPU, so that you can offload parts of the LLM into the much faster VRAM and run a ~27B dense, 4-bit quant, (much) faster VS a DDR5-only device.

[2]
Context: 16,384[KV cache context tokens per GB]*(32 [GB total memory] - 6 [GB for OS] - [quant filesize])[3]:

  • ~59,000 context tokens when using Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf = 16,384*(32-6-22.4) (or ~26,000 context tokens when giving the OS 8 GB, instead of 6 GB)
  • ~138,000 context tokens when using Qwen3.8-27B-UD-Q4_K_XL.gguf = 16,384*(32-6-17.6)

In other words, Qwen 27B more than doubles your available context tokens to approx. 76,800 (16,384*(22.4-17.6)).

[3] reddit.com/r/Qwen_AI/comments/1vo8pjz/qwen3827b_kv_cache_works_out_to_64_kibtoken_so/

For 1500 bucks you can get a desktop PC with more of total memory and it will be faster, too, because it will have a good gaming GPU with 12 GB of VRAM:
  • this: 3808 AI score = 32 GB RAM * 119 GB/s (=128*7467/1000/8)
vs
  • desktop PC build with a RTX 4070: 8915 AI score = 32 GB RAM * 89.6 GB/s (=128*5600/1000/8) + 12 GB VRAM * 504 GB/s
You may also be able to get 2*24 GB of RAM.
Posted by Redaktion
 - Today at 13:08:07
The Minisforum M2 Pro combines Intel's new Core Ultra X7 358H with the powerful Arc B390 and a wide range of features. In our review, this compact mini-PC demonstrates just how much power is packed into its small chassis, how quietly the system operates, and whether the overall package is worth the price of $1,439.

https://www.notebookcheck.net/Minisforum-M2-Pro-mini-PC-review-Plenty-of-performance-a-fair-price-and-a-well-rounded-package.1398885.0.html