160 GB can be allocated, but it may be more on Linux, just like with Strix Halo (search: "strix halo linux gtt").
Assuming:
32GB Mainboard 64GB Mainboard 128GB Mainboard
Windows 28GB* 56GB* 112GB*
Linux 28GB 56GB 112GB128 GB totally available - 112 GB allocateable = 16 GB minimum for the OS.
192 GB totally available - 16 GB minimum for the OS = 176 GB possible? vs official 160 GB.
That said, check this out:
Quote from: youtube.com/watch?v=Z26kN5VfyGAAMD and the industry are pushing the Ryzen AI Max+ 495 very hard as serious AI desktops, but the marketing is getting ahead of the actual performance. I own both a 128GB Framework Desktop and a 128GB HP Zbook laptop on this platform, so I know what they can do. Yes, the huge unified memory lets you load large models, but in many cases you're running heavily quantized versions, which inevitably sacrifice model quality. And with only ~273 GB/s bandwidth versus around 1.2 TB/s on an M5 Ultra at about €10,500 for 256GB ram, inference is simply too slow for what they are now charging. Last year, the 128GB Framework Desktop with the Ryzen AI Max+ 395 was around €2,500 now we're talking about €8,500 for 192GB. If you're a developer, you'd be far better off paying €80–€200/month for a serious AI coding subscription and getting much stronger models. And if you're a company that absolutely needs local AI for data privacy, then spend more and buy genuinely high-performance hardware. These are interesting machines, but at €8,500 they are being massively oversold. 👎
Gorgon Halo is (just) a software-patched Strix Halo to support more than up to 128 GB RAM. All Ryzen AI 400 are a rename/refresh of the Ryzen AI 300 series.
Wish AMD would allow for 256 GB RAM straight away and not artificially limit the maximum RAM config..
On desktop B850 motherboards, up to 256 GB RAM are possible, and this is on a 128-bit memory interface (aka dual-channel). Strix/Gorgon Halo use a 256-bit memory interface (and faster RAM, too, btw.), so by this argument, up to 512 GB RAM should be possible on Strix/Gorgon Halo, or even more, because the memory speed is ~1.52 times faster: 8533 MT/s vs desktop's 5600 MT/s.
But, 256-384 GB LPDDR6 RAM is where Medusa Halo comes in later and its speed / memory bandwidth, and therefore also probably compute, increases are very welcomed.
If more than 160 GB is allocateable:
You would be able to fit huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF at its native weights (quants perform over-proportionally worse, because it's a native mostly 4-bit release) (162 GB) + a decent amount of context tokens.
Of course, fitting 3- or 4-bit quants of any ~300B models becomes possible, like huggingface.co/unsloth/GLM-5.3-Flash-GGUF (IQ4_XS) (157 GB), tho, this is where 256 GB RAM would have been nice (see the YouTube quote), to be able to run UD-Q4_K_XL (200 GB).