The interesting part of AMD's Ryzen AI Max+ 395 is not the CPU, it is the memory trick. Strix Halo pools up to 128GB of LPDDR5X between its 16 Zen 5 cores and its big integrated Radeon GPU, then lets you hand up to 96GB of that pool to graphics. That single number is why a small mini PC can suddenly load a 70 billion parameter model that would choke most discrete cards.

Quick Answer

The Ryzen AI Max+ 395 carries up to 128GB of unified LPDDR5X memory, of which up to 96GB can be assigned as VRAM through AMD Variable Graphics Memory. With memory bandwidth around 256 GB/s, a Radeon 8060S iGPU built on 40 RDNA 3.5 compute units, and a 50 TOPS XDNA 2 NPU, it can run a quantised 70B model locally. Expect grey import pricing near R75,000 for the mini PC in South Africa.

How the 96GB VRAM Allocation Works

A normal desktop splits memory in two: system RAM for the CPU and dedicated VRAM soldered to the graphics card. Strix Halo throws that split out. It uses one shared LPDDR5X-8000 pool, and AMD's Variable Graphics Memory feature lets you carve off a large slice of that pool and present it to the GPU as VRAM. On a 128GB part, up to 96GB can become graphics memory, which leaves enough for the operating system while giving the iGPU a VRAM budget no consumer discrete card comes close to.

Bandwidth is the trade off. A 256-bit LPDDR5X bus delivers roughly 256 GB/s, which is generous for an integrated design but well short of a high end discrete card's GDDR. For large language models, where simply fitting the weights in memory is the hard limit, that capacity matters far more than raw bandwidth.

Why a 70B Model Suddenly Fits

A 70 billion parameter model at 4-bit quantisation needs roughly 40GB of memory just to hold its weights, before any context. Discrete cards top out at 24GB or 32GB on the consumer side, so that model normally demands multiple GPUs or a workstation card costing many times more. With up to 96GB addressable as VRAM, Strix Halo loads the whole thing in one device, leaving headroom for a longer context window.

It will not match a stack of data centre GPUs on speed. Token generation is steady rather than blistering, which suits local coding assistants, document analysis and private inference rather than serving many users at once. For a developer who wants a big model running on a desk without cloud bills, that is the appeal. Evetech's AI PC category is where this class of hardware lands as local AI machines reach the SA market.

What This Means for SA Buyers

Strix Halo mostly ships inside mini PCs and compact workstations rather than as a loose chip, with units like GMKtec's EVO-X2 leading the wave. Local availability is thin, so most early adopters are looking at grey import routes, which means adding shipping, import duties and 15 percent VAT on top of the dollar price. Budget realistically around R75,000 once everything lands, and confirm warranty cover before you commit. If you would rather build around proven, locally stocked parts, the PC best sellers list shows what is moving right now.

Frequently Asked Questions

Is the Ryzen AI Max+ 395 better than a discrete GPU for AI?

For fitting very large models, yes, because no consumer discrete card offers up to 96GB of usable VRAM. For raw speed on models that already fit in 24GB or 32GB, a strong discrete card is faster thanks to higher memory bandwidth. It comes down to capacity versus throughput.

Can I game on a Strix Halo machine?

Yes. The Radeon 8060S iGPU with 40 RDNA 3.5 compute units is genuinely capable for an integrated part and handles modern titles at 1080p comfortably. It is built for AI and productivity first, but gaming is a real secondary use.

How much memory should I assign as VRAM?

Match it to your workload. For large model inference, push the allocation high so the weights fit, but always leave enough for the operating system and background apps. AMD's Adrenalin software and the BIOS both expose the Variable Graphics Memory setting.

Why is the bandwidth lower than a graphics card?

It uses an LPDDR5X system bus rather than dedicated GDDR memory, so bandwidth sits around 256 GB/s instead of the much higher figures on discrete cards. That keeps power and cost down while still delivering the huge capacity that makes the chip interesting.

Curious where local AI hardware is heading? Keep an eye on the AI PC range at Evetech and talk to the team about a machine sized for the models you actually run.