A Strix Halo mini-PC earns its place on a coding desk for one reason: capacity. The Ryzen AI Max+ 395 can hand the iGPU up to 96GB of its 128GB unified memory as VRAM, so a 70B-class coding model loads and runs on a single chip that fits in your hand. The catch is bandwidth, and once you understand where the ceiling sits, the value of the box becomes clear.
Quick Answer
Yes, for capacity. The Ryzen AI Max+ 395 allocates up to 96GB as VRAM across a 256-bit bus rated at 256 GB/s, around 215 GB/s measured, which is enough to run a dense 70B coding model at 4-bit at roughly 5 tokens per second. That is reading pace, fine for batch refactors and long reasoning, not snappy autocomplete. SA stock is thin, so expect a grey import around R75,000.
What the silicon actually gives you
The 395 pairs 16 Zen 5 cores with a 40-compute-unit Radeon 8060S iGPU and a 50 TOPS XDNA 2 NPU, all sharing one 128GB LPDDR5X pool. For local AI the number that matters is addressable memory, and roughly 96GB of it can be handed to the GPU. That changes the question from whether a large model fits to which quantisation you pick. On a 128GB unit a 70B model at Q8 loads cleanly with close to 58GB left for context, system overhead, and a second model running alongside.
Where the bandwidth ceiling bites
Memory capacity is generous; memory speed is the constraint. A dense 72B model running every parameter on each token settles around 4.5 tokens per second, because the chip has to stream the whole model through that 215 GB/s pipe per token. A mixture-of-experts model tells the better story: a 35B MoE with around 4B active parameters runs near 43 tokens per second on the same box, because only a slice of the weights moves per token. For coding, that means the machine is excellent for explaining a codebase, generating a function, or running an overnight agent task, and less suited to instant line-by-line completion where latency is everything.
Thermals are the other practical note. In the small chassis these chips ship in, sustained inference runs warm, so a model with decent cooling and airflow holds clocks better under long sessions. If you are weighing it against a desktop with a discrete card, the broader AI PC selection at Evetech is the cleaner local route, since fast machines move quickly through the current PC best sellers.
Frequently Asked Questions
Can a Strix Halo mini-PC really run a 70B coding model locally?
Yes. With up to 96GB allocatable as VRAM, a 70B model at 4-bit loads on a single APU and produces usable output. Speed lands around 5 tokens per second for a dense model, so it suits batch and reasoning work more than instant completion.
Why is it slower than a Mac Studio at the same task?
The limit is memory bandwidth, not capacity. At roughly 215 GB/s measured, the chip streams a dense model more slowly than higher-bandwidth platforms, so large dense models read out at a steady but modest pace.
Should I choose a dense or MoE model on this hardware?
MoE models are the sweet spot. Because only a few billion parameters are active per token, a 35B MoE can run near 43 tokens per second, far faster than a dense 70B on the same box while still being capable for coding.
Is it easy to buy in South Africa?
Local availability is limited, so most units arrive as grey imports around R75,000. Check warranty and support terms carefully, since grey-import cover differs from locally stocked machines.
Want a local-AI machine you can buy and support in Rand? Compare the AI PCs available at Evetech for capable inference boxes with proper local warranty.