A discrete entry-level graphics card hits a hard wall the moment your AI model needs more memory than its fixed VRAM holds. An 8GB card runs out of room well before a useful local language model is loaded, and there is no way to add more. AMD's Strix Halo APU sidesteps that ceiling entirely: by pooling system memory and graphics memory into one shared block, a compact mini-PC can hold models that would never fit on a budget discrete GPU, all while sipping power around the clock.

Quick Answer

A Strix Halo mini-PC built on the Ryzen AI Max+ 395 pairs 16 Zen 5 CPU cores, a 40 compute-unit Radeon 8060S iGPU and an XDNA 2 NPU with up to 128GB of unified LPDDR5X memory. That shared pool lets the integrated GPU address tens of gigabytes for a model, far beyond the 8GB or 16GB an entry discrete card offers, making it a genuinely capable always-on host for local AI agents at roughly 65W.

What unified memory actually changes

On a normal desktop, your graphics card has its own sealed pool of VRAM. The model you want to run must fit inside that pool, full stop. Spill past it and performance collapses as data shuffles across the slow PCIe link to system RAM. An entry-level card with 8GB simply cannot hold a mid-sized language model, no matter how patient you are.

Strix Halo throws out that split. The CPU and the integrated Radeon 8060S graphics share a single LPDDR5X memory pool, configurable up to 128GB, both addressing it directly. Allocate a large slice to the GPU and a model that needs 30, 40 or 60GB suddenly fits in graphics-addressable memory on a machine you can hold in one hand. This is the same architectural trick Apple has shipped in its M-series chips since the M1, brought to the PC side with AMD silicon.

The trade-off is bandwidth. Soldered LPDDR5X-8000 on this platform delivers somewhere in the region of 210 to 256 GB/s, which is healthy for an integrated design but well below a high-end discrete card's dedicated memory. In practice that means token generation on very large models is steady rather than blistering. For an always-on agent host, where the job is to be available and to hold the model resident rather than to win benchmark races, that balance lands in the right place.

Why always-on changes the hardware maths

An AI agent host is not a gaming rig that spikes to full load for an evening and then sleeps. It runs continuously, waiting for a prompt, a webhook or a scheduled task, then responding and going quiet again. The metric that matters is not peak frames per second, it is watts at idle and watts under a light, frequent load, multiplied by twenty-four hours a day.

A full tower with a discrete GPU can pull 60 to 100W or more just idling with the card spun up, and that adds up on a power bill running every hour of every day. A Strix Halo mini-PC operates in a far tighter power envelope, with the whole package configured around the 65W class. Leaving it on permanently is realistic rather than wasteful, and a compact box tucks onto a shelf or behind a monitor without a tower's bulk or noise.

Sizing the memory to the model

The right configuration depends entirely on what you intend to run. A rough guide:

  • For small assistant models and embedding workloads, a 32GB to 64GB unified configuration leaves comfortable headroom.
  • For mid-sized models in the tens of billions of parameters at quantised precision, 64GB is the practical floor and 128GB removes the worry.
  • For the largest models you can realistically host locally, the 128GB option is what makes them fit at all.

Because the memory is soldered on this platform, you cannot upgrade it later. Choose the capacity you will want in a year, not just the one you need this week. The AI PC range at Evetech covers exactly this category, configured around unified-memory designs that match the workload directly.

Setting realistic expectations

Strix Halo is a remarkable fit for hosting a resident model and serving an agent, but it is not a substitute for a top-tier discrete GPU when raw inference speed on large models is the priority. If your workflow is heavy fine-tuning or you need the fastest possible generation on a big model and you have the power budget and the rands for a flagship card, a discrete setup still wins on throughput.

Where Strix Halo is hard to beat is the specific combination this title is about: holding a large model resident, staying available continuously, and doing it in a quiet, compact, power-frugal box. For a local agent that needs to be up at 3am as reliably as 3pm, that is the right set of priorities. If a complete balanced system is the starting point you prefer, the PC best sellers show which configurations SA buyers have settled on for this type of machine.

Frequently Asked Questions

How is Strix Halo different from a normal mini-PC?

A normal mini-PC has a modest integrated GPU and a small fixed share of memory for graphics. Strix Halo pairs a much larger 40 compute-unit iGPU with up to 128GB of unified memory the GPU can address directly, which is what lets it hold large AI models a standard mini-PC cannot.

Can I upgrade the memory later?

No. The LPDDR5X memory is soldered to the package for the high bandwidth this design needs, so whatever you order at purchase is what you have permanently. Pick your memory tier based on the largest model you expect to run.

Will it run a large language model as fast as a discrete GPU?

For holding and serving a model it is very capable, but generation speed on big models is slower than a high-end discrete card because the shared memory bandwidth is lower. For an always-on agent that prioritises availability over raw speed, that trade is usually worth it.

How much power does an always-on Strix Halo box use?

The platform is built around a 65W class power envelope and draws far less than a tower with a discrete GPU, especially at idle. That makes leaving it running continuously practical rather than costly.

Do I need it for cloud-based agents?

If your agents run entirely in the cloud you mainly need uptime and bandwidth, not local horsepower. Strix Halo earns its keep specifically when you want the model running locally on your own hardware rather than paying per token to an external service.

Want a compact, low-power box that can hold real AI models and stay up around the clock? Explore the AI PC range at Evetech to match unified memory capacity to the models you plan to run.