Every time a new Flux image model lands, the first question on local AI art forums is the same: will my card even load it? The honest reality is that new Flux image models are hungry at full precision, but the quantised builds that follow each release bring them down to cards real people own. Flux.1 Dev needs roughly 23.8GB at FP16, yet FP8 and NF4 variants drag newer releases onto 12GB and even 8GB GPUs.

Quick Answer

Flux.1 Dev wants about 23.8GB of VRAM at full FP16 precision, which means a 24GB card. FP8 roughly halves that to around 12GB, NF4 squeezes onto 8GB cards, and GGUF Q4 builds fit 6 to 8GB. Match the quantisation to your VRAM and almost any modern GPU can run some Flux variant.

The VRAM ladder for Flux

Think of each Flux release as having several download flavours, each trading a little quality for a lot less memory. At the top, FP16 is the reference build and demands the most VRAM, which is why it effectively requires a 24GB card to run cleanly. That is the realm of cards like the RTX 4090, where you get the model exactly as released.

Step down to FP8 and the memory need drops to roughly 12GB while image quality stays very close to FP16, which is why FP8 is the format most people actually run. NF4 goes further, fitting into the 6 to 12GB band and running quickly on smaller cards, which makes it the practical choice for a 8GB GPU. GGUF quantised builds (Q4 through Q8) give you fine-grained control, with Q4 landing on 6 to 8GB cards and Q8 producing near-identical output to FP16 at around half the VRAM.

Keep your tooling current

The reason a brand-new Flux model sometimes refuses to load on day one is tooling, not hardware. Quantised formats like GGUF rely on loader nodes that get updated as each new variant ships. Keeping ComfyUI and the GGUF custom nodes up to date is what lets you run the latest release the week it drops rather than waiting weeks for compatibility to catch up.

If you are assembling or upgrading a machine specifically for local generation, VRAM is the spec that decides which formats you can touch, so it is worth planning around. You can see how purpose-built configurations are specced in the AI PC range at Evetech, which is a useful reference even if you intend to build your own.

Picking a card by the format you want

Work backwards from the format. If you want to run FP16 exactly as released and never think about quality loss, budget for a 24GB card. If FP8 quality is good enough (and for most work it is), a 12GB card covers you and costs far less. If you are on 8GB, NF4 or a GGUF Q4 build keeps you in the game, just with smaller batch sizes and a bit more patience. Whichever tier you target, comparing current pricing and VRAM across the GPU bestsellers at Evetech is the fastest way to see what each rand band actually buys today.

Frequently Asked Questions

Can I run Flux on an 8GB GPU?

Yes, with the right build. NF4 and GGUF Q4 quantised versions of Flux are designed to fit 6 to 8GB of VRAM. You give up some quality and batch size compared to FP8 or FP16, but the model runs.

Is FP8 noticeably worse than FP16?

For most use it is hard to tell apart. FP8 keeps image quality very close to FP16 while using roughly half the VRAM, which is exactly why it has become the default format for people running Flux locally.

Why does a new Flux model fail to load on my rig?

Usually because your loader nodes are out of date. Quantised formats depend on ComfyUI and GGUF custom nodes that update per release, so refresh those first before assuming your hardware is the problem.

How much VRAM should I buy for future Flux models?

More headroom is safer. A 12GB card runs FP8 comfortably today, while 24GB lets you run FP16 and gives margin for heavier future releases. If AI art is a serious hobby, lean toward the larger VRAM tier you can afford.

Building a rig that loads the next Flux model on launch day? Compare VRAM tiers and pricing in the AI PC and graphics card range at Evetech and choose the card that matches the precision you want to run.