The single number that decides what AI image models you can run locally is VRAM, and the jump from 16GB to 24GB to 32GB is not a smooth gradient. Each tier unlocks a specific class of model and resolution, and spending up a tier you do not need is as wasteful as buying short and hitting a wall mid-workflow. The 24GB tier sits at the centre of the decision, because that is where full-quality Flux generation becomes comfortable.

Quick Answer

16GB runs Stable Diffusion and SDXL well but cannot hold Flux.1 Dev at full precision. 24GB clears Flux.1 Dev in practical FP16 settings and is the sweet spot for serious image work. 32GB adds headroom for higher resolutions and local video generation. For most image-gen workflows, 24GB is the pivot point.

Why VRAM is the hard limit

An image model has to fit in graphics memory to run. If the model weights plus the working data exceed your VRAM, generation either fails outright or falls back to painfully slow system-memory swapping. Unlike a slow CPU, which just makes things take longer, insufficient VRAM stops a model from running at all at its intended quality. That is why VRAM, not raw GPU speed, frames this entire conversation.

Quantisation muddies the picture slightly: a model can be compressed to FP8 or GGUF formats that shrink its memory footprint at some cost to quality. So the honest framing is what each tier runs at full quality, with quantised options as the fallback below it.

16GB: Stable Diffusion territory

A 16GB card is a capable starting point and runs the Stable Diffusion family comfortably. SD 1.5 needs only a few gigabytes, and SDXL base inference sits around 8 to 10GB, with the refiner adding a couple more. SD 3.5 Large in its FP8 form fits here too. You can also run Flux at lower precision, FP8 lands around 13GB and GGUF Q4 around 7GB, so Flux is technically usable on 16GB, just not at full fidelity.

Where 16GB struggles is full-precision Flux.1 Dev and high-resolution batch work. If your output is SDXL portraits and standard-resolution art, 16GB is genuinely enough and the most cost-effective choice.

24GB: the Flux full-precision tier

24GB is where serious image generation opens up. Flux.1 Dev runs in practical FP16 settings around the 24GB mark, so you get the model's full quality rather than a compressed approximation. This tier also gives you comfortable headroom for SDXL at higher resolutions, larger batches, and running a model alongside upscalers and ControlNet without constant memory juggling.

Why this is the pivot point

For anyone doing image generation as more than a casual hobby, 24GB removes the daily friction of working around memory limits. You stop choosing quantised fallbacks to fit the card and start running models as their creators intended. That shift, from working around VRAM to ignoring it, is exactly why 24GB is the tier most serious local image creators target. For builders sizing a machine around this, the AI PC range at Evetech pairs these cards with the CPU and memory to keep them fed.

32GB: high resolution and local video

32GB is for workflows that push past still images. The extra memory buys headroom for very high resolutions, heavier batch generation, and the increasingly popular local video generation models, which are far hungrier than image models. If you are experimenting with text-to-video, training or fine-tuning, or generating at large dimensions, 32GB stops being a luxury and becomes the working minimum.

For pure image generation, though, 32GB is often more than the workload demands, and the money is better spent elsewhere unless video or training is on your roadmap. Comparing what local AI builders are actually buying helps calibrate; the top-selling graphics cards at Evetech show where the practical sweet spot sits.

Where Flux.2 Fits Into the Equation

Flux.2, Black Forest Labs' 32-billion-parameter follow-up released in late 2025, shifts the goalposts again. Running it at native precision requires around 90GB of VRAM, far beyond any consumer card. The practical route is FP8 on a 32GB GPU, which is the only consumer configuration that handles Flux.2 Dev without falling back to aggressive GGUF compression. For anyone on 24GB, Flux.2 is GGUF territory only, and the quality gap at very low quant levels is more visible than it is with the lighter Flux.1 Dev model. Knowing where the next generation sits is useful when deciding whether to stop at 24GB or stretch to 32GB.

Generation Speed: What Each Tier Delivers in Practice

VRAM determines model fit, but architecture drives speed within that constraint. The RTX 4090 generates an SDXL image in around 3 seconds and a full Flux.1 Dev FP16 image in roughly 18 seconds. A 16GB card running Flux at FP8 is slower still because it has fewer compute units alongside its smaller memory. The practical takeaway: iterations per second varies not just with VRAM tier but with the model and precision you are running. Budget for the VRAM your workflow needs, then the card with the strongest compute in that tier.

Matching the tier to your work

Map the card to the output, not the marketing. SDXL art and standard resolutions, 16GB. Full-quality Flux.1 and serious image workflows, 24GB. High-res, batch-heavy work, local video, or early access to Flux.2 at a usable level, 32GB. The most expensive mistake is buying a 16GB card and immediately wishing you had Flux headroom, or buying 32GB for work that never leaves SDXL.

Frequently Asked Questions

Can I run Flux on a 16GB GPU?

You can, but not at full precision. Flux fits on 16GB in FP8 (around 13GB) or GGUF Q4 (around 7GB) form, which trades some quality for the smaller footprint. Full FP16 Flux.1 Dev really wants 24GB.

Is 24GB enough for Stable Diffusion and SDXL?

Comfortably. 24GB runs SDXL at high resolutions, large batches, and alongside upscalers and extra networks without memory pressure. It is overkill for SD 1.5 alone but ideal once you add Flux to the mix.

Do I need 32GB for AI image generation?

Not for still images. 32GB matters when you move into local video generation, very high resolutions, or model training and fine-tuning, all of which are far more memory-hungry than standard image generation.

Does a faster GPU help if it has less VRAM?

Only up to a point. If the model does not fit in VRAM, raw speed does not rescue it, generation either fails or crawls on system memory. Fit the model first, then chase speed within that tier.

What lowers a model's VRAM requirement?

Quantisation. Converting a model to FP8 or GGUF formats shrinks its memory footprint substantially, letting larger models run on smaller cards at some cost to output quality and sometimes speed.

Building a machine for AI image generation? Match the VRAM to your models before anything else, then browse current options in the AI PC range at Evetech to spec a system that runs your workflow without compromise.