Pick the wrong VRAM tier for AI generation and you do not get a slow result, you get no result. The model simply refuses to load, or it offloads to system RAM and crawls to a halt. Image and video models each have a hard floor, and those floors sit roughly at 12GB for Stable Diffusion, 24GB for Flux at full precision, and 32GB once you move into local video. Matching your card to the work you actually plan to do is the single decision that saves the most money and frustration.

Quick Answer

For local AI generation, 12GB of VRAM is the practical floor for Stable Diffusion, 24GB clears Flux.1 Dev comfortably (full FP16 wants closer to 24GB or more), and 32GB is where local video generation with models like Wan 2.2 14B becomes realistic. Buy for the heaviest workload you intend to run, because VRAM capacity is a hard wall, not a speed dial you can turn down.

The 12GB Tier: Stable Diffusion and Entry AI Art

Twelve gigabytes is where serious local image generation starts. It comfortably runs Stable Diffusion and SDXL, handles reasonable batch sizes, and leaves room for the extras that make the work fun, like LoRAs and ControlNet. For someone learning the ropes or generating stills as a hobby, this tier does the job without drama.

It even gets you onto Flux through the back door. Flux.1 Dev does not fit in 12GB at full precision, but GGUF quantisation compresses the model down to Q4 or Q8, which lets a 12GB (and sometimes 8GB to 10GB) card run it with a modest quality trade off. So 12GB is not just an SD tier, it is a flexible entry point into most of the image generation world.

The 24GB Tier: Flux at Full Quality

Flux.1 Dev is the model that pushes people up a tier. The diffusion transformer alone is around 12 billion parameters, and running it together with its text encoder at full FP16 wants roughly 24GB or more for maximum speed. At 24GB you can run Flux properly, and FP8 is widely regarded as the sweet spot here, delivering image quality almost indistinguishable from FP16 while using close to half the memory.

Why FP8 Changes the Maths

FP8 matters because it makes a 24GB card the practical home for Flux rather than a workstation requirement. You get near full quality, faster generation, and enough spare memory to stack LoRAs and run higher resolutions. If polished Flux output is your goal and you are not yet doing video, 24GB is the tier that hits the value sweet spot.

The 32GB Tier: Local AI Video

Video is a different scale of problem. Generating clips locally with a model like Wan 2.2 14B, even at modest resolutions like 720p, pushes well past 24GB and into 32GB territory once you account for the model weights, the frames being held in memory, and the temporal data that ties them together. This is the tier for people who are serious about local video generation rather than stills.

A 32GB card also future proofs an image workflow. It runs Flux at full FP16 with room to spare, handles large batches, and absorbs whatever the next generation of models demands. Cards at this tier sit at the top of the consumer stack, and you can see what is currently available in the GPU best sellers list. If a 32GB consumer card is not enough, unified memory systems in the AI PC category offer far larger memory pools at the cost of raw speed.

The Flux.2 and Next-Generation Consideration

Flux.2, released in late 2025, raised the bar further. The full 32-billion-parameter model needs around 90GB at native precision, placing it firmly in data-centre territory. Practically, an RTX 5090 with 32GB can run Flux.2 Dev in FP8 form, which places this model at the very edge of what current consumer cards handle. For anyone building now with future headroom in mind, 32GB is not just adequate for current video work, it is the entry point for running next-generation image models at quantised quality without yet needing a cloud render.

How to Choose Your Tier

Work backwards from your heaviest intended task. If you only generate stills with Stable Diffusion, 12GB is enough and quantisation extends it to Flux.1. If clean Flux.1 output is the goal, target 24GB and lean on FP8, where you get near-full quality at roughly half the VRAM demand. If you want to generate video locally or run Flux.2 at a usable quality level, plan for 32GB. VRAM cannot be added later, so it is the one spec worth buying ahead of your needs rather than catching up to them later.

Frequently Asked Questions

What software handles all these precision formats?

ComfyUI is the practical home for FP16, FP8, and the full GGUF range. It lets you load any quantised Flux variant through a dedicated node, swap between precision levels without rebuilding your workflow, and keep an eye on VRAM usage in real time. Most community model releases come with a ComfyUI workflow attached, so the setup overhead is lower than it used to be.

Can I run Flux on a 12GB card?

Yes, with quantisation. Full FP16 Flux needs around 24GB, but GGUF formats compress the model down to Q4 or Q8, letting 12GB (and even 8GB to 10GB) cards run it. The trade off is a small drop in quality and slower generation compared with a higher tier card.

Is FP8 noticeably worse than FP16 for Flux?

For most work, no. FP8 produces images almost identical to FP16 while using close to half the VRAM, which is why it is the recommended choice on 24GB cards. Only the most demanding side by side comparisons tend to reveal a difference.

How much VRAM do I need for AI video generation?

Plan for 32GB. Local video models such as Wan 2.2 14B, even at 720p, exceed 24GB once weights, frames and temporal data are loaded. A 32GB card is the realistic entry point for generating video on your own machine rather than renting cloud compute.

What happens if a model exceeds my VRAM?

It either fails to load outright or spills over into system RAM, which slows generation to a crawl because system memory is far slower than VRAM. Capacity is a hard wall, so the fix is a larger card rather than a setting change.

Does more VRAM make generation faster?

Not directly. VRAM determines whether a model fits at all and how large your batches can be, while raw generation speed comes from the GPU's compute cores. A bigger card helps most by letting you run heavier models and higher resolutions without offloading.

Building a machine for AI art or video? Match your VRAM to the work in the GPU best sellers range at Evetech and talk to the team about sizing a card that loads the models you actually plan to run.