There is a trap in choosing a graphics card for local AI: the model loads, you celebrate, and then it crawls or runs out of memory the moment you ask it to hold a real conversation. That gap is the difference between minimum vs recommended GPU memory. The minimum is what gets the model onto the card at all; the recommended is what lets it run well with room for the context you actually use. They are two different targets, and confusing them costs you.
Quick Answer
Minimum VRAM loads the model with no headroom; recommended VRAM runs it well with context to spare. A 7B model fits an 8GB card comfortably, but a 13B's 8GB to 10GB need leaves zero room on the same 8GB card, so 12GB is the recommended target for the next tier up.
Why the two numbers diverge
The minimum figure counts only the model's weights at a given quantisation. The recommended figure adds the working memory the model needs while running: the context window holding your conversation, plus the operating system's own overhead. That extra demand grows as the conversation lengthens, which is why a model that loads on the dot of your card's capacity stalls or fails once you actually use it.
So the rule is simple. Buy to the recommended number, not the minimum, unless you only ever run short, single prompts. The headroom is what turns a model that technically fits into one that is pleasant to use.
The tiers, minimum versus recommended
A 7B model at Q4 needs roughly 4GB to 6GB minimum, and an 8GB card is the comfortable recommendation, since it loads the model with plenty left for context. This is the safest entry tier.
A 13B model needs about 8GB to 10GB minimum. An 8GB card can technically load it, but with no headroom it struggles the moment context builds, so a 12GB card is the recommended floor. A 32B model needs around 20GB to 24GB, which makes a 24GB card the realistic recommendation rather than a tight squeeze on a 20GB part. A 70B model needs 40GB or more, so the recommendation moves to a workstation-class card or a pair of 24GB cards working together. Read across that table and the pattern is clear: every tier wants the card one notch above its bare minimum. The AI-ready desktops that house these cards sit in the AI PC range, and the higher-memory cards local-AI buyers favour tend to appear among the popular PC picks.
Frequently Asked Questions
Why can a 7B model run well on 8GB but a 13B cannot?
A 7B's weights take only 4GB to 6GB, leaving the 8GB card room for context. A 13B's weights sit near 8GB to 10GB, so on an 8GB card there is nothing left for the conversation, which makes it the recommended jump to 12GB.
Should I buy to the minimum or recommended VRAM?
Recommended, unless you only run short prompts. The minimum loads the model but leaves no room for a growing context window, so a card sized to the recommendation is what actually runs the model comfortably.
Does a longer conversation need more VRAM?
Yes. The context window holding the conversation grows as it lengthens, consuming more memory on top of the weights. That is the main reason the recommended figure sits above the bare minimum.
Can two smaller cards replace one big one for 70B models?
Often yes. A pair of 24GB cards can combine to host a 70B model that needs 40GB or more, which is a common alternative to a single very large workstation card.
Sizing a card for local models means buying to the recommended memory, not the minimum. Match your target model tier to a build in the AI PC range at Evetech and leave yourself the headroom to run it well.