Quick Answer

A local-LLM rig in Hermanus is sized by GPU VRAM: an RTX 4060 Ti 16 GB (~R10,500) runs quantised 7B-13B models like Llama 3 and Mistral well, and a 24 GB RTX 4090 handles 70B models. Pair the GPU with a Ryzen 5 7600 and 32 GB system RAM. For local AI, VRAM size sets the model ceiling - it matters far more than gaming frame rates.

VRAM tiers for local LLMs

Local language models live in GPU memory, so VRAM caps model size and speed:

  • 12 GB: 7B-8B models (Mistral 7B, Llama 3 8B)
  • 16 GB (RTX 4060 Ti 16 GB): 13B models, longer context
  • 24 GB (RTX 4090): 70B quantised models

System RAM (32 GB) lets you offload spillover layers to the CPU, but speed falls sharply when it does, so match the GPU to your largest model.

A Hermanus local-LLM build

  • GPU: RTX 4060 Ti 16 GB (~R10,500)
  • CPU: Ryzen 5 7600 (~R4,200)
  • RAM: 32 GB DDR5 (~R2,000)
  • Storage: 1 TB NVMe (~R1,200) - model files are 4-40 GB each

Hermanus is on the Overberg coast, so Evetech deliveries take 3-4 business days. Salt air means keeping the PC ventilated and off the floor; power-test before signing.

FAQ

Can I run a local LLM on a gaming GPU?

Yes. The same RTX 4060 Ti 16 GB that games well also runs 7B-13B LLMs locally via Ollama or LM Studio. Local AI just leans on VRAM and the CUDA cores rather than frame-rate performance.

How fast is local inference on a 16 GB card?

A quantised Llama 3 8B generates tokens quickly on an RTX 4060 Ti 16 GB - fast enough for interactive chat. 13B models run a little slower but remain usable; speed depends on quantisation and context length.

Does Evetech deliver to Hermanus?

Yes, Evetech ships to Hermanus on the Overberg coast in 3-4 business days. Power-test the GPU and confirm it's detected before signing.

TIP

files on the fast NVMe and load with GPU offload set to all layers in LM Studio - it minimises load time and avoids the slow CPU fallback that drags down token speed.