The marketing line on a new Copilot+ laptop reads 85 TOPS, and that number sounds like it should run a chatbot at home with ease. It will run a small quantised model, slowly and quietly, but the moment you want a 70 billion parameter model or a long context window, the NPU hits a wall a discrete GPU never sees. The reason is not raw compute. It is memory: how much the chip can hold and how fast it can move it. An RTX 5090 carries 32GB of dedicated VRAM at 1,792 GB/s, and for local LLM inference that combination wins decisively.

Quick Answer

For serious local LLM inference, a discrete GPU beats an 85 TOPS NPU outright. The RTX 5090's 32GB of GDDR7 at 1,792 GB/s lets it load large models and generate text at roughly 85 tokens per second, while a Copilot+ NPU manages around 25 tokens per second on small models because it shares slow system RAM. TOPS measures peak math, not the memory that actually feeds the model.

Why TOPS is the wrong number to shop on

TOPS counts trillions of operations per second the chip can theoretically perform. It is a peak compute rating, and for the image and audio tasks NPUs were designed for, background blur, voice cleanup, live captions, it is a fair guide. LLM text generation is a different problem. Generating each token means reading the entire model's weights from memory once. The bottleneck is not how fast the chip can multiply, it is how fast it can pull those weights in. That is memory bandwidth, and it is the spec almost nobody puts on the box.

This is why two devices with similar TOPS can differ wildly at running a chatbot. The one with faster, dedicated memory wins, because the maths units on either spend most of their time waiting for data to arrive.

The bandwidth gap, in plain numbers

A Copilot+ NPU runs on the laptop's shared LPDDR5X system memory. That memory is fast for a laptop but pools at roughly 100 to 135 GB/s and is split between the CPU, the integrated graphics and the NPU. The RTX 5090's 32GB of GDDR7 runs on a 512-bit bus at 1,792 GB/s, dedicated entirely to the GPU. That is well over ten times the throughput feeding the model.

What that means for tokens per second

Token generation speed in an autoregressive model scales almost linearly with memory bandwidth. In practice a Copilot+ NPU produces conversational speed, around 25 tokens per second on a small model at about 15W, silent and cool. The RTX 5090 produces roughly 85 tokens per second, faster than you can read, drawing far more power and spinning its fans hard. For a quick assistant the NPU is pleasant. For agentic workflows, code generation or anything with a long prompt, the GPU's pace changes what is workable.

VRAM decides which models you can run at all

Bandwidth sets speed; capacity sets the ceiling. A model has to fit in memory to run well. 32GB of VRAM comfortably holds a 32B parameter model at a sensible quantisation, or a smaller model with a very long context window. The shared pool an NPU draws from gets squeezed by everything else the laptop is doing, so it is practically limited to small 3B to 8B class models. Spill past available memory and inference slows to a crawl as data shuttles to and from disk.

Where the NPU genuinely earns its place

None of this makes the NPU pointless. It is the right tool for always-on, low-power inference: a small model summarising your day, on-device transcription, photo search, all running for the cost of a few watts with no fan noise and no battery penalty worth mentioning. For a thin laptop that needs to last a working day, that efficiency is the entire point. The category of machine built around it is worth understanding, and the AI PC range at Evetech shows what these chips target. The mistake is expecting that same silicon to replace a desktop GPU for heavy model work.

For anyone whose goal is local inference at speed and scale, the build is a desktop with a high-VRAM card. Browsing the systems our customers buy most gives a sense of where the RTX class GPUs land, and how much chassis and cooling a 575W card asks for.

Who should buy which

Choose an NPU-equipped Copilot+ laptop if you want portable, efficient, always-on AI features and your models stay small. Choose a discrete GPU desktop if you run large models, need long context, generate code, or want generation fast enough to feel instant. They are not competitors so much as different answers to different questions, and the spec that separates them is memory, not the TOPS figure on the sticker.

Frequently Asked Questions

Is 85 TOPS enough to run a large local LLM?

Not on its own. TOPS rates peak compute, but large-model inference is limited by memory bandwidth and capacity. An 85 TOPS NPU pulling from shared system RAM cannot feed a big model fast enough, so it is restricted to small quantised models.

How much VRAM do I need for serious local inference?

For comfortable headroom, 24GB to 32GB. The RTX 5090's 32GB holds a 32B class model at a usable quantisation or a smaller model with a long context window. Less VRAM forces you onto smaller models or aggressive compression.

Why is memory bandwidth more important than TOPS for LLMs?

Generating each token requires reading the whole model from memory once. The compute units finish quickly and then wait for data, so the speed you feel depends on how fast memory delivers weights. Bandwidth, not peak math, sets tokens per second.

Can an NPU and a GPU work together?

In practice you pick the one that suits the job. NPUs handle light, always-on tasks efficiently; the GPU handles heavy model work fast. Some toolchains can offload across both, but the GPU does the demanding inference.

What about power and noise?

This is the NPU's real advantage. It runs near 15W in silence, ideal for a laptop on battery. The RTX 5090 can draw 575W under full load with audible cooling, which is the trade for its speed and capacity.

Building a rig for local AI? Explore the AI PC range at Evetech and match the memory and GPU to the models you actually plan to run. https://www.evetech.co.za/PC-Components/ai-pcs-445