Skip to content

Local LLM Inference on a Laptop's Mobile GPU

9 min read · updated August 11, 2026

A laptop advertised with a familiar GPU name shares that name with a desktop card and little else. Two published specifications diverge, and both of them are ones that matter for inference specifically.

The same name is not the same card

Mobile GPUs are sold under the same model numbers as desktop parts and differ in three ways that all move in the same direction. They frequently carry less memory on a narrower bus. They run at a board power the laptop maker chooses within a range the GPU vendor allows, rather than a single published figure. And they share a cooling solution and a power budget with the CPU.

The practical consequence for buying: the model number tells you almost nothing. What you need from the specific machine’s specification sheet is the memory size, the memory interface width, the memory data rate, and the graphics power the manufacturer has configured. The first three give you bandwidth and capacity; the fourth tells you what happens after ten minutes.

Bandwidth is the spec that diverges

Single-stream generation is memory-bound: every token requires reading every weight once, so the ceiling is bandwidth divided by the byte count of the weights. That makes bus width and memory data rate the two numbers to find, and they are exactly the two that mobile parts cut.

bandwidth = (bus_width_bits / 8) * data_rate_Gbps

  256-bit @ 18 Gbps GDDR6  = 576 GB/s
  192-bit @ 18 Gbps GDDR6  = 432 GB/s
  128-bit @ 18 Gbps GDDR6  = 288 GB/s
   96-bit @ 18 Gbps GDDR6  = 216 GB/s

token ceiling (batch 1) = bandwidth / weight_bytes

  8B Q4_K_M is 4.92 GB (bartowski GGUF card, 11 Aug 2026)

  576 GB/s -> 117 tok/s ceiling
  288 GB/s ->  59 tok/s ceiling
  216 GB/s ->  44 tok/s ceiling

Those are ceilings implied by the arithmetic, not measurements, and real rates sit below them. The useful property of a ceiling is that it rules things out: a configuration whose ceiling is 44 tokens per second will not produce 60, whatever anyone claims, and a machine measurably far below its ceiling has a fixable problem rather than a hardware limit.

Note how much of the spread is bus width rather than generation. A 128-bit mobile part is at half the bandwidth of a 256-bit one carrying identical memory chips. This is the same reasoning applied to desktop cards on the 12GB page; the formula does not change, only the inputs do.

Configurable TGP and the sustained-load problem

Mobile GPU vendors publish a range of board powers and let the laptop maker pick a point in it, which is why two laptops with the same GPU name can differ substantially. The figure is variously called configurable TGP, dynamic boost or maximum graphics power, and it usually appears in the fine print of the machine’s specification page rather than next to the GPU name.

Two consequences specific to inference. First, inference is a sustained load, not a bursty one, so a machine that boosts well for thirty seconds and then settles will spend essentially all of its time at the settled figure. Any number you see quoted for short bursts is not the number you get. Second, the CPU and GPU share a thermal and power budget, which means the partial-offload configuration — some layers on the GPU, some on the CPU — is worse on a laptop than on a desktop, because loading both halves at once is exactly what the shared budget cannot do.

There is a third effect with no desktop equivalent: on battery, most machines apply a much lower power limit, and some disable the discrete GPU for compute entirely. A model that generates acceptably on mains can be several times slower unplugged, with no error and no notification. If you care about the unplugged case, measure it separately.

There is an asymmetry in how the two phases of inference respond to that squeeze, and it is worth planning around. Generation is memory-bound, so it is comparatively insensitive to a lower power limit: the cores are waiting on memory either way, and memory clocks move far less than core clocks when the budget tightens. Prompt processing is compute-bound and loses much more. The practical effect on a thermally constrained laptop is that short-turn chat degrades gracefully while long-document work degrades badly, and the two will not track each other from one machine to the next. Benchmark both, not one.

The last cost is simply that a sustained load runs the fans continuously for as long as the job does, which for a batch of documents means hours of it. That is a real reason to move this class of work to a desktop or an external enclosure even when the mobile part could technically carry it — the case for the Thunderbolt enclosure is partly thermal rather than purely about capacity.

A laptop VRAM budget, derived

The capacity arithmetic is the same as everywhere else — weights from the published GGUF size, cache from the model’s layer and head counts — with one addition that catches people: on a laptop the discrete GPU is nearly always also driving the display, so the desktop compositor’s framebuffer is inside the budget.

8 GB mobile GPU, ~7.4 GiB after CUDA context,
minus ~0.5 GiB for the desktop session -> ~6.9 GiB usable

  Llama 3.2 3B Q4_K_M  weights 1.88 GiB -> 5.0 GiB free
                       112 KiB/token    -> ~47k tokens of fp16 cache

  Llama 3.1 8B Q4_K_M  weights 4.58 GiB -> 2.3 GiB free
                       128 KiB/token    -> ~18k tokens of fp16 cache

An 8 GB mobile part runs an 8B at Q4 with a working but not generous context, and a 3B comfortably. A 6 GB part does not run an 8B at Q4 with any useful context at all, and the honest configuration there is a 3B or a 4B, not an 8B across the offload cliff.

Machines with unified memory rather than dedicated VRAM — where the GPU and CPU share one pool — change this calculation fundamentally, because the capacity available is a large fraction of system RAM rather than a fixed few gigabytes. That is a different architecture with different arithmetic; see unified memory systems.

Making a laptop usable for this

  • Read the machine’s spec sheet, not the GPU’s. Memory size, bus width, data rate and configured graphics power. If the manufacturer does not publish the bus width, that is itself informative.
  • Drive the display from the integrated GPU where the machine allows it. On a discrete-plus-integrated laptop this frees the framebuffer and, at 6 or 8 GB, that is thousands of tokens of context.
  • Fix the context length explicitly. A runtime that defaults to a model’s full advertised window will try to allocate a cache far larger than the card, and fail at load rather than at generation.
  • Prefer a smaller model fully resident over a larger one spilling. The offload cliff derived on the 12GB page is steeper on a laptop, because the CPU side of the split is also thermally constrained.
  • Measure on mains and on battery. llama-bench -m model.gguf -ngl 99 -p 512 -n 128 -r 5 in both states, and treat the battery figure as the one you actually have when travelling.
Mobile GPU memory configurations and configurable power ranges change with every generation, and manufacturers revise them within a generation. The formulas above are stable; the specific bus widths and data rates are not, so take them from the current spec sheet for the exact machine.