PCIe Lane Bandwidth in a Multi-GPU Local Setup, Derived
9 min read · updated August 11, 2026
A second graphics card in a consumer motherboard usually lands in a slot wired for four lanes, not sixteen. What that costs is three separate numbers, and only one of them is large.
The bandwidth table, and where it comes from
PCIe bandwidth is not a marketing figure; it falls out of the raw transfer rate and the encoding, both fixed by the base specifications published by PCI-SIG. From PCIe 3.0 onward the line rate is 8, 16 and 32 GT/s for generations 3, 4 and 5, with 128b/130b encoding — so 128 bits of payload for every 130 bits on the wire.
per-lane, per-direction throughput = GT/s * (128/130) / 8 bytes
PCIe 3.0 8 GT/s -> 0.985 GB/s per lane
PCIe 4.0 16 GT/s -> 1.969 GB/s per lane
PCIe 5.0 32 GT/s -> 3.938 GB/s per lane
x16 x8 x4
PCIe 3.0 15.75 7.88 3.94 GB/s
PCIe 4.0 31.51 15.75 7.88 GB/s
PCIe 5.0 63.02 31.51 15.75 GB/sTwo properties of that table are worth internalising. It is per direction, so a full-duplex transfer gets that in each direction simultaneously. And a generation step is exactly equivalent to a doubling of width: PCIe 4.0 x4 and PCIe 3.0 x8 are the same 7.88 GB/s. That equivalence is how you reason about an older board with a newer card, or the reverse.
Model loading: the one that hurts
Every byte of the weights crosses the bus once to get into VRAM. That makes load time directly proportional to link bandwidth, and it is the only part of inference where the lane count produces a number a human notices.
Take the 42.5 GB Q4_K_M repackage of Llama 3.3 70B published on bartowski’s model card (retrieved 11 August 2026):
42.5 GB over PCIe 4.0 x16 (31.5 GB/s) = 1.35 s minimum 42.5 GB over PCIe 4.0 x4 ( 7.88 GB/s) = 5.39 s minimum 42.5 GB over PCIe 3.0 x4 ( 3.94 GB/s) = 10.8 s minimum
Those are floors, and in practice the storage path is usually the binding constraint instead: a Gen4 NVMe drive reading at 5 GB/s cannot feed a 31.5 GB/s link, so on a x16 slot the disk is what you are waiting for. On a x4 slot the bus becomes the constraint and the disk stops mattering. The crossover is around 8 GB/s of sequential read, which is roughly where current Gen5 NVMe sits.
The second load matters less than the first: once the file is in the operating system’s page cache, the source-side read is memory speed and only the bus transfer remains. If you are reloading models frequently — switching between two of them, or restarting a server — a narrow slot turns a two-second operation into a ten-second one, repeatedly. That is the case where lane count is worth paying for.
Generation: two microseconds a token
Once the weights are resident, a single-stream layer split moves almost nothing across the bus. The hidden state at the split boundary is one vector of hidden_size values per token: 8,192 values for a 70B, 16 KiB at fp16.
16,384 B over PCIe 4.0 x16 (31.5 GB/s) = 0.52 us per token 16,384 B over PCIe 4.0 x4 ( 7.88 GB/s) = 2.08 us per token against a token step of, say, 25 ms: 0.008% of the step
Eight thousandths of one per cent. There is no configuration of lane widths available on a consumer board that makes this number matter for a layer split. The mechanism behind that split, and the tensor-split case where the arithmetic comes out very differently, is on the two-card page.
Prompt processing sits between the two cases. The prompt is uploaded once, and a token is a handful of bytes, so even a 100,000-token prompt is well under a megabyte on the wire — irrelevant. But if the runtime is streaming layers in and out of VRAM because the model does not fit, every one of those layers crosses the bus on every token, and then the bus is the entire performance story. That configuration is not “a bit slower”; it is two orders of magnitude slower, and the fix is to make the model fit rather than to widen the slot.
How you end up in a narrow slot without knowing
- Chipset slots are not CPU slots. A consumer desktop CPU provides a limited number of lanes, typically sixteen to the primary graphics slot plus four to an NVMe drive. Every other slot hangs off the chipset, behind a link that is itself only four or eight lanes wide and shared with USB, SATA and networking.
- Populating an M.2 slot can steal lanes. Many boards electrically share the second M.2 socket with the second PCIe slot, or drop the primary slot from x16 to x8 when a second card is installed. This is in the manual’s slot-configuration table and nowhere else.
- Physical width is not electrical width. An open-ended x16 connector wired for four lanes is common and looks identical from above.
- Some cards are natively narrow. Several recent mid-range GeForce parts expose only eight lanes regardless of the slot. In a PCIe 3.0 board that card runs at PCIe 3.0 x8 = 7.88 GB/s, which is a quarter of what the same card gets in a Gen4 x16 slot — and that quarter is entirely a load-time cost, not a generation-speed one.
Bifurcation, risers and link training down
There is a way to get two properly-wired slots out of one, and it is worth knowing because it changes the arithmetic above by a factor of two. PCIe bifurcation splits a single x16 root port into two x8 links or four x4 links, in hardware, without a switch. Where the firmware exposes it — usually as an option named something like x8/x8 under the PCIe slot configuration — two cards each get a genuine eight-lane link off the CPU rather than one card getting sixteen and the other getting four off the chipset.
On PCIe 4.0 that turns a 31.5/7.88 GB/s pair into a symmetric 15.75 GB/s pair, which for the model-loading case is a straightforward win for the second card and a modest loss for the first. Bifurcation requires firmware support, a physical adapter that breaks the slot out, and cards that will train at the reduced width — all three, not two of the three.
The failure mode that follows from any of this is link training down: the endpoint and the root port negotiate, fail to establish a reliable link at the advertised rate, and settle on a lower generation or a narrower width without reporting an error anywhere. It is common with riser cables, which is how most multi-card builds physically fit, and it gets worse with cable length and with cheap shielding. A riser rated for PCIe 3.0 in a Gen4 slot will frequently produce a Gen3 link and nothing will say so.
The symptom is not a crash. It is a load that takes four times as long as it did last month, or a card that behaves differently from an identical one in the next slot. Because generation is unaffected, the usual benchmark shows nothing wrong, which is exactly why this is worth checking explicitly rather than waiting for it to announce itself.
Reading the link you actually got
Do not trust the manual; read the negotiated link. Both the maximum the hardware supports and the current state are exposed directly:
nvidia-smi --query-gpu=index,name,pcie.link.gen.max,pcie.link.width.max,\ pcie.link.gen.current,pcie.link.width.current --format=csv # on Linux, the same from the bus itself sudo lspci -vv -s 01:00.0 | grep -E "LnkCap|LnkSta"
One trap: an idle GPU drops its link to a lower generation to save power, so a reading taken at idle will under-report. Take the measurement while a model is loading, or start a load and query in a second shell. A card that reports pcie.link.width.current of 4 while a model is loading and the board manual promises 16 has a seating or a lane-sharing problem worth finding before you conclude anything else.