Definition: An interconnect is the wiring between AI chips. NVLink is NVIDIA's high-speed interconnect, which lets many GPUs in one system share data fast enough to behave as a single machine. Other chip families have their own versions, such as the inter-chip interconnect (ICI) on Google's TPUs.
TL;DR: A frontier model does not fit on one chip, so its weights are split across many, and those chips must exchange results for every generated token. The speed of the links between chips then sets the speed of the whole system. NVIDIA says fifth-generation NVLink moves 1,800 GB/s per GPU, and a 72-GPU rack reaches 130 TB/s. Build an app free →
Why Interconnects Matter
A single accelerator has fast local memory. A big model has more weights than one chip can hold, so software slices the model across chips. That is called model parallelism. After each slice computes, the chips must combine partial results before the next layer can start.
If the links are slow, the chips wait. The arithmetic units sit idle, the same failure described in memory bandwidth but one level out. Memory bandwidth is the road inside the chip. The interconnect is the road between chips.
This gets worse for mixture-of-experts models, where each token is routed to different experts that can live on different chips. Every token then crosses the fabric.
How It Works
The exchange step repeats at every layer, so a slow link is paid for again and again.
Two Networks, Two Jobs
An AI cluster uses two kinds of network at once:
- Scale-up. A fast, tightly coupled fabric inside a rack or pod, so a group of chips shares memory access and acts as one large accelerator. NVLink and Google's ICI are scale-up fabrics.
- Scale-out. A wider data center network that links racks and pods to each other, usually with lower bandwidth per chip and more distance.
Work that needs constant exchange stays on the scale-up fabric. Work that can tolerate delay crosses the scale-out network.
The Numbers, Verified
| Link | Bandwidth | Source note |
|---|---|---|
| PCIe 5.0 x16 | 128 GB/s, both directions combined | PCI-SIG specification, 32 GT/s per lane |
| NVLink 5, per GPU | 1,800 GB/s | NVIDIA product page |
| NVLink domain, 72 GPUs | 130 TB/s total | NVIDIA product page |
| Google Ironwood ICI, per chip | 1,200 GB/s, bidirectional | Google Cloud documentation |
Divide 1,800 by 128 and NVLink 5 offers about 14 times the bandwidth of one PCIe 5.0 x16 slot. Different vendors count bandwidth differently, so treat cross-vendor comparisons as approximate.
Worked Example: A 72-GPU Rack
NVIDIA lists the GB200 NVL72 with 72 GPUs, 13.4 TB of HBM3E memory, and 576 TB/s of aggregate GPU memory bandwidth. Its NVLink fabric moves 130 TB/s. Check the arithmetic: 72 GPUs at 1.8 TB/s is about 130 TB/s.
Compare the two aggregate figures. The fabric carries about 23% of the bandwidth the memory delivers (130 divided by 576). That is why software works hard to keep data local and to overlap communication with computation. Even inside a rack built to act as one machine, moving data between chips costs more than reading local HBM.
Why the Rack Is the New Unit
For years the unit of AI compute was a server with 8 GPUs. Rack-scale designs put 72 accelerators in one fast domain, which changes what fits: larger models, longer context windows, and bigger KV caches can stay on the scale-up fabric instead of crossing slower links. The cost is density. Such a rack draws roughly 120 kW, which is why it needs liquid cooling and a redesigned data center.
Common Mistakes
- Comparing headline bandwidth across vendors. Some figures count both directions, some count one, and some count per chip versus per rack.
- Assuming more chips means proportional speed. Communication overhead grows with cluster size.
- Ignoring the scale-out network. A fast rack still slows down if racks connect through a weak network.
- Treating NVLink as only an NVIDIA topic. Every large accelerator cluster has the same problem and a comparable fabric.
Connection to Taskade
Interconnects are below the layer where you work. When you run an AI agent or an automation on Taskade, model serving happens behind a hosted model call on secure infrastructure, and you never size a cluster. The practical effect is indirect: better fabrics let providers serve larger models at lower cost per token.
What You Would Build in Taskade
A "latency log": a project where an automation records how long each agent reply takes, so you can see which model choices feel slow for your team. Describe yours and build it free →
Related Concepts
- GPU - the chip the fabric connects
- TPU - Google's chip with its own pod fabric
- HBM - the memory each chip reads locally
- Memory Bandwidth - the on-chip limit that interconnect extends
- Liquid Cooling - the thermal cost of dense racks
- Mixture of Experts - the architecture that stresses the fabric
- Tokens Per Second - the user-visible result
Frequently Asked Questions About NVLink
What is NVLink?
NVLink is NVIDIA's high-bandwidth interconnect for linking GPUs, and in some systems CPUs, so they can exchange data much faster than over a standard PCIe slot. In rack-scale systems it lets 72 GPUs act as one large accelerator.
How fast is NVLink?
NVIDIA lists fifth-generation NVLink at 1,800 GB/s per GPU, and 130 TB/s across a 72-GPU NVLink domain. For comparison, a PCIe 5.0 x16 slot provides 128 GB/s with both directions combined.
What is the difference between NVLink and PCIe?
PCIe is the general-purpose standard that connects cards, drives, and network adapters to a server. NVLink is a dedicated chip-to-chip link built for GPU traffic, with far higher bandwidth. GPUs still use PCIe for many host connections.
What is the difference between scale-up and scale-out networking?
Scale-up links chips inside a rack or pod so they share memory access and act as one machine. Scale-out links racks and pods across the data center. Scale-up is faster and shorter. Scale-out reaches further.
Do TPUs have something like NVLink?
Yes. Google TPUs connect through an inter-chip interconnect (ICI). Google's documentation lists 1,200 GB/s of bidirectional ICI bandwidth per Ironwood chip, and pods of up to 9,216 chips.
Why do AI models need fast chip-to-chip links?
Because large models are split across many chips, and the chips must combine partial results at every layer. Slow links leave the chips waiting. The interconnect therefore limits speed once a model no longer fits on one chip.
What is NVLink Fusion?
NVIDIA describes NVLink Fusion as a way to combine NVIDIA technology with semi-custom ASICs or CPUs, so other hardware designers can use the NVLink fabric. NVIDIA's page gives no bandwidth figures for it.