Definition: A TPU (tensor processing unit) is a custom AI accelerator that Google designed to run neural network math. It is an ASIC, a chip built for one job, and its core is a grid of multiply-and-add units that passes data from cell to cell instead of shuttling it through memory each time.
TL;DR: A GPU is a general parallel processor that turned out to suit AI. A TPU is a chip designed for AI from the start. Google's first TPU ran in its data centers from 2015, and its 2026 generation splits into one chip for training and one for inference. TPUs and GPUs now compete on memory, interconnect, and cost per token, not on raw arithmetic. Build an app free →
Why TPUs Matter in 2026
For most of the deep learning era, "AI chip" meant "NVIDIA GPU". That is no longer the whole picture. Google announced its seventh-generation TPU, Ironwood, in April 2025 and described it as the first TPU designed specifically for inference. At Cloud Next in April 2026, Google announced its eighth generation as two chips: TPU 8t for training and TPU 8i for inference and reinforcement learning. Both are due to be generally available later in 2026, according to Google.
The split is the news. Training and serving put different pressure on a chip. Training wants huge clusters that stay in sync. Serving reasoning models and agents wants low latency and enough fast memory to hold a large KV cache. A chip tuned for one job can beat a chip that must do both.
The Kitchen Analogy
A CPU is a head chef who can cook anything. A GPU is ten thousand line cooks who all chop at once. A TPU is a bread factory: one very long conveyor belt, built to make exactly one thing at very high volume. It cannot cook a steak. It makes an enormous number of identical loaves per second.
How a TPU Works
The heart of a TPU is a systolic array. In the first TPU, described in Google's 2017 paper, the matrix unit held 65,536 eight-bit multiply-accumulate cells, a 256 by 256 grid. Data flows through the grid in rhythm, like blood through a heart, which is where the word "systolic" comes from.
A value read from memory once is reused by many cells as it moves through the grid. That saves memory traffic, the usual limit described in memory bandwidth.
Three design choices follow from that:
- Fixed function. The chip spends its silicon on matrix math and little else.
- Data reuse. Results pass between neighbors, so fewer trips to memory.
- Custom networking. TPUs connect chip to chip in large pods with their own interconnect, so a whole pod can act as one machine. See NVLink and interconnects for the general problem.
The TPU Lineup, Verified
| Generation | Role | Headline fact (vendor-stated) |
|---|---|---|
| First TPU (2015 deployment) | Inference | 65,536 8-bit multiply units; 30x to 80x higher TOPS per watt than contemporary CPUs and GPUs, per the 2017 paper |
| Ironwood (7th gen) | Inference focus | 4,614 FP8 TFLOPs, 192 GiB HBM at 7,380 GB/s, 9,216 chips per pod |
| TPU 8t (8th gen) | Training | 9,600 chips per superpod, 121 ExaFlops, 2 petabytes shared memory |
| TPU 8i (8th gen) | Inference and RL | 288 GB HBM, 384 MB on-chip SRAM, 80% better performance per dollar than the prior generation |
Every number in the last two rows is Google's own claim. No neutral benchmark compares them with GPUs on the same workload.
Worked Example: One Ironwood Pod
Google lists 4,614 TFLOPs of FP8 compute per Ironwood chip and 9,216 chips in a full pod. Multiply them: 9,216 chips times 4.614 petaFLOPs is about 42,500 petaFLOPs, or 42.5 exaFLOPs. That matches Google's published pod figure.
Now look at memory. Each chip holds 192 GiB of HBM. A model with 70 billion parameters at 8-bit precision needs about 70 GB for its weights, so one chip holds the weights with room left for the KV cache. A model ten times larger must be split across chips, and then the chip-to-chip links decide the speed.
TPU vs GPU
| Question | TPU | GPU |
|---|---|---|
| Built for | Neural network math | Graphics first, then AI |
| Flexibility | Narrower | Broader |
| Software | Google's stack (JAX, XLA, and TensorFlow) | Mature CUDA ecosystem |
| Where you get it | Rented through Google Cloud | Many vendors and clouds |
| Main differentiator | Pod-scale networking and cost | Ecosystem and availability |
Neither wins in general. The right question for a buyer is cost per million tokens on the workload you run, measured on the actual model.
Common Mistakes
- Comparing peak FLOPs across chips. Precision, memory, and networking change real speed. A chip with fewer FLOPs can serve faster if it has more bandwidth.
- Treating vendor multiples as neutral. "80% better performance per dollar" compares a chip with its own predecessor.
- Assuming a TPU replaces a GPU one-for-one. Software, model code, and serving stacks differ.
- Forgetting the memory. A TPU still depends on HBM, and HBM supply limits every accelerator.
Connection to Taskade
You do not pick a chip when you use Taskade. On a paid plan, you pick a model for each AI agent from the OpenAI GPT, xAI Grok and open-weight models in the picker, or bring your own Anthropic key for Claude on Enterprise, and the platform handles capacity behind that choice. Chip economics still reach you indirectly: faster and cheaper serving hardware lowers the cost of agentic loops and automations.
What You Would Build in Taskade
A "model cost tracker": a project where an agent logs the model used for each task, an automation records the run, and a dashboard shows which workflows use the most model calls. Describe yours and build it free →
Related Concepts
- GPU - the general-purpose accelerator TPUs compete with
- HBM - the memory stacked beside every AI chip
- Memory Bandwidth - the limit a systolic array is designed to soften
- NVLink and Interconnects - how chips link into one machine
- Liquid Cooling - how pods stay cool
- Data Center - the building around the pod
- Inference · Model Training
- Inference Cost - where chip choice shows up on a bill
Frequently Asked Questions About TPUs
What does TPU stand for?
TPU stands for tensor processing unit. A tensor is a multi-dimensional array of numbers, the basic data type of neural networks. The chip is built to multiply and add tensors quickly.
What is the difference between a TPU and a GPU?
A GPU is a flexible parallel processor originally built for graphics. A TPU is a custom ASIC built for neural network math. The TPU trades flexibility for efficiency on that one job, while the GPU keeps a larger software ecosystem.
Who makes TPUs?
Google designs TPUs and offers them through Google Cloud. Google's 2017 paper says the first one had been in its data centers since 2015. Other companies sell TPU-style custom accelerators under different names.
Is a TPU faster than an NVIDIA GPU?
It depends on the model, precision, and cluster size. Vendors publish their own numbers, and no neutral benchmark in this page's sources compares Ironwood or the eighth generation with a GPU on the same workload. Compare cost per million tokens on your own model.
What is Ironwood?
Ironwood is Google's seventh-generation TPU, announced in April 2025. Each chip has 192 GiB of HBM and 4,614 FP8 TFLOPs, and a full pod links 9,216 chips. Google describes it as its first TPU designed specifically for inference.
What are TPU 8t and TPU 8i?
They are Google's eighth-generation TPUs, announced in April 2026. TPU 8t targets training, with 9,600 chips per superpod. TPU 8i targets inference and reinforcement learning, with 288 GB of HBM and 384 MB of on-chip SRAM. Google says both arrive later in 2026.
Can I buy a TPU?
Not as a card for your own server in the way you can buy many GPUs. You rent TPU capacity through Google Cloud, using its Kubernetes or Compute Engine services.
Is an ASIC always better than a GPU for AI?
No. An ASIC is faster and cheaper only on the work it was designed for. If the model architecture changes, a fixed-function chip can lose its advantage, which is why GPU flexibility still matters.