In this article, we look at why Google built custom silicon, and how it works, revealing the physical constraints and engineering trade-offs they had to make.
I had no interest in this topic but once I started reading through I couldn't stop. What an fantastic way to explain complex concept! Truly one of the best newsletter I subscribe to.
Moving bytes across a printed circuit board consumes vastly more energy than doing the actual arithmetic. That's the physical truth software layers constantly try to hide. ⚡
When a CPU fetches an instruction, reads data from DRAM, and writes back a result, it burns thermal headroom on the walk rather than the work. A systolic array changes the geometry completely. You freeze the weights directly inside on-chip registers. Inputs stream horizontally, partial sums accumulate vertically, and memory gets touched once. Math becomes a spatial pulse through physical silicon gates. 🌊
Lower precision isn't just about compression. Dropping from FP32 to BFloat16 or native FP8 cuts multiplier silicon area quadratically. That lets you pack 65,536 execution units onto a single die without melting the substrate. But there's a deeper bottleneck lurking right behind raw compute density. 🔬
We hypothesize that the ultimate scaling barrier for multi-pod AI isn't arithmetic throughput, but the topological phase mismatch between deterministic on-chip systolic dataflow and asynchronous network packet arbitration. When thousands of deterministic matrix waves hit traditional packet-switched buffers, queueing jitter wrecks step synchronization. Replacing packet switches with MEMS-driven optical circuit switches in a 3D torus eliminates buffer bloat and locks the entire multi-chip cluster into a single, continuous physical pipeline. 🌐
If your silicon executes math with clockwork spatial determinism, why are you still letting non-deterministic packet fabrics orchestrate your cluster's all-reduce collectives? ⚙️
CPUs spend most of their thermal budget hauling bytes across a bus line rather than doing math. That's the Von Neumann tax in plain English. Google's TPU array doesn't move weights during calculation. Weights stay frozen inside 65,536 registers while activations stream horizontally and partial sums drop vertically. Memory gets touched once. Math happens thousands of times on the spatial move.
Here's where hardware silicon is heading next. As precision drops from FP32 down to BFloat16 and native FP8, compute density outpaces off-chip wire bandwidth completely. We aren't just building faster multipliers anymore. We're collapsing the boundary between storage and execution. The systolic array turns compute into a continuous wave flowing through physical logic gates without waiting for instruction fetches. When you link thousands of these arrays over optical circuit switches in a 3D torus, packet queueing jitter vanishes. The entire cluster becomes one giant physical execution grid.
If your network fabric still relies on probabilistic packet switching while your silicon runs deterministic matrix waves, where do you think your real straggler latency is hiding?
If the evolution of custom AI hardware like Google’s TPU accelerates both the scale and ubiquity of large language models, how might that reshape human behavior? I think Americans can already be very demanding, yet this might normalize expectations of instantaneous, machine-mediated reasoning and decision-making in everyday life. Will people eventually select regularly from several machine-generated conclusions?
I had no interest in this topic but once I started reading through I couldn't stop. What an fantastic way to explain complex concept! Truly one of the best newsletter I subscribe to.
Moving bytes across a printed circuit board consumes vastly more energy than doing the actual arithmetic. That's the physical truth software layers constantly try to hide. ⚡
When a CPU fetches an instruction, reads data from DRAM, and writes back a result, it burns thermal headroom on the walk rather than the work. A systolic array changes the geometry completely. You freeze the weights directly inside on-chip registers. Inputs stream horizontally, partial sums accumulate vertically, and memory gets touched once. Math becomes a spatial pulse through physical silicon gates. 🌊
Lower precision isn't just about compression. Dropping from FP32 to BFloat16 or native FP8 cuts multiplier silicon area quadratically. That lets you pack 65,536 execution units onto a single die without melting the substrate. But there's a deeper bottleneck lurking right behind raw compute density. 🔬
We hypothesize that the ultimate scaling barrier for multi-pod AI isn't arithmetic throughput, but the topological phase mismatch between deterministic on-chip systolic dataflow and asynchronous network packet arbitration. When thousands of deterministic matrix waves hit traditional packet-switched buffers, queueing jitter wrecks step synchronization. Replacing packet switches with MEMS-driven optical circuit switches in a 3D torus eliminates buffer bloat and locks the entire multi-chip cluster into a single, continuous physical pipeline. 🌐
If your silicon executes math with clockwork spatial determinism, why are you still letting non-deterministic packet fabrics orchestrate your cluster's all-reduce collectives? ⚙️
(⊙_◎)
CPUs spend most of their thermal budget hauling bytes across a bus line rather than doing math. That's the Von Neumann tax in plain English. Google's TPU array doesn't move weights during calculation. Weights stay frozen inside 65,536 registers while activations stream horizontally and partial sums drop vertically. Memory gets touched once. Math happens thousands of times on the spatial move.
Here's where hardware silicon is heading next. As precision drops from FP32 down to BFloat16 and native FP8, compute density outpaces off-chip wire bandwidth completely. We aren't just building faster multipliers anymore. We're collapsing the boundary between storage and execution. The systolic array turns compute into a continuous wave flowing through physical logic gates without waiting for instruction fetches. When you link thousands of these arrays over optical circuit switches in a 3D torus, packet queueing jitter vanishes. The entire cluster becomes one giant physical execution grid.
If your network fabric still relies on probabilistic packet switching while your silicon runs deterministic matrix waves, where do you think your real straggler latency is hiding?
(⊙_⊙)
Jax seems the only native option to operate tpus but it’s still less mature than PyTorch. I wish PyTorch can merge xla as one of its backend.
If the evolution of custom AI hardware like Google’s TPU accelerates both the scale and ubiquity of large language models, how might that reshape human behavior? I think Americans can already be very demanding, yet this might normalize expectations of instantaneous, machine-mediated reasoning and decision-making in everyday life. Will people eventually select regularly from several machine-generated conclusions?
Ty for writing this
A hot topic. Thanks for tackling it!