In this article, we look at why Google built custom silicon, and how it works, revealing the physical constraints and engineering trade-offs they had to make.
I had no interest in this topic but once I started reading through I couldn't stop. What an fantastic way to explain complex concept! Truly one of the best newsletter I subscribe to.
CPUs spend most of their thermal budget hauling bytes across a bus line rather than doing math. That's the Von Neumann tax in plain English. Google's TPU array doesn't move weights during calculation. Weights stay frozen inside 65,536 registers while activations stream horizontally and partial sums drop vertically. Memory gets touched once. Math happens thousands of times on the spatial move.
Here's where hardware silicon is heading next. As precision drops from FP32 down to BFloat16 and native FP8, compute density outpaces off-chip wire bandwidth completely. We aren't just building faster multipliers anymore. We're collapsing the boundary between storage and execution. The systolic array turns compute into a continuous wave flowing through physical logic gates without waiting for instruction fetches. When you link thousands of these arrays over optical circuit switches in a 3D torus, packet queueing jitter vanishes. The entire cluster becomes one giant physical execution grid.
If your network fabric still relies on probabilistic packet switching while your silicon runs deterministic matrix waves, where do you think your real straggler latency is hiding?
If the evolution of custom AI hardware like Google’s TPU accelerates both the scale and ubiquity of large language models, how might that reshape human behavior? I think Americans can already be very demanding, yet this might normalize expectations of instantaneous, machine-mediated reasoning and decision-making in everyday life. Will people eventually select regularly from several machine-generated conclusions?
I had no interest in this topic but once I started reading through I couldn't stop. What an fantastic way to explain complex concept! Truly one of the best newsletter I subscribe to.
CPUs spend most of their thermal budget hauling bytes across a bus line rather than doing math. That's the Von Neumann tax in plain English. Google's TPU array doesn't move weights during calculation. Weights stay frozen inside 65,536 registers while activations stream horizontally and partial sums drop vertically. Memory gets touched once. Math happens thousands of times on the spatial move.
Here's where hardware silicon is heading next. As precision drops from FP32 down to BFloat16 and native FP8, compute density outpaces off-chip wire bandwidth completely. We aren't just building faster multipliers anymore. We're collapsing the boundary between storage and execution. The systolic array turns compute into a continuous wave flowing through physical logic gates without waiting for instruction fetches. When you link thousands of these arrays over optical circuit switches in a 3D torus, packet queueing jitter vanishes. The entire cluster becomes one giant physical execution grid.
If your network fabric still relies on probabilistic packet switching while your silicon runs deterministic matrix waves, where do you think your real straggler latency is hiding?
(⊙_⊙)
Jax seems the only native option to operate tpus but it’s still less mature than PyTorch. I wish PyTorch can merge xla as one of its backend.
If the evolution of custom AI hardware like Google’s TPU accelerates both the scale and ubiquity of large language models, how might that reshape human behavior? I think Americans can already be very demanding, yet this might normalize expectations of instantaneous, machine-mediated reasoning and decision-making in everyday life. Will people eventually select regularly from several machine-generated conclusions?
Ty for writing this
A hot topic. Thanks for tackling it!