Discussion about this post

User's avatar
Prasanna Narayanan's avatar

I had no interest in this topic but once I started reading through I couldn't stop. What an fantastic way to explain complex concept! Truly one of the best newsletter I subscribe to.

Latent Dynamics's avatar

Moving bytes across a printed circuit board consumes vastly more energy than doing the actual arithmetic. That's the physical truth software layers constantly try to hide. ⚡

When a CPU fetches an instruction, reads data from DRAM, and writes back a result, it burns thermal headroom on the walk rather than the work. A systolic array changes the geometry completely. You freeze the weights directly inside on-chip registers. Inputs stream horizontally, partial sums accumulate vertically, and memory gets touched once. Math becomes a spatial pulse through physical silicon gates. 🌊

Lower precision isn't just about compression. Dropping from FP32 to BFloat16 or native FP8 cuts multiplier silicon area quadratically. That lets you pack 65,536 execution units onto a single die without melting the substrate. But there's a deeper bottleneck lurking right behind raw compute density. 🔬

We hypothesize that the ultimate scaling barrier for multi-pod AI isn't arithmetic throughput, but the topological phase mismatch between deterministic on-chip systolic dataflow and asynchronous network packet arbitration. When thousands of deterministic matrix waves hit traditional packet-switched buffers, queueing jitter wrecks step synchronization. Replacing packet switches with MEMS-driven optical circuit switches in a 3D torus eliminates buffer bloat and locks the entire multi-chip cluster into a single, continuous physical pipeline. 🌐

If your silicon executes math with clockwork spatial determinism, why are you still letting non-deterministic packet fabrics orchestrate your cluster's all-reduce collectives? ⚙️

(⊙_◎)

5 more comments...

No posts

Ready for more?