1. Backing the clock out of the geometry
Each MXU retires 128 × 128 = 16,384 MACs per cycle, and a MAC is 2 FLOP:
v4: 275e12 / (8 x 16384 x 2) = 1.049 GHz
v5e: 197e12 / (4 x 16384 x 2) = 1.503 GHz
v5p: 459e12 / (8 x 16384 x 2) = 1.751 GHz
The v4 answer is the interesting one: its documented clock is 1,050 MHz, so the model closes to three digits. That is a real check, not a coincidence — it confirms the MXU count and array size.
v6e: 918e12 / (8 x 16384 x 2) = 3.502 GHz
3.5 GHz is not a plausible clock for a datacenter accelerator, so the assumption must be wrong: the v6e cannot have the same 8 × 128×128 geometry. Either it has more MXUs or larger ones. This is the most useful thing back-of-envelope arithmetic does — not confirming what you believed, but telling you precisely which belief is broken.
2. Ridge points
H100: 989e12 / 3.35e12 = 295
v4: 275e12 / 1.20e12 = 229
v5e: 197e12 / 0.819e12 = 241
v5p: 459e12 / 2.765e12 = 166
The v5p is the best-balanced part in this set: it needs an arithmetic intensity of 166 to saturate, where the H100 needs 295. Since decode intensity equals batch size, that is a statement that the v5p reaches its own ceiling at a batch of 166 while the H100 needs 295 to reach its. It says nothing about which ceiling is higher — the H100's is 2.2× the v5p's. A low ridge point means "easy to saturate," not "fast."
3. How many chips before it runs, and how fast then
weights = 70.6e9 x 2 = 141.2 GB
H100 (80 GB): ceil(141.2/80) = 2 chips -> 70.6 GB per chip / 3.35 TB/s = 21.07 ms (47 tok/s)
v5e (16 GB): ceil(141.2/16) = 9 chips -> 15.7 GB per chip / 0.819 TB/s = 19.16 ms (52 tok/s)
v5p (95 GB): ceil(141.2/95) = 2 chips -> 70.6 GB per chip / 2.765 TB/s = 25.53 ms (39 tok/s)
at 8 chips:
H100: 17.65 GB / 3.35 TB/s = 5.27 ms (190 tok/s)
v5e: 17.65 GB / 0.819 TB/s = 21.55 ms ( 46 tok/s)
v5p: 17.65 GB / 2.765 TB/s = 6.38 ms (157 tok/s)
Two things to notice. First, the minimum-fit configuration is nobody's answer — you shard past the capacity requirement because sharding buys aggregate bandwidth, which is the actual currency of decode. Second, the v5e needs 9 chips before it can run at all and is still slower on 8 than an H100 is on 8: it is a small, cheap, throughput-oriented part, and a 70B model in fp16 is simply not what it is for. That is a statement about part selection, not about TPUs.
4. Array fill
m = 1: 1/129 = 0.8%
m = 8: 8/136 = 5.9%
m = 128: 128/256 = 50.0%
m = 512: 512/640 = 80.0%
m = 2048: 2048/2176 = 94.1%
90%: m/(m+128) = 0.9 -> m = 9 x 128 = 1,152 tokens
Set that against the v5p's ridge point of 166. To be memory-bound-free you need a batch of 166; to run the array near its peak you need a batch of 1,152. Under this (deliberately pessimistic, un-pipelined) model the geometry binds long after the roofline does — so on a systolic machine the question "what batch do I need?" has two answers and you must take the larger. Real MXUs overlap the fill of one weight tile with the drain of the previous one, which pulls the second number down substantially; the code lab leaves that as an exercise precisely because how much it helps is the whole argument about how much the fill really costs.
5. Scratchpad per unit of compute
H100: 228 KB x 132 = 30.8 MB -> 30.8e6 / 989 = 31.1 KB per TFLOP/s
v5p: 128 MiB = 134.2 MB -> 134.2e6 / 459 = 292.4 KB per TFLOP/s
ratio: 9.4x
The prediction is exactly what the two ecosystems look like. On a GPU, fitting attention into 228 KB per SM required an algorithm — tile, recompute the softmax online, never materialise the N×N matrix — and that algorithm has a name and a citation. On a TPU, keeping a layer's activations on chip across several fused operations is a scheduling decision the compiler makes, because there is room. Same idea, one is a paper and one is a pass.
6. Padding, twice
257 tokens -> ceil(257/128) x 128 = 384 tiles' worth of arithmetic
wasted: 127/384 = 33.1%
A batch of 257 costs exactly what a batch of 384 costs. Sizing your server's step to land just past a tile boundary is one of the easiest large wins available, and one of the easiest to miss because nothing reports it.
50,257 -> 50,304 (= 128 x 393), extra 47 columns = 0.093% more LM-head FLOPs
The padding itself is negligible — yet padding the vocabulary is a well-known and large speedup. The reason is that the extra 0.093% is not what you are buying. An unaligned dimension pushes the matmul off the library's fast path: it selects a different, slower kernel, or falls back to one that does not use the tensor cores efficiently at all. You are not saving the 0.093%; you are buying back the fast kernel. Whenever a tiny alignment change produces a large speedup, that is the shape of the explanation.
7. Collectives at TP-8 and TP-64
per all-reduce, TP-8: 2 x 32 x 1 x 8192 x 2 x (7/8) = 0.918 MB
per token: x 2 all-reduces x 80 layers = 146.8 MB
per all-reduce, TP-64: 2 x 32 x 1 x 8192 x 2 x (63/64) = 1.032 MB
per token: = 165.2 MB
TP-8, NVLink 900 GB/s: 0.163 ms ~3% of a 5.27 ms step free
TP-64, ICI 600 GB/s: 0.275 ms ~5% free
TP-64, InfiniBand 50 GB/s: 3.30 ms ~63% fatal
TP-64 is unremarkable on a torus and unusable across a commodity network, and the reason is not link speed — a single NVLink domain is faster per link than ICI. It is reach: the TPU's fast tier extends to the whole pod, so "how many chips can I shard a layer across?" has a different answer on each machine. Note also that this compares bandwidth only; ring all-reduce latency grows with hop count, so the torus's diameter shows up as a per-collective fixed cost the bandwidth model here ignores — one more reason the real crossover is workload-specific.
8. What changes at prefill
The ridge points (2) do not change — they are properties of the chip. The array-fill numbers (4) change completely and in the TPU's favour: prefill processes thousands of tokens at once, so m is large, array_eff is near 1, and the geometry penalty that dominates decode disappears. The decode floors (3) are replaced by a compute-bound calculation entirely. The collective volumes (7) scale with the number of tokens in the batch, so they grow by three orders of magnitude and stop being free — which is why prefill is where interconnect actually gets tested. In short: decode measures your memory system, prefill measures your arithmetic and your network. The two chips are nearly identical under the first measurement and quite different under the second, which is the entire content of this level in one sentence.