The decode-step bill: active bytes and cache bytes
Module 1 established that a decode step reads the weights once to produce one token. Module 2 added the cache that grows with context. Module 4 turned both into time. Put the three together and one line prices any architecture:
bytes per decode step ~= W_active + B x S x kv_per_token
step time ~= bytes / achieved bandwidth
W_active is the weight bytes actually read, B is the number of sequences in flight, S is their mean length. The asymmetry between the two terms is the thing to internalise: the weight term is paid once per step no matter how many sequences you are serving, and the cache term is paid once per sequence. That is the entire reason batching works, and it is why the two terms have to be tracked separately rather than added into a single "model size".
For Llama-3-8B at fp16, the two numbers are 15.0 GB and 128 KiB. At batch 1 and 8k context:
weights 15.0 GB
cache 8192 x 131072 = 1.07 GB
---------
16.1 GB -> 4.8 ms on an H100 -> 208 tok/s
At batch 64 the weight term does not move and the cache term multiplies by 64, to 68.7 GB. The model has not changed. Which term you are paying has.
Here is the whole level in one table — four families, four different bets:
| family | what it changes | term attacked | what it costs |
|---|---|---|---|
| Dense transformer | nothing; the baseline | neither | it is the thing everything else is measured against |
| Mixture of experts | one MLP becomes E MLPs, k read per token | W_active | memory capacity — every expert must be resident |
| MQA · GQA · MLA · cross-layer sharing | fewer or smaller cached vectors | kv_per_token | quality, in amounts that vary from negligible to real |
| Sliding window · local-global | old tokens fall out of the window | S in most layers | exact recall of distant tokens on those layers |
| SSM · linear attention · hybrid | a fixed-size state replaces the cache | S entirely, in most layers | recall again, harder — Module 7 has the argument |
Two warnings before the details. First, the axis an architecture does not attack tends to get worse, because designers spend the savings: MoE models are enormous, MLA models are dense, and hybrids that keep a few attention layers keep a cache for them. Second, quantization scales both terms and is not on this list, because it is a deployment choice rather than an architectural one. Module 6 covers it, and the two compose — a fp8 MoE moves half of a small number.
ONE DECODE STEP, batch B, mean context S
┌──────────────────────────────┬──────────────────────────────────┐
│ W_active │ B x S x kv_per_token │
│ read once per STEP │ read once per SEQUENCE │
│ independent of B and S │ grows with BOTH │
└──────────────────────────────┴──────────────────────────────────┘
^ ^
│ │
mixture of experts GQA · MLA · sliding window
cuts this term · recurrent state cut this one
│ │
pays in CAPACITY pays in RECALL
REMEMBEREvery architectural choice lands in one of two terms: weight bytes read per step, or cache bytes read per token per sequence.