LLM INFERENCE SCORE000000 CLEARED00/15

REFERENCE

GLOSSARY

164 terms from across all sixteen levels, each linked back to where it is explained.

acceptance rate (alpha)
LEVEL 08
Probability a drafted token survives verification. Equals sum min(p, q), i.e. 1 minus the total variation distance between draft and target.
achieved bandwidth
LEVEL 04
The fraction of peak HBM bandwidth a real kernel reaches. Typically 70-85%.
activation ratio
LEVEL 11
Active parameters divided by total parameters. DeepSeek-V3 is 5.5%, Qwen3-235B-A22B is 9%, Mixtral-8x7B is 28%.
active parameters
LEVEL 11
The parameters read to produce one token. Equal to the total in a dense model; a small fraction of it in a sparse one.
adapter (Houlsby adapter)
LEVEL 14
A small bottleneck feed-forward block inserted inline into each transformer sublayer. Cannot be merged away, so it adds latency at every inference call.
admission control
LEVEL 05
Deciding how many sequences to admit given that their eventual lengths, and therefore memory needs, are unknown.
all-reduce
LEVEL 09
Collective that sums a tensor across all GPUs and returns the result to all. Ring implementation moves 2(N-1)/N of the data per GPU.
answer evidence
LEVEL 15
Whether a downstream probe shows the information a cached context should have carried, independent of whether the model's exact next-token continuation matches — the "memory" axis, distinct from trajectory fidelity.
arithmetic intensity
LEVEL 01
FLOPs performed per byte moved from memory. Compare it against the hardware ratio to find your bottleneck.
attention hybrid
LEVEL 07
A stack mixing mostly SSM or linear-attention layers with a few full-attention layers, trading some memory saving for exact recall.
attention sink
LEVEL 06
The first few tokens of a sequence, which receive large attention weight regardless of content. Keeping them enables indefinite streaming.
automatic prefix caching
LEVEL 12
Resolving a new request’s prefix to existing physical blocks via a hash chained over each block’s tokens and its predecessor’s hash, without the request explicitly declaring a shared prefix.
Activation-aware weight quantization: scale up the ~1% of channels identified as salient from activation magnitude, then quantize.
batch invariance
LEVEL 03
The property that a kernel produces bit-identical results regardless of batch composition. Required for true determinism; costs performance.
block table
LEVEL 06
Per-sequence map from logical token positions to physical KV cache blocks. What allows non-contiguous allocation.
bonus token
LEVEL 08
The extra token available when all gamma drafts are accepted, since the verification pass already computed p at the final position.
bytes per token
LEVEL 02
2 x n_layers x n_kv_heads x head_dim x dtype_bytes. The most useful derived number in inference planning.
capacity factor
LEVEL 09
A cap on tokens routed to any single MoE expert, bounding load imbalance at the cost of dropping or rerouting overflow.
catastrophic forgetting
LEVEL 14
A model loses previously-held capabilities while being fine-tuned on new data. PEFT methods generally cause less of it than full fine-tuning does.
chance-normalized task-quality retention
LEVEL 15
(mapped_score − chance_floor) / (oracle_score − chance_floor) — a benchmark-accuracy comparison that accounts for different multiple-choice baselines, making retention comparable across benchmarks.
chunked prefill
LEVEL 05
Splitting a long prefill into fixed-size pieces processed one per iteration, mixed with ongoing decode.
closed-loop load
LEVEL 10
Fixed concurrency; a new request is sent only when one finishes. Self-throttling, so it hides the capacity cliff.
column-parallel / row-parallel
LEVEL 09
Megatron's arrangement: split the first matmul by columns and the second by rows, so the nonlinearity between them needs no communication.
constrained decoding
LEVEL 03
Masking logits each step so only tokens consistent with a grammar or schema can be sampled.
content space
LEVEL 15
A source model's K/V representation with RoPE's positional rotation removed, leaving only position-independent information about the token — the space a cross-model map is fit in.
continuous batching
LEVEL 05
Iteration-level scheduling: evict finished sequences and admit waiting ones at every decode step. From Orca.
copy-on-write
LEVEL 06
Sequences sharing a prefix point at the same physical blocks; a block is duplicated only when one of them writes to it.
copy-on-write (KV cache)
LEVEL 12
Sharing a physical block between sequences via a reference count, and copying it only at the moment one sequence’s write would make it diverge from another’s.
cross-attention
LEVEL 11
Decoder attention over an encoder output. Its K and V are computed once per input and never grow, unlike self-attention KV.
cross-layer KV sharing
LEVEL 11
Letting neighbouring layers reuse one set of K and V, cutting n_layers in the cache formula rather than the head count.
cross-model KV cache transfer
LEVEL 15
Reusing a KV cache built by one model to skip or shorten prefill on a different, larger model in the same family, via a learned coordinate mapping rather than a direct copy.
crossover point
LEVEL 02
The number of cached tokens in flight at which KV traffic overtakes weight traffic. About 114k tokens for Llama-3-8B at fp16.
CUDA Graph
LEVEL 01
A captured, replayable sequence of kernel launches. Removes per-launch CPU overhead, which is significant when a decode step is hundreds of tiny kernels.
decode
LEVEL 01
The per-token forward passes that follow prefill, one position at a time. Memory-bound; sets TPOT.
decode context parallelism (DCP)
LEVEL 09
Sharding the KV cache by token position across ranks so per-GPU cache keeps shrinking past the KV-head count. One extra collective per layer: AllGather the query, attend locally, combine the partials by their log-sum-exp.
degeneration
LEVEL 03
The repetitive-loop failure mode of likelihood-maximizing decoders. Named and diagnosed by Holtzman et al.
dense vs sparse FLOPs
LEVEL 04
Vendors quote 2:4 structured sparsity figures that are double the dense number. Production LLMs use dense.
disaggregation
LEVEL 09
Running prefill and decode on separate machine pools, transferring the KV cache between them.
Weight-Decomposed LoRA. Splits a frozen weight into magnitude and direction, trains both, and still merges for free — usually a small quality step up from plain LoRA.
Direct Preference Optimization — reparameterizes RLHF's objective so the reward model cancels out, turning preference learning into a classification loss over the policy's own log-probabilities.
draft model
LEVEL 08
A small, cheap model that proposes candidate tokens for the target model to verify.
driver process
LEVEL 12
In a multi-GPU vLLM deployment, the single process holding the scheduler and authoritative block tables, which broadcasts each iteration’s decisions to the worker processes.
Self-speculation by autoregressing over the target's hidden features rather than tokens. The strongest general method.
expert parallelism (EP)
LEVEL 09
Distributing MoE experts across GPUs. Requires all-to-all routing and suffers from load imbalance.
external fragmentation
LEVEL 02
Free memory that exists but is broken into pieces too small to satisfy a contiguous allocation.
fine-grained experts
LEVEL 11
Many small experts instead of a few large ones, so routing is finer-grained at the same active parameter count. The DeepSeekMoE design.
FlashAttention
LEVEL 07
Tiled exact attention that never materializes the N x N score matrix in HBM. Cuts HBM accesses to O(N^2 d^2 / M).
FlashDecoding
LEVEL 07
Decode-phase attention that parallelizes along the KV dimension (split-K) because query-length-1 offers no query parallelism.
frequency penalty
LEVEL 03
Subtracts a term proportional to how many times a token has appeared. Additive in logit space, better behaved than division.
Draft length: how many tokens are proposed before each verification pass.
goodput
LEVEL 04
Throughput counting only requests that met their latency SLO. The metric that matches what you are actually selling.
Post-training weight quantization using approximate second-order information, layer by layer. Works to 3-4 bits.
Grouped-query attention. Fewer K/V heads than Q heads, with each K/V head shared by a group of query heads. Shrinks the KV cache proportionally.
Group Relative Policy Optimization — drops PPO's critic network by scoring each sampled completion's advantage against the mean/std of a group sampled for the same prompt.
head_dim
LEVEL 00
The width of one attention head. Usually d_model / n_heads, though it can be set independently.
hidden state
LEVEL 00
One d_model-length slice of the residual stream, at one position and one layer.
ICI (Inter-Chip Interconnect)
LEVEL 13
The direct chip-to-chip links forming a TPU pod's 2D or 3D torus. Around 100 GB/s per link on a v5p, and the same bandwidth whether the pod has 64 chips or thousands.
image tokens
LEVEL 11
The positions a vision encoder and projector contribute to the language model sequence. 576 for LLaVA-1.5, thousands under tiling schemes.
internal fragmentation
LEVEL 02
Memory reserved inside a sequence allocation that the sequence never uses. The dominant waste in naive allocators.
iteration-level scheduling
LEVEL 12
Scheduling once per forward pass rather than once per request, so a sequence can join or leave the batch on any step. Introduced by Orca; the mechanism continuous batching runs on.
k (source-layer width)
LEVEL 15
How many top-predictive source layers are concatenated as input to one target layer's mapper. Larger k improves output-distribution fidelity (prefix NLL) but costs more to compute and can hurt intermediate-representation fidelity past a point.
kernel fusion
LEVEL 07
Combining adjacent operations into one kernel so intermediates never round-trip to HBM.
KL divergence between a mapped-cache continuation and the native-prefill continuation, measured over a 32-step forced rollout rather than a single token — designed to catch drift that single-step agreement misses.
The batch size beyond which throughput gains become small relative to the latency cost. The right default operating point.
KV cache
LEVEL 02
The stored key and value tensors for all previous positions, kept so they are not recomputed each step. Per-sequence, and it grows with every token.
linear attention
LEVEL 07
Replacing softmax with a kernel feature map so attention can be reassociated into O(N d^2). Constant state, weaker exact recall.
Little's Law
LEVEL 01
concurrency = arrival rate x latency. Tells you how many requests must be in flight to hit a throughput target.
llama.cpp
LEVEL 10
Local-inference engine optimized for the weight-dominated regime: aggressive weight quantization, CPU and partial-offload support.
local-global interleaving
LEVEL 11
A stack of mostly windowed layers with occasional full-attention layers. Gemma 2 used 1:1 at 4096; Gemma 3 uses 5:1 at 1024.
log-sum-exp (LSE) combination
LEVEL 09
Merging attention outputs computed over disjoint slices of the sequence by reweighting them with each slice’s softmax normalizer. FlashAttention’s online softmax, applied across devices instead of across tiles.
logits
LEVEL 00
The unnormalized [.., vocab_size] scores produced by the LM head, before any softmax or sampling.
Low-Rank Adaptation. Freezes W0 and trains a rank-r product BA added beside it; merges back into W0 for zero added inference cost.
Medusa
LEVEL 08
Self-speculation with multiple independent decoding heads predicting several positions ahead, verified as a tree.
Keep tokens with probability at least min_p times the top token's probability. Fixed ratio, adaptive to confidence.
Multi-head latent attention. Compresses K and V into a shared low-rank latent per token; the up-projections are absorbed into neighbouring matrices.
model routing (cross-model transfer)
LEVEL 15
A deployment pattern where a cheap model handles a request by default and escalates to a larger model only when needed, carrying its KV cache forward instead of discarding the reading already done.
modified rejection sampling
LEVEL 08
Accept with probability min(1, p/q); on rejection sample from normalized max(0, p - q). Provably yields p.
Multi-query attention: a single KV head shared by all query heads. Maximum cache savings, largest quality cost.
A TPU's Matrix Multiply Unit: a 128×128 weight-stationary systolic array. A v5p TensorCore has four; a chip has two TensorCores.
n-gram lookup
LEVEL 08
Drafting by searching the context for the current suffix and proposing what followed it. Free, and excellent on repetitive text.
next-token top-1 agreement
LEVEL 15
Whether a target model's single most-likely next token, continuing from a transferred cache, matches what it would have produced from its own native prefill.
4-bit NormalFloat — a quantization data type whose bins are spaced to match a roughly-Gaussian weight distribution, rather than uniform INT4 bins.
NVLink
LEVEL 09
High-bandwidth GPU-to-GPU interconnect, ~900 GB/s on H100. Roughly 14x faster than PCIe Gen5.
occupancy
LEVEL 13
How many warps are resident on an SM relative to the maximum. It is the GPU's mechanism for hiding memory latency: with nothing else resident, a stall is an idle cycle.
online softmax
LEVEL 07
Computing softmax incrementally with a running max and sum, rescaling earlier partials when a larger value appears. Exact, not approximate.
open-loop load
LEVEL 10
Arrivals follow an external schedule (usually Poisson) regardless of server state. The only way to observe overload.
optical circuit switch (OCS)
LEVEL 13
The reconfigurable optical fabric connecting TPU v4 cubes, allowing a job's topology to be chosen at schedule time and failed cubes to be routed around.
PagedAttention
LEVEL 06
Fixed-size block allocation for the KV cache with per-sequence block tables, borrowed from OS virtual memory. Cuts fragmentation waste from 60-80% to under 4%.
partial recomputation (cross-model)
LEVEL 15
Recomputing a target model's last one or two layers exactly from a shared boundary, instead of mapping them, to eliminate the nonlinear-drift ceiling a pure map cannot cross.
Parameter-efficient fine-tuning — the family of methods that update a small fraction of parameters, or none of the original ones, instead of every weight.
per-channel quantization
LEVEL 06
A separate scale for each channel of a tensor. Necessary for keys, which have strong per-channel outliers.
pipeline bubble
LEVEL 09
Idle stage time while the pipeline fills and drains. Fraction is (N-1)/(M+N-1) for M microbatches and N stages.
pipeline fill (array)
LEVEL 13
The cycles before a systolic array's first result emerges, while the wavefront crosses it. Costs a matmul of m rows roughly m/(m+128) of peak on a 128-wide array.
pipeline parallelism (PP)
LEVEL 09
Assigning different layers to different GPUs. One activation hand-off per stage boundary. Tolerates slow links.
pre-norm
LEVEL 00
Normalizing the input to each block rather than its output, leaving the residual path unbroken. Required in practice for deep transformers.
preemption
LEVEL 05
Reclaiming a running sequence's KV memory, either by swapping it to host memory or discarding and recomputing it.
prefill
LEVEL 01
The forward pass over the entire prompt, done in one shot. Compute-bound; sets TTFT.
prefill/decode interference
LEVEL 05
A long compute-bound prefill blocking every decoding sequence, spiking their inter-token latency.
prefix caching
LEVEL 02
Reusing cached KV state across requests that share a leading token sequence. Exact, not approximate.
LoRA with the frozen base model stored in 4-bit NF4, dequantized on the fly per matmul — cuts base-weight storage 4x on top of what freezing already saves.
RadixAttention
LEVEL 02
SGLang's radix-tree index over cached prefixes, enabling automatic longest-prefix reuse across arbitrary sharing patterns.
rank (r) / alpha
LEVEL 14
r sets the dimensionality of a LoRA update; alpha scales it (the update is multiplied by alpha/r) so quality is stable as r changes.
rank-r residual adapter
LEVEL 15
A small trained correction added on top of a frozen ridge map, factored through a low-rank bottleneck (rank r) to keep its parameter count small — linear (BAz) or a factorized-quadratic, genuinely nonlinear variant.
reasoning workload
LEVEL 10
Requests that generate thousands of hidden chain-of-thought tokens, shifting nearly all wall-clock into memory-bound decode.
recomputation
LEVEL 12
Preemption that drops a sequence’s KV blocks and reruns prefill over its tokens-so-far when it is re-admitted, at the cost of redone compute.
recurrent state
LEVEL 11
The fixed-size per-sequence memory of an SSM or linear-attention layer. Constant in context length, so it scales with concurrency alone.
reference count
LEVEL 12
The count of sequences whose block table points at a given physical block, used to decide whether releasing a sequence actually frees the block.
repetition penalty
LEVEL 03
Divides logits of already-seen tokens by r > 1. Blunt, and harmful for code and structured output.
residual distribution
LEVEL 08
The normalized positive part of p - q. The mass the target wanted that the draft undersupplied.
residual stream
LEVEL 00
The [batch, seq, d_model] tensor that every block reads from and adds back into. Its width is constant through the network.
reward model
LEVEL 14
A model trained on pairwise human preference comparisons to output a scalar score, used as the reward signal in RLHF's PPO stage.
ridge mapping
LEVEL 15
A per-attention-head linear transform, fit by ridge regression (least squares with an L2 penalty) from source-model K/V vectors onto target-model K/V vectors, using an offline calibration set.
ridge point
LEVEL 01
The arithmetic intensity at which a machine transitions from memory-bound to compute-bound. About 295 FLOP/byte for an H100.
Reinforcement learning from human feedback: train a reward model on preference pairs, then run PPO against it with a KL penalty back to the SFT policy.
RL with verifiable rewards — reward comes from a deterministic checker (does the math answer match, do the tests pass) instead of a learned reward model.
RMSNorm
LEVEL 00
Normalization by root-mean-square only, with no mean subtraction and no bias. Cheaper than LayerNorm and works as well.
roofline model
LEVEL 04
attainable = min(peak FLOP/s, intensity x bandwidth). A two-line model that identifies your bottleneck.
Rotary position embedding. Rotates Q and K by an angle proportional to absolute position so their dot product depends only on relative distance.
router
LEVEL 11
The small linear layer in an MoE that scores every expert for a token, so the top k can be selected. Runs per token, per layer.
A long-context benchmark testing multi-needle retrieval, aggregation and tracing. Far more informative than perplexity for evaluating KV compression.
selective batching
LEVEL 05
Batch the position-independent operations across sequences while handling attention per-sequence. What makes continuous batching implementable.
self-speculation
LEVEL 08
Generating drafts from the target model itself via extra heads or a small feature-level head, avoiding a separate draft model.
sequence / context parallelism
LEVEL 09
Splitting the sequence dimension across devices, as in Ring Attention. Used when one sequence exceeds a single GPU.
SFT (supervised fine-tuning)
LEVEL 14
Fine-tuning a pretrained model on curated (instruction, response) pairs with the ordinary next-token loss, masked so gradient only flows through the response tokens.
SGLang
LEVEL 10
Serving engine organized around structured, prefix-sharing workloads: RadixAttention plus a frontend programming language.
shape bucketing
LEVEL 13
Quantizing dynamic dimensions (batch, context length) to a small set of values and padding up, so a serving loop needs only a few cached compiled programs instead of one per distinct shape.
shared expert
LEVEL 11
An expert every token reads, alongside its routed ones. Absorbs the common knowledge that would otherwise be duplicated into every expert.
shared memory / SRAM
LEVEL 07
On-chip memory, ~228 KB per SM on Hopper, with roughly 6x the bandwidth of HBM. The tier FlashAttention keeps its tiles in.
Single instruction, multiple threads — the GPU execution model where a warp of 32 threads issues one instruction together, with divergent branches serialized under lane masks.
sliding-window attention
LEVEL 11
Attending only to the last w positions, which makes the cache constant in context rather than linear. Bounds the cache, not the receptive field.
slot mapping
LEVEL 12
The per-token mapping from a batch position to a physical KV-cache slot, which the attention kernel reads through to gather scattered K and V.
SmoothQuant
LEVEL 06
Migrates activation outliers into the weights via per-channel scaling, enabling int8 for both operands.
sparse upcycling
LEVEL 11
Turning a dense checkpoint into an MoE by copying its MLP into every expert and continuing training. Costs roughly half the original pretraining.
split-K
LEVEL 07
Partitioning a reduction across parallel workers and combining their partial results afterward.
State space model. Maintains a fixed-size recurrent state, so there is no KV cache; Mamba is the leading example.
static batching
LEVEL 05
Fixed batch that runs to completion. Utilization collapses when sequence lengths vary, and gets worse as batch grows.
streaming multiprocessor (SM)
LEVEL 13
The GPU's independent execution unit: warp schedulers, registers, shared memory and tensor cores. An H100 SXM has 132 of them.
Preemption that copies a sequence’s KV blocks to host memory and back, preserving the cache at the cost of a PCIe round trip.
SwiGLU
LEVEL 00
The gated feed-forward block used by Llama-family models: down(silu(gate(x)) * up(x)). Three weight matrices instead of two.
systolic array
LEVEL 13
A mesh of multiply-accumulate cells wired to their neighbours, through which operands flow and partial sums accumulate. One memory read feeds many multiplications, which is where its energy advantage comes from.
teacher forcing
LEVEL 00
Training on the known target sequence so all positions are computed in one parallel pass. The reason training is far more hardware-efficient than generation.
temperature
LEVEL 03
Divisor applied to logits before softmax. Below 1 sharpens, above 1 flattens.
tensor core
LEVEL 13
The matrix-multiply unit inside an SM, where essentially all of a GPU's FLOPs live. Not a systolic array in the TPU sense, despite the name.
tensor parallelism (TP)
LEVEL 09
Splitting weight matrices within a layer across GPUs. Two all-reduces per transformer layer. Needs NVLink.
TensorRT-LLM
LEVEL 10
NVIDIA serving engine that compiles a model into a hardware-specific engine ahead of time. Fastest and least flexible.
tied embeddings
LEVEL 11
Reusing the input embedding matrix as the LM head. Saves vocab x d_model parameters — rounding error at 70B, a quarter of a 0.6B model.
TMA (Tensor Memory Accelerator)
LEVEL 13
Hopper's asynchronous DMA engine for moving tiles between global and shared memory without spending warps on address arithmetic.
token budget
LEVEL 05
The cap on tokens processed per scheduler iteration. Decode tokens are admitted first, prefill chunks fill the rest.
Keep the k highest-probability tokens. Fixed count, does not adapt to the distribution.
top-k routing
LEVEL 11
Sending each token to the k highest-scoring experts. Switch Transformer uses k=1, GShard and Mixtral k=2, DeepSeek-V3 k=8 of 256.
top-p / nucleus
LEVEL 03
Keep the smallest set whose cumulative probability reaches p. Fixed mass, adaptive count.
TPOT / ITL
LEVEL 01
Time per output token / inter-token latency. The gap between successive tokens once streaming starts.
tree attention
LEVEL 08
Verifying many candidate continuations in one pass using a mask where each node attends only to its ancestors.
Time to first token. How long the user waits before anything appears.
Serving engine organized around memory management: PagedAttention, continuous batching, broad model support.
A TPU's on-chip scratchpad, roughly 128 MiB on recent generations, entirely managed by the compiler. There is no cache behind it.
The TPU's vector unit, an (8, 128) lane array that handles the elementwise work — softmax, normalization, activations, residual adds — that the MXU cannot.
The 32-thread group a GPU schedules as a unit. Hopper adds the warpgroup — four warps issuing one wgmma matrix instruction together.
warp specialization
LEVEL 07
Assigning different warps within a block to different roles (producer/consumer) so data movement overlaps computation. Central to FlashAttention-3.
weight-stationary
LEVEL 13
A dataflow in which a tile of weights is loaded into the array and held while activations stream through, rather than both operands moving.
width-pruned sibling
LEVEL 15
A smaller model derived from a larger one by pruning (reducing width, depth, attention or MLP dimensions) and distillation-based retraining — e.g. NVIDIA's Minitron family — structurally closer to its parent than an independently-trained model of the same size.
worker process
LEVEL 12
One per GPU in a multi-GPU deployment; holds a shard of the model and executes the batch the driver assigned, exchanging activations with other workers via NCCL under tensor parallelism.
workspace
LEVEL 02
GPU memory reserved for activations, temporary buffers, and communication. Typically 10-20 GB, and easy to forget when budgeting.
The compiler that turns a JAX or TensorFlow graph into a statically scheduled TPU program. It specializes on tensor shapes, which is why shape variety costs compilations.