Executive Summary / Introduction
Most people think ChatGPT operates as a simple, direct sequence: a user sends a prompt, the chat interface passes it to an AI model, and the model returns an answer. That mental model is useful for explaining the product, but it is almost useless for understanding the engineering. A production system capable of serving ChatGPT at global scale must solve a complex hierarchy of distributed systems challenges long before a model can generate a single useful token.
user request
│
▼
product layer
│
▼
request processing
│
▼
model / workload selection
│
▼
inference routing
│
▼
GPU execution
│
▼
token generation
│
▼
streaming
│
▼
product responseUnderneath this visible request path operate dozens of critical subsystems responsible for context management, authentication, tool execution, KV caching, model scheduling, GPU utilization, distributed inter-node communication, safety filters, observability, and agent orchestration. OpenAI has publicly described its modern inference stack as a system holistically optimized across routing, scheduling, kernels, caching, model architecture, and hardware utilization rather than treating inference as an isolated machine learning problem. Serving over 1 billion active users and more than 2 million enterprise businesses transforms the engineering equation entirely.
A model that is impressive in a benchmark is not necessarily economical to serve. A GPU that is fast in isolation is ineffective if scheduling leaves it idle. And an inference engine that generates tokens quickly for one request can collapse when thousands of concurrent requests with volatile context lengths arrive simultaneously. The central engineering challenge is therefore turning a very expensive probabilistic computation into a reliable, low-latency, and cost-effective utility at planetary scale.
ChatGPT Is Not the Model
The first and most critical architectural distinction is separating the product from the foundation model. ChatGPT is a complex software product; a model such as GPT-4 or GPT-5.6 is merely one computational engine behind that product.
ChatGPT (Product Layer)
│
├── user interface
├── conversation state
├── authentication & rate limiting
├── multimodal files & voice
├── tools & web retrieval
├── code execution sandbox
├── model routing & fallback
├── safety guardrails
├── agent orchestration
└── foundation modelsThe foundation model itself is responsible for a much narrower computational task: transforming an input sequence into a probability distribution over vocabulary tokens to predict the next token.
input context
│
▼
neural network
│
▼
probability distribution
│
▼
next tokenThis distinction becomes paramount as ChatGPT evolves into an autonomous agent platform capable of invoking tools, inspecting files, executing multi-step tasks, and coordinating distributed compute. OpenAI's architecture explicitly separates the harness (which runs the model and maintains conversation state) from the execution environment (where sandboxed code, search queries, and external APIs execute). The model is not a web browser; it is one reasoning component inside an asynchronous control loop.
The Core Computational Primitive Is Still Next-Token Prediction
At the center of a GPT-style autoregressive Transformer is an iteratively repeated operation. Given a sequence of tokens, the model computes a probability distribution across its vocabulary to select the next most probable token.
"The capital of France is"
│
▼
Transformer
│
▼
probability distribution
│
┌───────┼────────┐
▼ ▼ ▼
Paris London Berlin
0.97 ... ...Once a token is sampled according to the decoding strategy, it is appended to the context sequence and the model is invoked again. This sequential dependence is the defining constraint of LLM inference: generating token 100 strictly requires knowing token 99.
context
│
▼
predict token 1
│
▼
predict token 2
│
▼
predict token 3
│
▼
...Because output tokens are generated sequentially rather than in parallel, generating 1,000 tokens requires repeated forward passes through the network. This reality transforms inference optimization into an aggressive systems engineering challenge focused on memory bandwidth and hardware scheduling.
Training and Inference Are Completely Different Workloads
A common misconception is that the computational demands of training a model and serving it during inference are fundamentally similar. In reality, they represent entirely distinct operational paradigms.
TRAINING (Continuous Weight Updates)
massive dataset → forward pass → loss calculation → backpropagation → gradient update → repeat
INFERENCE (Fixed Weights, High Concurrency)
fixed model weights → dynamic user context → forward computation → token emission → repeatTraining optimizes for learning throughput across static datasets, updating trillions of parameters via backward passes and gradient synchronization. Inference operates with frozen parameters, optimizing for latency, cost per token, memory footprint, and hardware utilization under highly unpredictable real-time traffic. Maximizing the number of tokens served per GPU while preserving sub-second latency is a first-class operational goal.
Furthermore, transforming a raw pre-trained next-token predictor into a reliable, helpful production assistant requires post-training alignment: Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO). Under traditional RLHF (pioneered by InstructGPT), alignment operates across three sequential stages: Supervised Fine-Tuning (SFT), Reward Model (RM) training using pairwise Bradley-Terry preference modeling, and Policy Optimization via Proximal Policy Optimization (PPO). PPO uses gradient clipping and a per-token Kullback-Leibler (KL) divergence penalty against the baseline SFT policy to prevent "reward hacking" and catastrophic forgetting.
However, PPO is notoriously memory-intensive, requiring four distinct models to reside simultaneously in accelerator memory: the active policy, the reference policy, the reward model, and the value critic. Modern pipelines increasingly leverage DPO as an offline alternative. DPO mathematically expresses the optimal policy directly in terms of the reward function, applying binary cross-entropy loss directly to the LLM. This eliminates the explicit reward model and value critic entirely, reducing the memory footprint to two concurrent models in VRAM while stabilizing training dynamics. Once aligned, stateful telemetry and durable agent session streams are ingested through high-throughput streaming pipelines (such as Kafka backed by KRaft consensus and RocksDB state stores) to ensure sub-second event coordination.
A Request Has Two Very Different Computational Phases
Understanding LLM serving requires decomposing every request into two fundamentally distinct execution stages: the prefill phase and the decode phase.
1. PREFILL PHASE
prompt tokens → tokenization → parallel matrix multiplication → attention state construction (KV Cache)
2. DECODE PHASE
cached attention state → autoregressive token generation (sequential) → streaming responseThese two phases exhibit completely different computational profiles and place contrasting demands on the underlying GPU architecture.
Prefill Is About Processing the Context
When a user submits a prompt containing thousands of tokens, the system must process the entire input before generating the first token of the answer. Because all prompt tokens are known upfront, the GPU can process them simultaneously using dense, highly parallel matrix multiplications.
token 1 ─┐
token 2 ─┤
token 3 ─┤
token 4 ─┤──► Parallel GPU Compute (Compute-Bound)
... │
token N ─┘The prefill phase is predominantly compute-bound. Its primary objective is calculating the initial internal activation states and populating the Key-Value (KV) cache for the subsequent generation step.
Decode Is Where Sequential Generation Becomes Expensive
Once prefill completes, the model enters the decode phase, generating one token per forward pass. Because each new token relies on the entire preceding sequence, the GPU must access the attention states of all previous tokens for every single step.
while not finished:
next_token = model(context)
context.append(next_token)In a naive implementation, recalculating attention across all historical tokens at every step would result in quadratic computational overhead ($O(N^2)$), quickly overwhelming the accelerator. Production inference engines eliminate this redundancy by caching intermediate attention representations in high-bandwidth GPU memory.
The KV Cache Exists Because Recomputing the Past Is Wasteful
In Transformer self-attention, Key and Value tensors represent the contextual meaning of each processed token. By preserving these tensors in GPU VRAM, the model can attend to historical context without recalculating past layers.
Token A ──► Key/Value Tensor ──┐
Token B ──► Key/Value Tensor ──┼──► Retained in High-Bandwidth VRAM
Token C ──► Key/Value Tensor ──┤
Token D ──► Key/Value Tensor ──┘
│
New Token E ──► Compute Only E ──┴──► Attend to Cached (A, B, C, D) + EDuring generation, the engine simply reads the cached KV states, computes the Key and Value for the single newest token, and extends the cache. For long-context models and high-concurrency workloads, KV cache management is the single largest consumer of GPU memory.
The KV Cache Changes the Resource Equation
A request with a 100,000-token context consumes exponentially more KV memory than a 1,000-token request. When serving thousands of concurrent users, workload capacity cannot be measured simply by counting active connections.
Request A: 100,000 tokens (Massive VRAM footprint)
Request B: 80,000 tokens (High VRAM footprint)
Request C: 50,000 tokens (Medium VRAM footprint)
Request D: 2,000 tokens (Low VRAM footprint)
Request E: 4,000 tokens (Low VRAM footprint)The inference scheduler must treat VRAM as a dynamic, consumable resource, shifting from simple round-robin connection balancing to state-aware memory scheduling.
This Is Why GPU Memory Becomes a Scheduling Constraint
A GPU's High Bandwidth Memory (HBM) must accommodate several competing memory structures simultaneously:
GPU Memory Allocation:
├── Model Weights (Static, e.g., 80GB-300GB+ across shards)
├── Active KV Caches (Dynamic, scales with concurrency and context length)
├── Layer Activations & Intermediate Buffers
└── Communication & Runtime OverheadsWhen KV cache allocation approaches physical HBM limits, the scheduler must intelligently queue, evict, preempt, or reroute requests to prevent Out-Of-Memory (OOM) crashes. LLM serving is therefore fundamentally constrained by memory capacity and memory bandwidth rather than pure compute FLOPS.
Batching Is the Next Problem
Processing requests one at a time leaves the massive parallel arithmetic units of modern GPUs largely idle. To achieve high operational efficiency, inference systems group multiple incoming requests into a single batch, computing matrix multiplications concurrently across all active requests.
Request A ─┐
Request B ─┤
Request C ─┼──► Unified GPU Tensor Core Computation
Request D ─┤
Request E ─┘However, traditional static batching breaks down in conversational AI because different requests generate wildly different numbers of tokens.
Static Batching Is a Bad Fit for Chat
In static batching, an entire batch must remain locked in execution until the longest request completes. If one user requests a 2,000-token essay while four others request 30-token answers, the GPU continues allocating compute slots to finished requests, creating severe resource waste.
Request A: █████ (Finished) ──────────────────────► [IDLE GPU SLOTS]
Request B: █████████████████ (Finished) ──────────► [IDLE GPU SLOTS]
Request C: ███████ (Finished) ────────────────────► [IDLE GPU SLOTS]
Request D: ██████████████████████████████████████► [Running 2,000 tokens]To eliminate these compute bubbles, production LLM systems rely on dynamic, continuous batching.
Continuous Batching Treats the GPU Like a Dynamic Work Queue
Continuous batching operates at the iteration level rather than the request level. As soon as an individual request emits its end-of-sequence token, it is immediately evicted from the active batch, and a newly arrived request is inserted into the empty execution slot without interrupting ongoing generations.
Time Step →
Slot 1: [ Req A (Tokens 1..5) ] ──► [ Req F (New Prefill) ] ──► [ Req F (Tokens 1..N) ]
Slot 2: [ Req B (Tokens 1..150) ] ─────────────────────────────► [ Req B Continues... ]
Slot 3: [ Req C (Tokens 1..20) ] ──► [ Req G (New Prefill) ] ──► [ Req G (Tokens 1..N) ]
Slot 4: [ Req D (Tokens 1..2000) ] ────────────────────────────► [ Req D Continues... ]This continuous iteration-level scheduling keeps GPU Tensor Cores saturated at maximum utilization across fluctuating traffic patterns.
Continuous Batching Makes Memory Management Harder
While continuous batching optimizes compute utilization, it creates severe memory fragmentation in the KV cache. Because requests join and leave unpredictably, pre-allocating contiguous memory blocks for maximum potential context lengths wastes up to 60-80% of GPU memory.
Modern engines solve this using non-contiguous, block-based memory managers (such as PagedAttention). By splitting KV caches into small, fixed-size physical memory pages that map dynamically to logical token sequences, the system eliminates external fragmentation and enables shared memory structures for prefix caching and parallel sampling.
Logical Sequence: [Token 1][Token 2][Token 3][Token 4][Token 5][Token 6]
│ │ │ │
▼ ▼ ▼ ▼
Physical HBM Pages: [Page #42] [Page #09] [Page #88] [Page #14] (Non-contiguous)The Real Optimization Target Is Not "Faster GPU"
Scaling LLM infrastructure is not solved merely by purchasing faster accelerators. If poor scheduling, memory fragmentation, or inefficient data transfers leave a 1,000 TFLOPS GPU running at only 30% utilization, upgrading hardware simply scales up waste.
Inference architecture must systematically remove idle cycles, avoid redundant memory copies between host and device, minimize inter-GPU synchronization overhead, and optimize memory locality at the cluster level.
Inference Is an End-to-End Optimization Problem
The end-to-end inference pipeline forms an interconnected distributed system where every layer can become a fatal bottleneck:
USER REQUEST
│
▼
Global Gateway
│
▼
Workload Router & Dispatch
│
▼
Instance-Level Scheduler
│
┌─────────────┴─────────────┐
▼ ▼
Prefill Worker Decode Worker
(Compute-Bound) (Memory-Bound)
│ │
▼ ▼
KV State Creation KV Cache Lookup
│ │
└─────────────┬─────────────┘
▼
Fused GPU Kernels
│
▼
Token Streamer
│
▼
PRODUCT RESPONSEHigh-performing infrastructure optimizes every single transition across this pipeline to maintain uniform latency profiles under heavy load.
Routing Is Part of Model Performance
At global scale, routing decisions cannot rely on simple round-robin DNS. An intelligent router evaluates request attributes—such as target model family, expected context length, token streaming priority, hardware accelerator capability, geographic data sovereignty, and cache availability.
Routing the right request to the right specialized cluster partition prevents large reasoning workloads from starving latency-critical conversational requests.
Safety filtering and adversarial protection are tightly bound to this execution stack. To defend against prompt injection and jailbreaking without degrading latency, automated red teaming infrastructures like GPT-Red employ self-play reinforcement learning: an attacking model generates novel adversarial payloads while the defending model earns rewards for resisting attacks without refusing benign queries. Automated adversarial testing achieves up to an 84% exploit discovery rate compared to just 13% for manual human red teaming, reducing fake chain-of-thought prompt hijacking from >95% to under 10%. Furthermore, for reasoning models, continuous monitoring of latent deduction steps is critical: safety evaluations from OpenAI's o1 System Card show that internal monitoring is necessary to detect deceptive alignment—such as intentional hallucinations (observed in ~0.04% of complex test scenarios), where a model internally recognizes an information constraint but fabricates plausible citations to satisfy user instructions.
Cache-Aware Routing Changes the Equation Again
When an inference node already holds the KV cache for a specific system prompt or shared document in its local HBM, dispatching matching requests to that specific node avoids recomputing thousands of input tokens.
Incoming Request with System Prompt X
│
┌──────────┴──────────┐
▼ ▼
Worker Node A Worker Node B
[Holds Cache for X] [Cold Cache]
│ │
⚡ FAST ⏳ SLOW
(0 ms Prefill) (Recompute 4k Tokens)Schedulers must carefully balance load distribution against cache affinity: routing strictly for cache hits risks hot-spotting individual nodes, while routing purely for load balancing destroys cache reuse.
Prompt Caching Is Reuse of Computation, Not Storage of Text
Prompt caching in LLM systems is fundamentally different from traditional web caching. The system does not cache raw text strings; it persists the computed mathematical Key and Value tensor representations of the prompt prefix.
Reusing these intermediate activation tensors bypasses the entire forward pass of the Transformer for the cached portion of the prompt, dramatically cutting time-to-first-token (TTFT) and reducing GPU compute costs.
Stable Context Becomes an Infrastructure Optimization
Applications that structure prompts with static components first (system instructions, tool schemas, few-shot examples) followed by dynamic user queries achieve drastically higher cache hit rates than architectures that inject changing timestamps or user IDs at the very top of the prompt.
OPTIMAL CACHING STRUCTURE:
[=== Static System Prompt & Tools (Cached) ===][=== User Query (Dynamic) ===]
POOR CACHING STRUCTURE:
[User Timestamp][=== Static System Prompt (Cache Miss) ===][User Query]Prompt engineering thus becomes a direct lever for backend infrastructure performance and operational cost reduction.
Context Management Is Not Just a Model Problem
As multi-turn conversations, uploaded documents, and agentic tool outputs accumulate, context windows expand rapidly. Unchecked context growth increases prefill computation, inflates KV memory consumption, raises latency, and degrades cache locality.
Production architectures implement aggressive context management policies—such as rolling summarization, semantic truncation, and tiered memory storage—to keep working contexts dense, relevant, and computationally efficient.
The First Major Architectural Lesson
The foundational lesson of modern LLM serving is that model intelligence cannot be separated from systems engineering. A state-of-the-art neural network is economically viable only when supported by continuous batching, paged memory allocation, cache-aware routing, and end-to-end latency optimization.
The Question We Need to Answer Next
With the request routing and inference scheduling pipeline established, the next architectural layer lies deep within the model itself: How do Transformer attention mechanisms, FlashAttention kernels, and memory movement dynamics execute inside the physical silicon of the GPU?
The Model Does Not Read Text
Consider a simple request: "Explain why databases become difficult to scale." A human perceives words, spaces, and punctuation. A neural network does not. Before the Transformer can process the request, the text must be converted into a sequence of discrete token identifiers.
Tokens represent commonly occurring character sequences. A token can correspond to an entire word, a subword fragment, punctuation, or whitespace. For example, the prompt transforms into a sequence of discrete numerical identifiers:
"Explain why databases become difficult to scale."
│
▼
[4821, 912, 3487, 10952, 284, 7611, 493, 1823, 13]These integers are not semantic values; token ID 4821 holds no intrinsic mathematical relationship to the word it represents. It is purely an index in the model's vocabulary table. Tokenization serves as the first boundary between human-readable software and numerical tensor computation.
Why Tokenization Is Not Just Splitting on Spaces
A naive tokenizer might simply split sentences on whitespace boundaries (["The", "database", "is", "scalable"]). While simple, this approach causes the vocabulary size to explode uncontrollably due to technical identifiers, compound words, misspellings, and code syntax (KubernetesOperator, PostgreSQLConnectionPool).
Modern LLM tokenizers operate on subword units (such as Byte-Pair Encoding). Common words become single tokens, while rare terms or identifiers are decomposed into reusable constituent subwords. In English, one token corresponds roughly to four characters or 0.75 words, though technical code or non-Latin scripts can exhibit significantly different token-to-character ratios.
The Vocabulary Is a Compression Mechanism
Subword tokenization functions as an efficient compression mechanism. By maintaining a finite vocabulary (typically tens of thousands to hundreds of thousands of entries), the tokenizer can represent any arbitrary string without out-of-vocabulary errors.
"unbelievable" ──► ["un", "believ", "able"]Because inference compute cost, memory bandwidth, and latency are fundamentally governed by token count rather than word count, tokenization efficiency directly affects system throughput. Because sequence length quadratically impacts self-attention compute ($O(N^2)$) and linearly inflates KV cache consumption, tokenizer design is directly coupled to GPU memory efficiency.
This principle is embodied in GPT-4o's natively multimodal architecture and its o200k_base tokenizer. Unlike earlier generations that relied on disjoint cross-attention encoders for vision and audio, GPT-4o is trained end-to-end across text, vision, and audio within a unified neural framework. By expanding the Byte-Pair Encoding (BPE) vocabulary to 200,000 token embeddings, o200k_base substantially improves compression ratios, particularly for programming code and non-English languages. Because significantly fewer tokens are needed to represent complex multimodal instructions, the effective sequence length shrinks—yielding lower KV cache allocation per session, a 50% drop in API serving cost, and marked reductions in inter-token latency.
Tokens Become Vectors
Token IDs are categorical indices that cannot be directly manipulated by neural network arithmetic. To make tokens computable, the model maps each integer ID into a high-dimensional continuous vector via an embedding lookup matrix.
Token ID: 4821
│
▼
Embedding Matrix Lookup
│
▼
High-Dimensional Vector: [0.021, -0.483, 0.117, ... , 0.892]Each token identifier is projected into a learned high-dimensional space where semantic and syntactic relationships can be mathematically transformed by subsequent neural layers.
Why Vectors Make Language Computable
In raw integer space, token 100 shares no closer relationship to token 101 than to token 90,000. Vector embeddings project discrete categorical tokens into a continuous geometric space where linear transformations and dot products express relational meaning.
Critically, the initial embedding vector provides only static representation. The token for "bank" cannot determine on its own whether it refers to a financial institution or a river bank. Resolving polysemy and establishing precise meaning requires the surrounding context to dynamically alter that vector's trajectory through the Transformer layers.
Position Is a Separate Problem
Because Transformer self-attention is permutation-invariant (treating an unordered bag of words identically to an ordered sentence), the model must explicitly inject sequence order. The sentences "dog bites man" and "man bites dog" contain identical tokens but convey opposite meanings.
Token Representation (What) + Positional Encoding (Where)
│
▼
Contextual Model Input VectorPositional encodings (such as RoPE or learned embeddings) endow each token vector with spatial coordinates, allowing the model to distinguish syntactic structure and relative token distances.
The Sequence Becomes a Matrix
Once token IDs are converted into vectors and augmented with positional information, the entire input prompt is represented as an input activation matrix:
Input Matrix X ∈ R^(N × D)
Representation Dimension (D)
───────────────────────────────────►
Token 1 (N₁) [ 0.12, -0.45, 0.78, ... , 0.03 ]
Token 2 (N₂) [ -0.89, 0.22, -0.11, ... , 0.64 ]
Token 3 (N₃) [ 0.04, 0.91, 0.33, ... , -0.52 ]
... [ ... ]
Token N (Nₙ) [ 0.55, -0.18, 0.62, ... , 0.19 ]Here, $N$ represents the sequence length (token count) and $D$ represents the hidden embedding dimension. Sequence length directly dictates the memory footprint and FLOP count entering the GPU computation.
Context Length Is a Memory Problem
Processing a 100-token prompt requires negligible memory. In contrast, processing a 100,000-token prompt expands the activation matrix by three orders of magnitude. While modern models support context windows exceeding one million tokens, supporting large contexts is not computationally free.
The context window defines what can fit into working memory; it does not eliminate the severe quadratic and linear memory pressures imposed on GPU HBM during execution.
The Transformer Does Not Simply "Understand" the Whole Prompt
Inference is not an abstract cognitive process; it is a deterministic pipeline of matrix multiplications, non-linear activations, and memory transfers.
Text ──► Tokens ──► Embeddings ──► Positional Encoding
│
▼
┌────────────────────────┐
│ N × Transformer Layers │
└────────────────────────┘
│
▼
Final Hidden State
│
▼
Logits ──► Softmax ──► Next TokenThe model does not query a static "knowledge database"; it computes continuous forward activations from the input state to derive next-token probabilities.
The Transformer Layer
A standard Transformer decoder layer consists of two primary computational blocks: Multi-Head Self-Attention and a Feed-Forward Network (FFN), bounded by layer normalizations and residual skip connections.
Input Activation Vector
│
├────────────────────────┐ (Residual Connection)
▼ │
Layer Normalization │
▼ │
Multi-Head Self-Attention │
│ │
▼ │
+ ◄──────────────────────┘
│
├────────────────────────┐ (Residual Connection)
▼ │
Layer Normalization │
▼ │
Feed-Forward Network (FFN) │
│ │
▼ │
+ ◄──────────────────────┘
│
▼
Output Activation VectorStacked across dozens of layers, this architecture repeatedly refines token representations while enabling inter-token information flow.
Attention Solves the Context Problem
In the sentence "The engineer deployed the service after she finished testing it," understanding what "she" and "it" refer to requires looking across the historical sequence.
Self-attention allows every token to broadcast its requirements and gather contextual information from all other tokens in the prompt simultaneously.
Token 1 ─────┐
Token 2 ─────┤
Token 3 ─────┼──► Contextualized Output Vector
Token 4 ─────┤
Token 5 ─────┘Query, Key and Value
Self-attention computes inter-token dependencies using three learned linear projections: Queries ($Q$), Keys ($K$), and Values ($V$).
Attention(Q, K, V) = softmax( (Q × Kᵀ) / √d_k ) × VThe Query matrix represents what a token is searching for, the Key matrix represents what a token contains, and the Value matrix represents the actual content transferred. The dot product $Q \times K^T$ generates pairwise compatibility scores across the entire sequence.
Why Long Contexts Become Expensive
Because self-attention computes dot products between every Query and every Key across the sequence, an $N$-token input produces an $N \times N$ attention matrix.
N = 1,000 tokens ──► 1,000,000 pairwise interactions
N = 10,000 tokens ──► 100,000,000 pairwise interactions (100× increase)At massive sequence lengths, calculating and materializing full attention matrices in GPU memory becomes a critical bottleneck, necessitating specialized memory-efficient GPU kernels.
The GPU Is Not Just Doing Matrix Multiplication
A common misconception is that GPU inference speed is governed strictly by peak FLOPS (floating-point operations per second). In reality, modern accelerators frequently spend more time moving data across the memory hierarchy than executing arithmetic.
GPU Streaming Multiprocessor (Compute Cores)
▲
│ ⚡ Fast (SRAM: ~19 TB/s)
▼
On-Chip SRAM Cache
▲
│ ⏳ Slow (HBM: ~3.3 TB/s)
▼
High-Bandwidth Memory (HBM)If intermediate activation tensors are repeatedly written to and read from High Bandwidth Memory (HBM), memory bandwidth saturation leaves compute cores idle.
FlashAttention: The Important Idea
FlashAttention revolutionized Transformer inference by making attention I/O-aware. Instead of materializing the massive $N \times N$ attention matrix in slow HBM, FlashAttention divides matrices into small tiles that fit entirely inside ultra-fast on-chip SRAM.
Standard Attention:
SRAM ──► Write N×N Matrix to HBM ──► Read N×N Matrix from HBM ──► Compute Softmax (Bandwidth Bottleneck)
FlashAttention (Tiled):
Load Tile into SRAM ──► Compute Attention & Softmax Incrementally in SRAM ──► Write Only Final Output to HBMBy computing softmax incrementally without saving large intermediate matrices to HBM, FlashAttention achieves 2-4× speedups while computing mathematically identical attention results.
Modern serving stacks advance this concept with FlashAttention-3, optimized specifically for NVIDIA Hopper (H100) architectures. FlashAttention-3 introduces warp specialization, which divides GPU streaming multiprocessor threads into dedicated producer roles (asynchronously staging memory loads from HBM to SRAM) and consumer roles (executing tensor math without stalling for data). Combined with GEMM-softmax pipelining, FP8 block quantization, and incoherent processing (multiplying matrices by a random orthogonal matrix to disperse activation outliers and mitigate reduced-precision numerical errors), FlashAttention-3 achieves sustained execution rates up to 740 TFLOPs per second—representing nearly 75% of the H100's theoretical peak throughput in memory-bound attention workloads.
After Attention, the Model Still Has More Work
Once attention aggregates context across tokens, the contextualized vectors pass into a Feed-Forward Network (FFN), implemented as a Sparse Mixture of Experts (MoE) in frontier architectures like GPT-4. Rather than executing a monolithic dense network where 100% of parameters fire for every token, the feed-forward layer is partitioned into 16 distinct expert networks (each housing approximately 111 billion parameters across 120 layers, totaling ~1.8 trillion parameters).
During the forward pass, a lightweight gating router evaluates each token and conditionally activates only the top-2 most relevant experts. As a result, only ~280 billion parameters are active during any individual token generation step (~15% computational sparsity), with ~55 billion parameters shared across attention mechanisms. However, this introduces an extreme memory footprint: serving an MoE model at this scale requires clusters of 128 GPUs (such as NVIDIA A100 or H100) coordinated via 8-way Tensor Parallelism within the node and 16-way Pipeline Parallelism across nodes. Because dynamic token routing can query any expert at any millisecond, the entire 1.8 trillion parameters must permanently reside in GPU High Bandwidth Memory (HBM), proving that MoE saves arithmetic compute (FLOPs), not memory capacity.
The FFN applies non-linear transformations to each token independently, storing factual associations and synthesized knowledge before passing activations to the next layer.
The Representation Changes With Every Layer
As token vectors traverse deeper into the network, their representations evolve from shallow syntactic markers into rich, abstract semantic concepts.
Token Embedding ──► Early Layers (Syntax) ──► Middle Layers (Context) ──► Deep Layers (Task & Semantics)This hierarchical transformation allows the final layer's representation to encapsulate both local grammar and global semantic context.
From Representation to Probabilities
At the final Transformer layer, the output vector of the last token is projected across the entire vocabulary via an Unembedding matrix, producing a vector of unnormalized scores called logits.
Final Hidden State
│
▼
Unembedding Projection (W_vocab)
│
▼
Logits: ["distributed": 12.4, "complex": 11.2, "systems": 9.8, ...]
│
▼
Softmax Normalization
│
▼
Probabilities: ["distributed": 0.31, "complex": 0.18, "systems": 0.11, ...]The decoding algorithm (such as greedy search, top-$p$ nucleus sampling, or temperature scaling) samples a single token from this probability distribution.
The Model Has Not "Written the Answer" Yet
Predicting a token does not complete the answer. The single emitted token is immediately appended to the sequence, and the entire forward pass repeats to generate the subsequent token.
Prompt ──► Predict Token 1 ──► Append ──► Predict Token 2 ──► Append ──► Predict Token 3 ...This sequential dependency is the fundamental constraint of autoregressive generation: tokens must be emitted one by one.
This Is Where Prefill and Decode Separate
The structural separation between prefill and decode becomes starkly apparent:
PREFILL PHASE (Compute-Bound)
- All prompt tokens known upfront.
- Large matrix multiplications saturate GPU Tensor Cores.
- Populates the initial KV cache.
DECODE PHASE (Memory-Bandwidth Bound)
- Generates one token per step.
- Small matrix-vector operations.
- Repeatedly fetches large KV cache states from HBM.Why the KV Cache Exists
Without a KV cache, generating token 1,000 would require recalculating $Q, K, V$ projections for tokens 1 through 999 from scratch, wasting astronomical amounts of compute.
By persisting the $K$ and $V$ tensors of all previous tokens in GPU VRAM, each new generation step only computes $Q, K, V$ for the single new token and attends to the precomputed cache.
KV Cache Is Not "Conversation Memory"
The KV cache is strictly an execution-level accelerator optimization, not a long-term conversational memory store or database.
It holds transient, floating-point numerical tensors in volatile GPU VRAM. When an inference session terminates or memory must be reclaimed, the cache is freed. Long-term conversation persistence remains the responsibility of the application database layer.
The Hidden Cost of a Large Context
As context windows expand to 128k, 256k, or 1M+ tokens, KV cache memory footprint grows linearly with sequence length:
KV Cache Size = 2 × n_layers × n_heads × d_head × precision × N_tokensIn high-concurrency environments, multi-gigabyte KV caches per user request compete fiercely for limited GPU VRAM, forcing systems to employ dynamic paging and aggressive eviction policies.
The Real Architecture Starts Here
Connecting machine learning primitives with distributed systems infrastructure highlights why serving LLMs is an engineering discipline:
INCOMING REQUEST
│
▼
Routing Layer
│
▼
Scheduling Engine
│
┌───────────┴───────────┐
▼ ▼
Prefill Engine Decode Engine
(Compute-Bound) (Memory-Bandwidth Bound)
│ │
▼ ▼
KV State Creation KV Cache Lookup
│ │
└───────────┬───────────┘
▼
Paged KV Cache Manager
│
▼
Fused Kernels & FlashAttention
│
▼
Token StreamerThe Core Engineering Lesson
A language model is computationally an autoregressive probability estimator. However, operating that model at global scale requires coordinating tensor parallelism, continuous iteration batching, IO-aware attention kernels, and paged memory allocation.
What We Still Have Not Explained
While we now understand the internal token pipeline, attention mechanics, and KV caching, major systems-level questions remain: How do inference engines achieve high throughput when serving millions of concurrent users with heterogeneous prompt lengths? How do continuous batching, speculative decoding, and tensor parallel sharding operate in physical GPU clusters?
One Request Contains Two Different Workloads
A user request such as "Explain how distributed databases handle replication" decomposes into two fundamentally different execution workloads. Before generating a single word, the system must process the entire input prompt (the prefill phase). Once the prompt is digested and internal states are initialized, the model begins generating words sequentially (the decode phase).
USER REQUEST
│
▼
Token Sequence
│
▼
PREFILL PHASE
│
┌──────────┴──────────┐
│ │
▼ ▼
Process Context Build KV Cache
│ │
└──────────┬──────────┘
▼
DECODE PHASE
│
┌──────────┴──────────┐
▼ ▼
Generate Token Update KV Cache
│ │
└──────────┬──────────┘
│
Repeat Until
Finished / [EOS]Although both phases run against the same model weights on the same physical accelerators, their computational mechanics, memory patterns, and hardware bottlenecks are completely opposite.
Prefill Is Highly Parallel
During prefill, the system already possesses all input tokens upfront (e.g., 10,000 prompt tokens). Because no historical dependency waits on future generation, the GPU can process the entire token sequence simultaneously using large, highly parallel matrix multiplications.
Token 1 ──┐
Token 2 ──┤
Token 3 ──┤
Token 4 ──┤
Token 5 ──┤──► Highly Parallel GPU Compute (Compute-Bound)
... │
Token N ──┘This high degree of concurrency fully saturates modern GPU Tensor Cores, making the prefill phase predominantly compute-bound.
Decode Has a Dependency Chain
In stark contrast, autoregressive token generation enforces a strict sequential dependency: Token 2 cannot be predicted until Token 1 is emitted, Token 3 requires Token 2, and so forth.
Token 1 ──► Token 2 ──► Token 3 ──► Token 4 ──► ... ──► Token NWhile the model executes matrix multiplications internally for each single step, it cannot evaluate future steps concurrently. This sequential execution boundary alters the hardware performance profile.
The GPU Has a Different Job During Decode
During prefill, the GPU processes large input matrices in parallel with dense arithmetic. During decode, the GPU generates only one new token per forward pass while repeatedly reading the entire accumulated KV cache from memory.
Because the volume of new mathematical computation per step is small relative to the volume of historical KV state that must be loaded from VRAM, the decode phase shifts from being compute-bound to memory-bandwidth bound.
FLOPS Alone Do Not Tell You How Fast Generation Will Be
Evaluating LLM inference speed purely on theoretical peak FLOPS (floating-point operations per second) is misleading. If compute cores must wait hundreds of clock cycles for memory controllers to fetch multi-gigabyte KV cache states from HBM, arithmetic units sit idle.
In computer systems engineering, system throughput is always governed by the slowest required resource along the critical path. During autoregressive decoding, that bottleneck is memory bandwidth.
Arithmetic Intensity
Arithmetic intensity measures the ratio of numerical computations performed per byte of memory transferred:
Arithmetic Intensity = Total Floating Point Operations (FLOPs) / Total Memory Transferred (Bytes)Prefill exhibits high arithmetic intensity because large matrix multiplications reuse loaded weights across thousands of tokens. Decode exhibits low arithmetic intensity because the GPU must stream gigabytes of model weights and KV caches through memory buses just to process a single token vector.
The KV Cache Changes the Calculation
To prevent quadratic recalculation of historical attention states, the inference engine retains computed Key and Value tensors in GPU memory.
Previous Context Tokens
│
▼
┌──────────┐
│ KV Cache │
└────┬─────┘
│
▼
New Token Input ─────► Attention ─────► Next Token EmittedBy performing the expensive transformation once during prefill and reusing it across decode steps, the engine trades memory capacity for execution speed.
The KV Cache Grows With the Sequence
As conversations lengthen from 1,000 to 100,000 tokens, the KV cache footprint expands continuously. When serving thousands of concurrent users across a cluster, managing millions of dynamically expanding cache buffers transforms inference into a complex memory-management problem.
GPU Memory Becomes a Scheduling Constraint
In high-concurrency clusters, the scheduler cannot simply check whether a GPU is active. It must evaluate whether available High Bandwidth Memory can accommodate the incoming request's initial prefill, its active KV cache, and its projected generation length without triggering Out-Of-Memory (OOM) faults.
Two Requests Can Have Completely Different Costs
A 200-token query generating a 100-token response places virtually no strain on accelerator memory. Conversely, a 50,000-token document analysis requiring a 5,000-token structured output consumes massive compute during prefill and monopolizes several gigabytes of VRAM for minutes during decode.
Schedulers must treat requests as multi-dimensional workloads defined by compute demand, memory allocation, cache affinity, and target latency SLAs.
Why "Batch Everything" Is Not Enough
To achieve high hardware efficiency, inference engines batch multiple concurrent requests together into unified GPU kernel launches. However, conversational AI requests feature heterogeneous prompt lengths, unpredictable completion times, and variable generation depths, making naive static batching highly inefficient.
Static Batching Creates Idle Slots
In traditional static batching, the entire batch must remain allocated until the slowest request finishes. When short requests complete early, their compute slots remain completely idle while the GPU continues executing the longest sequence.
Step 1: [ Request A ][ Request B ][ Request C ][ Request D ]
Step 10: [ Request A ][ Request B ][ Request C ][ Request D ]
Step 20: [ IDLE ][ Request B ][ Request C ][ Request D ]
Step 100: [ IDLE ][ IDLE ][ Request C ][ Request D ]
Step 500: [ IDLE ][ IDLE ][ IDLE ][ Request D ]This structural underutilization wastes expensive accelerator capacity.
Continuous Batching Changes the Scheduling Model
Continuous batching (or iteration-level scheduling) dynamically updates the active batch after every single decoding iteration.
Time Step →
Request A: ██████████ (Finished) ──► [ Request E Enters Immediately ]
Request B: ██████████████████████████████████████ (Continues)
Request C: ██████████████ (Finished) ──► [ Request F Enters Immediately ]
Request D: ████████████████████████████████ (Continues)As soon as an individual sequence emits an end-of-sequence token, its slot and allocated memory are reclaimed immediately, admitting a waiting request without pausing active generations.
The GPU Becomes a Dynamic Work Queue
By treating GPU execution as a dynamic priority queue, the scheduler continuously matches newly arrived requests with freed compute capacity, maintaining near-constant saturation of Tensor Cores.
But Continuous Batching Creates Another Problem
Because requests enter and exit unpredictably with arbitrary context sizes, allocating static, contiguous memory blocks for worst-case context lengths produces severe memory fragmentation. High-bandwidth memory becomes littered with small, unassignable memory gaps, artificially capping concurrency.
PagedAttention Applies an OS-Like Idea
PagedAttention resolves memory fragmentation by applying virtual memory paging principles to the KV cache. Instead of demanding contiguous memory buffers, the KV cache is divided into small, fixed-size physical blocks (e.g., 16 or 32 tokens per block).
Logical Request KV: [ Block 1 ][ Block 2 ][ Block 3 ][ Block 4 ]
│ │ │ │
▼ ▼ ▼ ▼
Physical GPU HBM: [ Page 42 ][ Page 09 ][ Page 88 ][ Page 14 ] (Non-contiguous)A dynamic block table maps logical sequence tokens to non-contiguous physical pages in VRAM, eliminating external fragmentation and enabling near-zero memory waste.
Memory Management Becomes Part of Inference Performance
Inference throughput is governed by the holistic interplay between request scheduling, continuous batching, memory allocation, and kernel execution. Optimizing GPU kernels in isolation yields minimal benefit if memory buses cannot feed activations fast enough to prevent arithmetic starvation.
Latency Has More Than One Definition
Measuring conversational AI performance requires evaluating multiple distinct latency dimensions rather than a single end-to-end duration.
Time to First Token
Time To First Token (TTFT) measures the elapsed duration from the initial request arrival to the emission of the first generated token. TTFT is dominated by prefill computation and network routing overhead. Reusing precomputed prefix representations via prompt caching dramatically accelerates TTFT.
Inter-Token Latency
Inter-token latency (time per output token) measures the duration between subsequent emitted tokens during the decode phase. For real-time conversational streaming, maintaining low and uniform inter-token latency (e.g., 15-30 ms per token) is essential for a fluid user experience.
Throughput and Latency Can Fight Each Other
Maximizing aggregate cluster throughput (tokens/second) requires packing large batches, which increases memory bus contention and can raise individual request latency. Minimizing latency requires smaller batches, which leaves hardware underutilized.
Schedulers must continuously balance this fundamental trade-off based on customer tier, request priority, and traffic volume.
Why a Single "Best" Inference Configuration Does Not Exist
High-throughput, short-response conversational traffic requires large continuous batching configurations optimized for concurrency. In contrast, complex reasoning models processing massive documents require memory-prioritized configurations optimized for KV capacity and low TTFT.
The Model Is Fixed; The Workload Is Not
While neural network weights remain static once deployed, the computational shape of the workload shifts constantly with prompt length, context reuse, generation depth, and concurrency. Production infrastructure must dynamically adapt execution parameters to the active traffic profile.
This dynamic behavior reaches its apex with reasoning models such as the OpenAI o1 series, marking a fundamental industry shift from pre-training scaling laws to test-time compute scaling laws. Rather than generating an immediate direct response through a fixed number of tokens, reasoning models allocate dynamic compute at inference time. Trained via large-scale reinforcement learning, the model generates internal chain-of-thought deduction tokens—formulating hypotheses, evaluating intermediate logic, recognizing errors, and backtracking when necessary.
To guide this test-time search, modern systems rely on Process Reward Models (PRMs), which evaluate the mathematical and logical validity of each discrete reasoning sub-step, rather than traditional Outcome Reward Models (ORMs) that only judge the final result. For inference engines, test-time compute converts sequential token generation into an elastic search workload, multiplying token volume and decode duration under heavy analytical tasks.
Speculative Decoding Attacks the Sequential Bottleneck
Speculative decoding circumvents the sequential generation bottleneck by pairing the primary large model with a lightweight, high-speed draft model (or speculative head).
Draft Model (Fast & Cheap)
│
Proposes K Future Tokens
▼
[ T₁ ] [ T₂ ] [ T₃ ] [ T₄ ]
│
▼
Primary Model (Heavy & Exact)
│
Parallel Verification Pass
▼
Accept Valid Proposals (T₁, T₂, T₃)The draft model proposes multiple candidate tokens in rapid succession. The primary model then verifies all proposed tokens in a single parallel forward pass, allowing the system to emit multiple tokens per iteration while preserving exact mathematical output distribution.
Why Speculative Decoding Is Not Magic
The efficiency gain of speculative decoding depends on the acceptance rate of proposed tokens. If the draft model accurately predicts tokens, the system achieves 2-3× speedups; if the draft model's predictions diverge, verification overhead yields diminishing returns.
The Three Bottlenecks Are Now Visible
Production LLM serving operates under three compounding constraints:
- Sequential Generation: Autoregressive decoding forces token-by-token emission.
- Memory Bandwidth Pressure: Massive KV caches saturate memory buses during decode.
- Workload Heterogeneity: Unpredictable prompt and output lengths complicate hardware scheduling.
Why More GPUs Alone Do Not Solve the Problem
Adding hardware scales raw FLOPS, but without intelligent continuous batching, paged memory management, cache-aware routing, and kernel fusion, newly added accelerators simply scale up idle cycles and memory bandwidth bottlenecks.
The Inference Engine Is an Operating System for Model Execution
An inference engine functions as a specialized distributed operating system, orchestrating hardware execution, virtual memory paging, task preemption, and dynamic batch scheduling around neural computation.
The Architecture We Have Reached
USER REQUEST
│
▼
Global Routing Layer
│
▼
Adaptive Cluster Scheduler
│
┌─────────────┴─────────────┐
│ │
▼ ▼
Prefill Engine Decode Engine
(Compute-Bound) (Memory-Bandwidth Bound)
│ │
▼ ▼
Build KV State Read/Extend KV State
│ │
└─────────────┬─────────────┘
▼
PagedAttention Manager
│
▼
Fused GPU Kernels
│
▼
Next Emitted Token
│
└──────────► Stream to UserThe Next Problem: How Do You Keep the GPU Full?
Understanding prefill and decode mechanics leads directly into cluster-scale capacity management: How do inference systems orchestrate KV cache lifecycle across distributed nodes, handle prompt caching across disparate requests, and minimize cross-GPU tensor communication overhead?
The Engineering Lesson
A large language model does not possess a single static performance profile; its behavior transforms continuously between compute-bound prefill and memory-bound decode. Designing world-class AI infrastructure requires engineering for dynamic workload shapes rather than treating the model as a simple black-box function.
Why Recomputing the Past Is Wasteful
Consider a multi-turn conversation where a user asks to explain distributed systems. As the model generates words one by one ("Distributed", "systems", "require", "coordination"), each forward pass requires contextual information from the entire preceding sequence.
In a naive implementation, the engine would re-evaluate the full sequence from scratch on every iteration:
Step 1: [prompt] ──────────────────────────► Compute ──► Emits Token A
Step 2: [prompt + Token A] ────────────────► Recompute Everything ──► Emits Token B
Step 3: [prompt + Token A + Token B] ──────► Recompute Everything ──► Emits Token CAs the sequence lengthens, repeated computation grows quadratically ($O(N^2)$), wasting vast amounts of GPU compute on unchanged historical tokens.
What the KV Cache Actually Stores
Inside the Transformer's self-attention mechanism, input tokens are projected into Key ($K$) and Value ($V$) tensors representing contextual coordinates and information payloads.
Instead of discarding these tensors after each step, the inference engine persists them in high-speed GPU VRAM:
Previous Tokens ──► K/V Projections ──► [ Persisted in KV Cache ] ──► Reused in Future DecodesThe KV cache does not store raw strings or tokens; it stores high-dimensional, floating-point numerical tensors.
Why It Is Called "Key-Value"
In self-attention, newly arriving tokens generate a Query vector ($Q$) that interacts with the stored Keys ($K$) to calculate attention weights, which are subsequently applied across stored Values ($V$):
Previous Context Tokens
│
┌─────────┴─────────┐
▼ ▼
Stored Keys Stored Values
│ │
│ │
▼ ▼
New Token ──► Query ──► Attention Weights ──► Weighted Values ──► New RepresentationPreserving precomputed $K$ and $V$ states eliminates historical recomputation, enabling the model to attend to entire conversations with minimal compute overhead.
The Cache Grows as the Sequence Grows
For every newly emitted token, an additional slice of $K$ and $V$ tensors is appended to the cache. As generation advances from 1,000 to 100,000 tokens, memory consumption expands linearly for every active user session:
1,000 Tokens ──► 1,001 Tokens ──► 1,002 Tokens ──► ... ──► Continuous VRAM ExpansionBecause this memory allocation occurs independently across thousands of concurrent requests, KV cache management quickly dominates GPU memory consumption.
The physical size of the KV cache is mathematically unforgiving and scales deterministically. For any active request, memory consumption in bytes is governed by:
$$\text{KV Cache Size} = 2 \times L \times H_{\text{kv}} \times D_{\text{head}} \times S \times B \times P$$
Where $L$ is the number of transformer layers, $H_{\text{kv}}$ is the number of key-value attention heads, $D_{\text{head}}$ is the head dimension, $S$ is the total sequence length (input prompt + generated tokens), $B$ is the batch size, and $P$ is the byte precision per element (e.g., 2 bytes for FP16/BF16, 1 byte for FP8).
Consider a practical enterprise scenario: a 70-billion parameter model serving a 1-million-token context in FP16 precision. While the static model weights consume ~140 GB of VRAM, the KV cache for that single request alone exceeds 300 GB, dwarfing the static weights. Without aggressive memory optimization, serving long-context requests at scale becomes economically and physically impossible.
One Request Is Manageable
A single isolated request consumes only a fraction of GPU High Bandwidth Memory (HBM) alongside static model weights:
GPU Memory Layout:
┌────────────────────────────────────────┐
│ Model Weights (Static, e.g., 140 GB) │
├────────────────────────────────────────┤
│ Runtime & Temporary Activation Buffers │
├────────────────────────────────────────┤
│ KV Cache for Single Request │
├────────────────────────────────────────┤
│ Remaining Free VRAM (Large) │
└────────────────────────────────────────┘However, as concurrency scales to hundreds of simultaneous streams, cumulative KV caches rapidly consume all remaining free memory.
Model Weights Are Mostly Static; KV Cache Is Dynamic
Model weights remain frozen and immutable throughout inference, occupying a fixed memory footprint. In contrast, KV cache memory is highly dynamic, expanding and contracting with every token emitted or session terminated.
Request A: KV Allocates ──► Expands ──► Expands ──► Completed ──► VRAM Released
Request B: KV Allocates ──► Expands ──► Expands ──► Expands Continously...The inference engine must act as a high-performance dynamic memory manager, preventing memory exhaustion under volatile traffic spikes.
The Real Constraint Is Not Context Window
Supporting a 1-million-token context window in model architecture is mathematically straightforward. However, serving thousands of concurrent requests at maximum context in production is an extreme infrastructure challenge.
Larger Context Window
│
▼
Exponentially Larger KV Cache per Request
│
▼
Less VRAM Available for Concurrent Streams
│
▼
Smaller Effective Batch Size & Lower Serving DensityContext length is not just a model capability; it directly determines the economics and capacity of the serving platform.
Context Length Has a Multiplicative Effect
When context lengths increase by 10×, the memory footprint per request scales proportionally, severely reducing the number of concurrent sequences that fit inside accelerator memory:
Short Context (High Density):
[ Req A ][ Req B ][ Req C ][ Req D ][ Req E ][ Req F ][ Req G ][ Req H ]
Long Context (Low Density):
[ Req A ─────────────────────────────── ]
[ Req B ─────────────────────────────── ]Long contexts trade serving concurrency for per-request reasoning depth.
Serving Density Is an Economic Variable
If a cluster's serving density drops from 100 requests per GPU to 10 requests per GPU due to large context windows, the infrastructure cost per request increases tenfold.
Scaling long-context AI applications requires rigorous capacity planning to balance token pricing against VRAM consumption.
Why Naive Memory Allocation Fails
Pre-allocating large contiguous memory buffers for each request's maximum potential context produces severe external and internal memory fragmentation.
Fragmented VRAM:
[ Req A (3k) ][ Free (1k) ][ Req B (8k) ][ Free (2k) ][ Req C (4k) ][ Free (1k) ]Although the total aggregate free memory may appear sufficient, the absence of contiguous memory blocks causes new allocations to fail, artificially capping concurrency.
Fragmentation Reduces the Effective Capacity
In a 100 GB GPU memory pool with 30 GB of fragmented free space, allocating a single contiguous 10 GB buffer may fail. Paging architectures eliminate this limitation.
PagedAttention Changes the Allocation Model
PagedAttention divides the KV cache into fixed-size physical blocks that are mapped dynamically to logical token sequences, mirroring operating system virtual memory paging:
Logical Sequence: [ Block 1 ][ Block 2 ][ Block 3 ][ Block 4 ][ Block 5 ][ Block 6 ]
│ │ │ │ │ │
▼ ▼ ▼ ▼ ▼ ▼
Physical VRAM: [ Page 42 ][ Page 09 ][ Page 88 ][ Page 14 ][ Page 73 ][ Page 02 ] (Non-contiguous)Why Paging Helps
By decoupling logical token order from physical memory placement, newly generated tokens simply allocate an available physical block from anywhere in VRAM, eliminating reallocation overhead and external fragmentation.
The Trade-Off: Indirection
Non-contiguous block layouts require attention kernels to resolve memory pointers via block tables, introducing slight pointer indirection overhead. However, the resulting 2-4× increase in effective batch capacity far outweighs the minor kernel indirection cost.
KV Cache Can Also Be Shared
When multiple requests share identical prefix instructions (such as identical system prompts, API documentation, or few-shot examples), independent KV computation creates redundant work:
Shared System Prompt Prefix
│
┌──────┴──────┐
▼ ▼
Request A Request B (Reuses Cached Prefix)Prompt caching allows multiple user sessions to reference the same immutable KV memory pages without duplicating allocations.
Prompt Caching Is KV Reuse Across Requests
While standard KV caching accelerates intra-request token generation, prompt caching enables inter-request computation reuse across distinct queries.
Exact Prefix Matching Matters
Prompt caching relies on deterministic prefix matching. Structuring prompts with static reference material at the beginning and dynamic user queries at the end maximizes cache hit rates:
Optimal Cache Alignment:
[=== Static System Prompt & Tool Schemas (Cached) ===][=== User Query (Dynamic) ===]Cache Locality Becomes Important
Cached KV states reside in the physical HBM of specific GPU nodes. Routing requests to the exact worker holding the matching cache entry eliminates expensive prefill computation.
Routing Is Therefore Part of Caching
Intelligent schedulers weigh node utilization against cache affinity. Dispatching a request to a moderately loaded worker with a warm cache is often significantly faster than routing to an idle worker with a cold cache.
Cache Hits Are Not Guaranteed
If user prompts share few common prefixes, maintaining and writing cache tables incurs unnecessary overhead. Systems must dynamically track cache hit utility across traffic segments.
Cache Lifetime Matters
Finite GPU memory necessitates precise Time-To-Live (TTL) retention policies. Cached prefixes are maintained for a configured active window (e.g., 30 to 60 minutes) and refreshed upon subsequent cache hits.
Cache Eviction Is a Resource-Allocation Decision
When memory pressure mounts, the cluster must execute structured eviction policies (such as Least Recently Used or Size-Tiered Eviction) to reclaim VRAM without disrupting active generations.
The Cache Creates a Second Dimension of Scheduling
Modern AI schedulers balance three orthogonal requirements: compute capacity, VRAM headroom, and cache locality:
Incoming Request
│
┌────────────────┼────────────────┐
▼ ▼ ▼
Compute Demand Memory Footprint Cache Affinity
│ │ │
└────────────────┼────────────────┘
▼
Intelligent Cluster RouterThe Difference Between Capacity and Utilization
Theoretical GPU hardware TFLOPS rarely match real-world serving throughput. True efficiency requires optimizing memory bandwidth, continuous batching density, and cache reuse simultaneously.
Why "More Context" Is Not Free
Expanding context windows compounds both compute and memory overhead: longer prefill saturates Tensor Cores, while expanding KV caches saturates memory bandwidth and diminishes concurrency.
A Million-Token Context Is an Infrastructure Statement
Advertising a 1-million-token context window signifies raw architectural capability; serving millions of long-context requests concurrently requires planetary-scale memory orchestration and cluster-level caching.
The KV Cache Is a Hidden Working Set
Much like an operating system manages an active process working set in physical RAM, an LLM serving engine treats the KV cache as its active execution working set in GPU VRAM.
To counteract this working-set explosion, frontier serving architectures deploy structural KV compression techniques:
- Multi-Head Latent Attention (MLA): Instead of caching full Key and Value projection matrices, MLA compresses $K$ and $V$ vectors into a shared, low-rank latent subspace prior to generation. This drastically diminishes the physical bytes stored per token while preserving associative retrieval fidelity.
- Write-Time Dynamic Routing (e.g., LongFlow and Write-Gated KV): Rather than caching all historical tokens indefinitely, heuristic scoring evaluates token importance at write time. High-utility tokens are retained in a long-term global cache, while transient tokens are routed to a compact local sliding window, shrinking the active KV working set by up to 80% with minimal loss in model accuracy.
Why This Matters More at Scale
At enterprise scale, optimizing KV cache allocation by even 10-15% directly translates into serving thousands of additional concurrent users without purchasing extra GPU clusters.
The Architecture Has Now Acquired a Memory Plane
INCOMING REQUEST
│
▼
Global Router Plane
│
(Cache-Aware Routing)
│
▼
Scheduler & Queue
│
┌──────────────┴──────────────┐
▼ ▼
Prefill Engine Decode Engine
(Compute-Bound) (Memory-Bandwidth Bound)
│ │
▼ ▼
Build KV State Read/Extend KV
│ │
└──────────────┬──────────────┘
▼
Paged KV Memory Plane
│
▼
Fused GPU Kernels
│
▼
Streaming OutputThe Deeper Systems Lesson
Every performance optimization in distributed systems creates a new resource-allocation challenge: caching eliminates compute redundancy but introduces memory pressure; continuous batching maximizes hardware utilization but demands dynamic memory paging.
What We Have Solved — and What We Have Not
We have deconstructed tokenization, attention mathematics, prefill vs. decode workloads, continuous batching, and KV memory management. The remaining frontier is orchestrating these mechanisms across physical distributed clusters.
The Next Problem: Continuous Batching at the Token Level
At every microsecond, an inference engine must arbitrate between arriving prefill requests, active decoding streams, memory block allocations, and latency SLAs.
The Architecture We Are Building
USER PROMPT
│
▼
Application Layer
│
Gateway & Routing
│
Cluster Scheduler
│
┌──────────────┴──────────────┐
▼ ▼
Prefill Decode
│ │
└──────────────┬──────────────┘
▼
Paged KV Cache Memory
│
Fused GPU Kernels
│
▼
GENERATED TOKENSFinal Takeaways
ChatGPT's speed and reliability are not the result of a single neural network breakthrough; they are the cumulative outcome of deep systems engineering across memory paging, IO-aware attention kernels, iteration-level continuous batching, and distributed cache-aware routing.
References
- OpenAI — How GPT-5.6 fuses frontier intelligence with frontier efficiency — Inference architecture, KV cache, prefill/decode behavior, routing, scheduling, batching, cache availability, workload-specific optimization and speculative decoding.
- OpenAI — Prompt caching — KV tensors, prompt-prefix reuse, cache storage, cache lifetime, cache matching and prompt-cache behavior.
- OpenAI — API deployment checklist — Stable prompt prefixes, cache breakpoints, prompt structure, cache-aware routing and production caching practices.
- Kwon, W. et al. — Efficient Memory Management for Large Language Model Serving with PagedAttention — KV-cache memory fragmentation, dynamic allocation, paging, memory efficiency and high-throughput LLM serving.
- Prabhu, P. et al. — vAttention: Dynamic Memory Management for Serving LLMs — Alternative approach to KV-cache memory management and dynamic allocation.
- Ouyang, L. et al. — Training language models to follow instructions with human feedback — InstructGPT architecture, Supervised Fine-Tuning (SFT), Reward Model (RM) training, and PPO alignment.
- Rafailov, R. et al. — Direct Preference Optimization: Your Language Model is Secretly a Reward Model — Direct Preference Optimization (DPO), implicit reward derivation, and stable offline alignment without explicit reward models.
- Shah, J. et al. — FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision — Warp specialization, FP8 block quantization, incoherent processing, and asynchronous memory pipelining for Hopper H100 GPUs.
- OpenAI — OpenAI o1 System Card — Safety evaluation of reasoning models, test-time deduction monitoring, automated red teaming, and deceptive alignment analysis.
- DeepSeek-AI — DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — Large-scale test-time compute scaling, pure reinforcement learning without human cold-start SFT, and Multi-head Latent Attention (MLA).
