Appearance
Inference vs. Training vs. Fine-Tuning
This is the most frequently asked — and most easily confused — cluster of concepts. Here is the conclusion up front: training is "teaching the model," fine-tuning is "a lightweight epilogue of training," and inference is "putting the graduated model on the job." All three run a neural network forward, but engineering-wise they are nearly three different fields — compute distribution, memory-access profile, batching shape, precision choices, and hardware utilization all differ.
Many people equate "can run transformers to train a model" with "understands inference engineering," and end up answering an inference-role interview as if it were a training-role interview. This page draws the three apart across seven dimensions, each with a quantifiable difference.
1. The Essential Difference: FLOPs, Memory Access, Activations
Break down the "forward + backward" compute of a single Transformer layer:
| Dimension | Training (forward + backward + optimizer update) | Inference (forward only) |
|---|---|---|
| FLOPs / token | forward ≈ 2P, backward ≈ 4P, optimizer ≈ 2P (Adam), total ≈ 8P | 2P (forward only) |
| Activation memory | Must keep all forward activations for backward (O(seq_len × hidden)) | No activations kept; discarded after use |
| Memory-access profile | Compute-bound (large batch, high arithmetic intensity) | Memory-bound (small batch, low arithmetic intensity, especially LLM decode) |
| GPU utilization (MFU) | 50-60% (training large models on A100) | 5-50% (<5% at batch=1 decode, 50%+ at prefill) |
Here P is the parameter count. A direct corollary: training 1 token costs 4x the FLOPs of inference — but that is only the tip of the iceberg. The real gap lies in the shape of the compute distribution:
- Training: compute within a batch is dense and uniform, every token is treated equally, and GPU utilization can be pinned at its ceiling.
- Inference: the prefill phase is compute-dense (processing the prompt), the decode phase is compute-sparse (one token per step), and the arithmetic intensity of the two phases of the same request differs by more than 100x.
This is why training engineers care about "what is your MFU" while inference engineers care about "what is your tokens/s, how much HBM bandwidth are you using" — the two are not even watching the same kind of metric. See Latency, Throughput, and Concurrency and The GPU Memory Hierarchy and the Bandwidth Wall.
Backprop is not "forward computed twice"
Many assume backward = two forward passes, so training costs 2x inference. In reality backward must compute gradients for both weights and activations, plus Adam's first- and second-moment updates — total FLOPs are 4-6x inference. This gap dictates that training is always "compute-first" while inference is always "memory-access-first."
2. Batching: Static for Training, Continuous for Inference
Batching is the biggest structural difference between training and inference engineering.
Training: Static Batching, the Bigger the Better
A training batch is fixed within a step: sample N examples from the dataset, pad to equal length, run one forward+backward, update parameters, then sample the next N in the next step. The larger N, the higher the arithmetic intensity and the higher the GPU utilization — so a training engineer's daily work is "find a way to enlarge the batch" (gradient accumulation, ZeRO sharding, tensor parallelism to make it fit in memory).
Training batch sizes range from dozens to thousands, and samples within one batch are usually padded to the longest, which is an acceptable waste.
Inference: Requests Arrive Any Time, at Any Length — Static Batching Explodes
Inference requests arrive whenever users come, and prompt lengths range from dozens to thousands of tokens. With static batching:
- Option A: wait until the batch fills before computing — latency explodes (users wait seconds to see the first token);
- Option B: run each request immediately on its own — throughput explodes (GPU utilization <5%);
- Option C: fill a fixed batch and pad to the longest — short requests are dragged down by long ones, wasting 50%+ of compute.
vLLM's continuous batching solves this: let requests join and leave the batch dynamically — the moment a request generates its last token it exits, a new request immediately takes its slot, and the other requests keep decoding. Combined with PagedAttention paging the KV cache, GPU memory utilization goes from ~20% to ~90% and throughput improves 10-20x.
text
t=0: [A][B][C] <- three requests decoding together
t=1: [A][B][C][D] <- D arrives and joins the batch
t=2: [A][B]__[D] <- C finishes generating and exits; A, B, D continue
t=3: [A][B][D][E] <- E arrives and joinsContinuous batching is the core mechanism of vLLM, TGI, and SGLang; see Batching and Request Scheduling.
3. KV Cache: Absent in Training, Mandatory in Autoregressive Inference
The KV cache is the lifeline of autoregressive LLM inference, yet it does not exist in training at all.
Why Training Needs No KV Cache
In training, a whole sequence enters at once and attention computes Q×K^T×V over the entire sequence in one shot — there is no concept of a "next step," so nothing needs caching. Training's attention memory footprint is O(seq_len²) (the attention matrix), which FlashAttention compresses to O(seq_len).
Why Inference Must Have a KV Cache
Autoregressive inference "generates one new token at a time," and at step N it must compute attention(Q_N, K_{1..N}, V_{1..N}). Without a KV cache, every generated token would recompute the K and V of all previous tokens — O(N²) redundant compute; for 200 tokens on a 1000-token prompt, the recomputation is roughly 1000x what is actually needed.
The essence of the KV cache: cache each layer's and each head's K and V tensors, and at each subsequent step compute Q/K/V only for the new token, appending to the cache. This drops per-step compute from O(N) to O(1), but memory footprint becomes O(N × layers × hidden × heads) — for a 70B model with 4k of context, a single request's KV cache needs ~5 GB.
The Engineering Cost of the KV Cache
A rough formula for KV cache memory:
KV_cache_size = 2 (K and V) × num_layers × seq_len × hidden_dim × num_kv_heads × dtype_bytesTake Llama-2-70B (80 layers, hidden 8192, 64 heads, 8 KV heads, FP16) with 4k of context for a single request:
2 × 80 × 4096 × 8192 × 8 × 2 bytes ≈ 8.6 GBThis means an 80 GB A100 that holds the 70B weights (140 GB in FP16 does not fit, so INT4 quantization down to 35 GB is required) has only ~45 GB left — enough for the KV caches of about 5 concurrent requests. This is exactly the problem PagedAttention solves — paging the KV cache to eliminate fragmentation and raise concurrency 5-10x.
4. Precision: BF16/FP32 for Training; INT8/INT4/FP8 for Inference
Precision choice is one of the most visible splits between training and inference.
| Stage | Main precision | Why |
|---|---|---|
| Training | BF16 / FP32 (partly mixed) | Backpropagation is sensitive to the numerical range of gradients; BF16's 8-bit exponent + 7-bit mantissa balances range and precision |
| Inference | FP16 (baseline) → INT8 / INT4 / FP8 (optimized) | Forward needs no gradients; weights and activations can be quantized to low precision, gaining memory access and compute at once |
Training almost never uses INT8/INT4 — quantization error explodes through backpropagation. Inference uses only the forward pass, where low precision's "small error" is acceptable. See Model Quantization Fundamentals.
The key differences in how precision evolves:
- Training precision evolves slowly: FP32 → FP16 → BF16, each taking 3-5 years to spread (BF16 only got hardware support on Ampere in 2020).
- Inference precision evolves fast: INT8 was hardware-supported on Volta in 2018, INT4 was engineered for LLMs in 2023 (GPTQ/AWQ), FP8 spread on H100 in 2024 — each hardware generation brings a new precision tier.
Why inference prefers FP8
After H100 introduced FP8 (E4M3 / E5M2), FP8 became more popular than INT8 for LLM inference — because it keeps the floating-point exponent range, is insensitive to activation outliers, loses less accuracy than INT8, and is natively supported by Tensor Cores. See Weight-Only Quantization and Mixed Precision.
5. Hardware Utilization: Training Is Large-Cluster and Communication-Dense; Inference Is Single-Machine Multi-GPU + Batching + KV Cache
Training and inference have almost opposite preferences in hardware topology.
Training: Large Clusters, Communication-Dense
Training large models requires many machines and many GPUs, and communication is the bottleneck — data parallelism needs to AllReduce gradients, tensor parallelism needs to AllReduce/AllGather activations, and pipeline parallelism needs to pass activations P2P. Training bottlenecks often sit in communication, not compute, so the training cluster's NICs (400G InfiniBand) and topology (NVLink, NVSwitch) are critical.
Training cluster scale: GPT-3 was trained on thousands of V100s, Llama-3-405B on 16k+ H100s; cluster communication design is the core of training engineering.
Inference: Mostly Single-Machine Multi-GPU, Fed by Batching + KV Cache
The vast majority of inference scenarios do not need a large cluster — a 70B model quantized to INT4 is 35 GB, fits on a single 80 GB A100, and still has room to serve 10-20 concurrent requests with the KV cache. Inference hardware optimization focuses not on communication but on:
- Batching: pack multiple requests into one forward to raise arithmetic intensity.
- KV cache management: PagedAttention reduces fragmentation.
- Operator optimization: FlashAttention, FlashInfer raise single-operator efficiency.
- Quantization: less memory, less memory traffic, more compute.
Only when a single GPU cannot hold the model (e.g. a 405B model) or a single machine cannot handle the traffic does inference use tensor parallelism (TP) / pipeline parallelism (PP). Inference TP/PP is simpler than training's — no gradient synchronization and no pipeline bubbles; you only need to shard weights during the forward pass. See Distributed Inference (TP/PP).
6. Where Fine-Tuning Sits: LoRA/QLoRA Are Training, but Often the Last Step Before Inference Serving
Fine-tuning is training, but in the modern LLM engineering pipeline it often appears "right before the inference service goes live," so it is worth pinning down its position.
The Essence of Fine-Tuning
Fine-tuning is "continuing to train on an already pretrained model with a small learning rate and a small dataset." Its compute profile sits between training and inference — there are gradient updates, but the batch is small, the step count is low, and the compute is far below pretraining.
LoRA and QLoRA: Making Fine-Tuning Fit on Inference-Grade GPUs
- LoRA (Low-Rank Adaptation): freeze the base model weights and train only a low-rank matrix A×B (rank typically 8-64). Trainable parameters drop from 70B to a few tens of millions, so a 24 GB GPU can fine-tune a 7B model.
- QLoRA: quantize the base model to 4-bit (NF4) and train only the LoRA part. Lets a 70B model be fine-tuned on a single 48 GB GPU.
QLoRA is the typical case of "quantization on the training side" — but its quantization applies only to the frozen base model, while the LoRA part is still trained in BF16/FP32, avoiding the low-precision gradient explosion. This is quantization's "boundary usage on the training side," not inference-specific.
Where Fine-Tuning Sits in the Deployment Pipeline
A typical "model weights → production service" pipeline:
text
Pretrained weights (public repo)
|
Fine-tuning (LoRA/QLoRA, optional) <- training, but often the customization step before deployment
|
Merge LoRA / convert format
|
Quantization (GPTQ/AWQ/FP8) <- first step of inference optimization
|
Engine conversion (TensorRT-LLM/vLLM deployment)
|
Online serving (with batching/KV cache/routing)Fine-tuning is training, but it is often the last customization step before deployment — business-domain adaptation, style alignment, safety fine-tuning (RLHF/DPO). Understanding this pipeline explains "why LoRA repos so often sit next to inference engines."
Fine-tuning is not inference optimization
Note: fine-tuning itself is not an inference-acceleration technique. It changes model weights, not the inference engine. But when fine-tuning is combined with quantization, the quantization parameters need recalibration — otherwise you get "fine-tuned weights degrading badly under INT4 quantization." This is one of the most common hidden traps in the deployment pipeline; see Common Pitfalls and Anti-Patterns.
7. Comparison Table + Summary
| Dimension | Training | Fine-Tuning | Inference |
|---|---|---|---|
| Goal | Teach the model the patterns | Adapt an already-trained model to a domain | Turn model weights into a production service |
| FLOPs / token | ~8P (forward+backward+optimizer) | ~8P (same as training, but few steps) | ~2P (forward only) |
| Compute distribution | Dense and uniform | Dense and uniform | Dense prefill, sparse decode |
| Batching | Static batching, the bigger the better | Static batching, medium batch | Continuous batching (vLLM/TGI/SGLang) |
| KV cache | Not needed (whole sequence at once) | Not needed | Mandatory (autoregressive generation) |
| Precision | BF16/FP32 | BF16/FP32 (QLoRA excepted) | FP16/INT8/INT4/FP8 |
| Activation memory | Keeps all forward activations | Same as training | Not kept |
| Hardware topology | Large cluster + high-speed interconnect | Single/multi-GPU, low requirements | Mostly single-machine multi-GPU, fed by batching |
| Typical tools | Megatron-LM, DeepSpeed, FSDP | PEFT, Unsloth, Axolotl | vLLM, TensorRT-LLM, SGLang, TGI |
| Metrics watched | MFU, loss, convergence | Convergence, domain metrics | Latency, throughput, memory, SLO |
| Position in the pipeline | First step (pretraining) | Pre-deployment customization (optional) | Goes live to serve |
Summary in one sentence: training cares about "can compute be pinned to its ceiling," inference cares about "can memory access be fed fast enough"; fine-tuning is a lightweight epilogue of training, but often the last customization step before the inference service goes live. The three are nearly three different fields in engineering terms, and understanding this boundary is the first watershed in telling whether an engineer "knows training or knows deployment."
Further Reading
- What Is Inference Acceleration? — the conceptual foundation this page builds on
- Anatomy of the Overall Architecture — the five-layer architecture of an inference system
- Latency, Throughput, and Concurrency — the core metric family on the inference side
- The GPU Memory Hierarchy and the Bandwidth Wall — why LLM inference is memory-bound
- Model Quantization Fundamentals — the engineering logic behind precision choices
- Batching and Request Scheduling — the full mechanism of continuous batching
- vLLM and PagedAttention — the benchmark case of KV cache management
- Distributed Inference (TP/PP) — the simplified forms of TP/PP on the inference side
References
- Kwon et al. Efficient Memory Management for LLM Serving with PagedAttention (SOSP 2023) — the engineering paper behind continuous batching and PagedAttention
- Hu et al. LoRA: Low-Rank Adaptation of Large Language Models (ICLR 2022) — the original LoRA paper
- Dettmers et al. QLoRA: Efficient Finetuning of Quantized LLMs (NeurIPS 2023) — QLoRA pushes training-side quantization to 4-bit
- NVIDIA. H100 Tensor Core GPU Architecture White Paper — the hardware foundation of FP8 and the Transformer Engine
- Dao et al. FlashAttention (NeurIPS 2022) — the operator-optimization benchmark shared by training and inference
- Narayanan et al. Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM (SC 2021) — compute and communication utilization on the training side