Skip to content

Inference vs. Training vs. Fine-Tuning

At a glance Training counts FLOPs and activations; inference is memory-access-heavy, batching-sensitive, and requires a KV cache. Fine-tuning is a lightweight epilogue of training. This page draws the three apart across seven dimensions — compute, memory access, batching, KV cache, precision, hardware, and where fine-tuning sits.

Inference vs. Training vs. Fine-Tuning ​

This is the most frequently asked — and most easily confused — cluster of concepts. Here is the conclusion up front: training is "teaching the model," fine-tuning is "a lightweight epilogue of training," and inference is "putting the graduated model on the job." All three run a neural network forward, but engineering-wise they are nearly three different fields — compute distribution, memory-access profile, batching shape, precision choices, and hardware utilization all differ.

Many people equate "can run transformers to train a model" with "understands inference engineering," and end up answering an inference-role interview as if it were a training-role interview. This page draws the three apart across seven dimensions, each with a quantifiable difference.

1. The Essential Difference: FLOPs, Memory Access, Activations ​

Break down the "forward + backward" compute of a single Transformer layer:

DimensionTraining (forward + backward + optimizer update)Inference (forward only)
FLOPs / tokenforward ≈ 2P, backward ≈ 4P, optimizer ≈ 2P (Adam), total ≈ 8P2P (forward only)
Activation memoryMust keep all forward activations for backward (O(seq_len × hidden))No activations kept; discarded after use
Memory-access profileCompute-bound (large batch, high arithmetic intensity)Memory-bound (small batch, low arithmetic intensity, especially LLM decode)
GPU utilization (MFU)50-60% (training large models on A100)5-50% (<5% at batch=1 decode, 50%+ at prefill)

Here P is the parameter count. A direct corollary: training 1 token costs 4x the FLOPs of inference — but that is only the tip of the iceberg. The real gap lies in the shape of the compute distribution:

  • Training: compute within a batch is dense and uniform, every token is treated equally, and GPU utilization can be pinned at its ceiling.
  • Inference: the prefill phase is compute-dense (processing the prompt), the decode phase is compute-sparse (one token per step), and the arithmetic intensity of the two phases of the same request differs by more than 100x.

This is why training engineers care about "what is your MFU" while inference engineers care about "what is your tokens/s, how much HBM bandwidth are you using" — the two are not even watching the same kind of metric. See Latency, Throughput, and Concurrency and The GPU Memory Hierarchy and the Bandwidth Wall.

Backprop is not "forward computed twice"

Many assume backward = two forward passes, so training costs 2x inference. In reality backward must compute gradients for both weights and activations, plus Adam's first- and second-moment updates — total FLOPs are 4-6x inference. This gap dictates that training is always "compute-first" while inference is always "memory-access-first."

2. Batching: Static for Training, Continuous for Inference ​

Batching is the biggest structural difference between training and inference engineering.

Training: Static Batching, the Bigger the Better ​

A training batch is fixed within a step: sample N examples from the dataset, pad to equal length, run one forward+backward, update parameters, then sample the next N in the next step. The larger N, the higher the arithmetic intensity and the higher the GPU utilization — so a training engineer's daily work is "find a way to enlarge the batch" (gradient accumulation, ZeRO sharding, tensor parallelism to make it fit in memory).

Training batch sizes range from dozens to thousands, and samples within one batch are usually padded to the longest, which is an acceptable waste.

Inference: Requests Arrive Any Time, at Any Length — Static Batching Explodes ​

Inference requests arrive whenever users come, and prompt lengths range from dozens to thousands of tokens. With static batching:

  • Option A: wait until the batch fills before computing — latency explodes (users wait seconds to see the first token);
  • Option B: run each request immediately on its own — throughput explodes (GPU utilization <5%);
  • Option C: fill a fixed batch and pad to the longest — short requests are dragged down by long ones, wasting 50%+ of compute.

vLLM's continuous batching solves this: let requests join and leave the batch dynamically — the moment a request generates its last token it exits, a new request immediately takes its slot, and the other requests keep decoding. Combined with PagedAttention paging the KV cache, GPU memory utilization goes from ~20% to ~90% and throughput improves 10-20x.

text
t=0:  [A][B][C]          <- three requests decoding together
t=1:  [A][B][C][D]       <- D arrives and joins the batch
t=2:  [A][B]__[D]        <- C finishes generating and exits; A, B, D continue
t=3:  [A][B][D][E]       <- E arrives and joins

Continuous batching is the core mechanism of vLLM, TGI, and SGLang; see Batching and Request Scheduling.

3. KV Cache: Absent in Training, Mandatory in Autoregressive Inference ​

The KV cache is the lifeline of autoregressive LLM inference, yet it does not exist in training at all.

Why Training Needs No KV Cache ​

In training, a whole sequence enters at once and attention computes Q×K^T×V over the entire sequence in one shot — there is no concept of a "next step," so nothing needs caching. Training's attention memory footprint is O(seq_len²) (the attention matrix), which FlashAttention compresses to O(seq_len).

Why Inference Must Have a KV Cache ​

Autoregressive inference "generates one new token at a time," and at step N it must compute attention(Q_N, K_{1..N}, V_{1..N}). Without a KV cache, every generated token would recompute the K and V of all previous tokens — O(N²) redundant compute; for 200 tokens on a 1000-token prompt, the recomputation is roughly 1000x what is actually needed.

The essence of the KV cache: cache each layer's and each head's K and V tensors, and at each subsequent step compute Q/K/V only for the new token, appending to the cache. This drops per-step compute from O(N) to O(1), but memory footprint becomes O(N × layers × hidden × heads) — for a 70B model with 4k of context, a single request's KV cache needs ~5 GB.

The Engineering Cost of the KV Cache ​

A rough formula for KV cache memory:

KV_cache_size = 2 (K and V) × num_layers × seq_len × hidden_dim × num_kv_heads × dtype_bytes

Take Llama-2-70B (80 layers, hidden 8192, 64 heads, 8 KV heads, FP16) with 4k of context for a single request:

2 × 80 × 4096 × 8192 × 8 × 2 bytes ≈ 8.6 GB

This means an 80 GB A100 that holds the 70B weights (140 GB in FP16 does not fit, so INT4 quantization down to 35 GB is required) has only ~45 GB left — enough for the KV caches of about 5 concurrent requests. This is exactly the problem PagedAttention solves — paging the KV cache to eliminate fragmentation and raise concurrency 5-10x.

4. Precision: BF16/FP32 for Training; INT8/INT4/FP8 for Inference ​

Precision choice is one of the most visible splits between training and inference.

StageMain precisionWhy
TrainingBF16 / FP32 (partly mixed)Backpropagation is sensitive to the numerical range of gradients; BF16's 8-bit exponent + 7-bit mantissa balances range and precision
InferenceFP16 (baseline) → INT8 / INT4 / FP8 (optimized)Forward needs no gradients; weights and activations can be quantized to low precision, gaining memory access and compute at once

Training almost never uses INT8/INT4 — quantization error explodes through backpropagation. Inference uses only the forward pass, where low precision's "small error" is acceptable. See Model Quantization Fundamentals.

The key differences in how precision evolves:

  • Training precision evolves slowly: FP32 → FP16 → BF16, each taking 3-5 years to spread (BF16 only got hardware support on Ampere in 2020).
  • Inference precision evolves fast: INT8 was hardware-supported on Volta in 2018, INT4 was engineered for LLMs in 2023 (GPTQ/AWQ), FP8 spread on H100 in 2024 — each hardware generation brings a new precision tier.

Why inference prefers FP8

After H100 introduced FP8 (E4M3 / E5M2), FP8 became more popular than INT8 for LLM inference — because it keeps the floating-point exponent range, is insensitive to activation outliers, loses less accuracy than INT8, and is natively supported by Tensor Cores. See Weight-Only Quantization and Mixed Precision.

5. Hardware Utilization: Training Is Large-Cluster and Communication-Dense; Inference Is Single-Machine Multi-GPU + Batching + KV Cache ​

Training and inference have almost opposite preferences in hardware topology.

Training: Large Clusters, Communication-Dense ​

Training large models requires many machines and many GPUs, and communication is the bottleneck — data parallelism needs to AllReduce gradients, tensor parallelism needs to AllReduce/AllGather activations, and pipeline parallelism needs to pass activations P2P. Training bottlenecks often sit in communication, not compute, so the training cluster's NICs (400G InfiniBand) and topology (NVLink, NVSwitch) are critical.

Training cluster scale: GPT-3 was trained on thousands of V100s, Llama-3-405B on 16k+ H100s; cluster communication design is the core of training engineering.

Inference: Mostly Single-Machine Multi-GPU, Fed by Batching + KV Cache ​

The vast majority of inference scenarios do not need a large cluster — a 70B model quantized to INT4 is 35 GB, fits on a single 80 GB A100, and still has room to serve 10-20 concurrent requests with the KV cache. Inference hardware optimization focuses not on communication but on:

  • Batching: pack multiple requests into one forward to raise arithmetic intensity.
  • KV cache management: PagedAttention reduces fragmentation.
  • Operator optimization: FlashAttention, FlashInfer raise single-operator efficiency.
  • Quantization: less memory, less memory traffic, more compute.

Only when a single GPU cannot hold the model (e.g. a 405B model) or a single machine cannot handle the traffic does inference use tensor parallelism (TP) / pipeline parallelism (PP). Inference TP/PP is simpler than training's — no gradient synchronization and no pipeline bubbles; you only need to shard weights during the forward pass. See Distributed Inference (TP/PP).

6. Where Fine-Tuning Sits: LoRA/QLoRA Are Training, but Often the Last Step Before Inference Serving ​

Fine-tuning is training, but in the modern LLM engineering pipeline it often appears "right before the inference service goes live," so it is worth pinning down its position.

The Essence of Fine-Tuning ​

Fine-tuning is "continuing to train on an already pretrained model with a small learning rate and a small dataset." Its compute profile sits between training and inference — there are gradient updates, but the batch is small, the step count is low, and the compute is far below pretraining.

LoRA and QLoRA: Making Fine-Tuning Fit on Inference-Grade GPUs ​

  • LoRA (Low-Rank Adaptation): freeze the base model weights and train only a low-rank matrix A×B (rank typically 8-64). Trainable parameters drop from 70B to a few tens of millions, so a 24 GB GPU can fine-tune a 7B model.
  • QLoRA: quantize the base model to 4-bit (NF4) and train only the LoRA part. Lets a 70B model be fine-tuned on a single 48 GB GPU.

QLoRA is the typical case of "quantization on the training side" — but its quantization applies only to the frozen base model, while the LoRA part is still trained in BF16/FP32, avoiding the low-precision gradient explosion. This is quantization's "boundary usage on the training side," not inference-specific.

Where Fine-Tuning Sits in the Deployment Pipeline ​

A typical "model weights → production service" pipeline:

text
Pretrained weights (public repo)
    |
Fine-tuning (LoRA/QLoRA, optional)   <- training, but often the customization step before deployment
    |
Merge LoRA / convert format
    |
Quantization (GPTQ/AWQ/FP8)          <- first step of inference optimization
    |
Engine conversion (TensorRT-LLM/vLLM deployment)
    |
Online serving (with batching/KV cache/routing)

Fine-tuning is training, but it is often the last customization step before deployment — business-domain adaptation, style alignment, safety fine-tuning (RLHF/DPO). Understanding this pipeline explains "why LoRA repos so often sit next to inference engines."

Fine-tuning is not inference optimization

Note: fine-tuning itself is not an inference-acceleration technique. It changes model weights, not the inference engine. But when fine-tuning is combined with quantization, the quantization parameters need recalibration — otherwise you get "fine-tuned weights degrading badly under INT4 quantization." This is one of the most common hidden traps in the deployment pipeline; see Common Pitfalls and Anti-Patterns.

7. Comparison Table + Summary ​

DimensionTrainingFine-TuningInference
GoalTeach the model the patternsAdapt an already-trained model to a domainTurn model weights into a production service
FLOPs / token~8P (forward+backward+optimizer)~8P (same as training, but few steps)~2P (forward only)
Compute distributionDense and uniformDense and uniformDense prefill, sparse decode
BatchingStatic batching, the bigger the betterStatic batching, medium batchContinuous batching (vLLM/TGI/SGLang)
KV cacheNot needed (whole sequence at once)Not neededMandatory (autoregressive generation)
PrecisionBF16/FP32BF16/FP32 (QLoRA excepted)FP16/INT8/INT4/FP8
Activation memoryKeeps all forward activationsSame as trainingNot kept
Hardware topologyLarge cluster + high-speed interconnectSingle/multi-GPU, low requirementsMostly single-machine multi-GPU, fed by batching
Typical toolsMegatron-LM, DeepSpeed, FSDPPEFT, Unsloth, AxolotlvLLM, TensorRT-LLM, SGLang, TGI
Metrics watchedMFU, loss, convergenceConvergence, domain metricsLatency, throughput, memory, SLO
Position in the pipelineFirst step (pretraining)Pre-deployment customization (optional)Goes live to serve

Summary in one sentence: training cares about "can compute be pinned to its ceiling," inference cares about "can memory access be fed fast enough"; fine-tuning is a lightweight epilogue of training, but often the last customization step before the inference service goes live. The three are nearly three different fields in engineering terms, and understanding this boundary is the first watershed in telling whether an engineer "knows training or knows deployment."

Further Reading ​

References ​