Appearance
GPUs and Hardware Selection
The one-sentence definition: hardware selection is finding a machine that satisfies "model size × latency requirement × throughput demand × budget"—VRAM sets the capacity floor, bandwidth sets the speed ceiling, and compute sets peak throughput.
Industry insight: the vast majority of deployment incidents aren't code problems—they're hardware mismatches: cramming a 7B model into a 16GB T4 and OOM-restarting all day; requiring 5ms risk-control latency but picking a bare CPU to save budget. The opposite failure is just as common—over-provisioning: an A100 serving a model that gets 10,000 calls a day sits below 5% utilization, burning serious money in electricity every year. A deployment engineer who knows how to do the math spends half their time calculating VRAM and bandwidth, and the other half load-testing to verify their own estimates.
1. VRAM: The Hard Constraint—Do This Math First
1.1 The Model VRAM Estimation Formula
text
Model weight VRAM ≈ parameter count × bytes per parameter
Runtime VRAM ≈ weight VRAM × 1.2 headroom (activations + intermediate tensors + engine overhead)Bytes per parameter by precision: FP32=4, FP16/BF16=2, INT8=1, INT4=0.5.
Inference VRAM for common LLMs (weights only, excluding the KV cache):
| Model | Parameters | FP16 | INT8 | INT4 |
|---|---|---|---|---|
| Llama-7B | 7B | ~14 GB | ~7 GB | ~3.5 GB |
| Llama-13B | 13B | ~26 GB | ~13 GB | ~6.5 GB |
| Llama-70B | 70B | ~140 GB | ~70 GB | ~35 GB |
Don't Forget the KV Cache and Activations
For LLM inference, the KV cache grows linearly with concurrency and can end up larger than the weights (a 7B model serving 128 concurrent requests can consume 40GB+ of KV cache alone). The complete "VRAM ledger" is covered in LLM Inference Optimization.
1.2 Three Ways Out When VRAM Falls Short (In Priority Order)
- Quantize: FP16 → INT8 saves half outright, INT4 saves 75%—see Quantization;
- Offload: keep inactive weights in CPU memory/on disk, trading latency for capacity;
- Bigger or more cards: tensor parallelism splits the weights across GPUs—see LLM Inference Optimization.
2. Key GPU Specs: Three to Check First, One to Check Last
| Metric | Unit | Meaning | Impact on inference |
|---|---|---|---|
| VRAM capacity | GB | How large a model fits | Hard constraint—fall short and you OOM |
| Memory bandwidth | GB/s | How much data moves per second | Speed of memory-bound operators (inference's main battleground) |
| Compute | TFLOPS (FP16/INT8) | Operations per second | Ceiling for compute-bound operators (training/large batches) |
| NVLink interconnect | GB/s | Multi-GPU communication bandwidth | Efficiency of multi-GPU parallelism (tensor parallelism) |
So what's the reading order? For online inference: VRAM capacity first, memory bandwidth second, compute last. Inference is bandwidth-intensive (see Inference: From Forward Pass to Inference Engines), so a mid-range card with high bandwidth often beats a compute-heavy flagship for small-batch online serving.
Inference Is Bandwidth-Intensive; Training Is Compute-Intensive
Through the Roofline lens: every operator has an "arithmetic intensity" (FLOP/byte). Small-batch inference has extremely low arithmetic intensity (a few operations per byte of weights read), so the bottleneck is bandwidth; large-batch training has high arithmetic intensity, so the bottleneck is compute. This explains two observations:
- The gap between a T4 (320 GB/s) and an A100 (2 TB/s) on small-batch inference is far smaller than the gap in their compute specs (65 vs 312 TFLOPS);
- INT8 quantization gives inference nearly a 2× speedup (half the bytes to move) but means little for training.
3. Comparing Mainstream Inference GPUs
Using models on the market as of 2024 (specs are approximations from public datasheets—defer to official whitepapers before purchasing):
| GPU | VRAM | Memory bandwidth | FP16 compute | Typical use |
|---|---|---|---|---|
| T4 | 16 GB | 320 GB/s | 65 TFLOPS | Entry-level online inference, batch small models |
| L4 | 24 GB | 300 GB/s | 121 TFLOPS | T4 upgrade path, video/vision |
| A10 | 24 GB | 600 GB/s | 125 TFLOPS | Mid-size online inference, multiple small models |
| L40S | 48 GB | 864 GB/s | 362 TFLOPS | Both LLM inference and training |
| A100 40G/80G | 40/80 GB | 1.5–2 TB/s | 312 TFLOPS | LLMs, tensor parallelism, training |
| H100 | 80 GB | 3.35 TB/s | 989 TFLOPS (with sparsity) | Flagship inference and training |
| RTX 4090 (consumer) | 24 GB | 1 TB/s | 330 TFLOPS | Development, internal tools, small services (no ECC/no vGPU) |
Consumer vs Professional Cards
The RTX 4090 is exceptional value (nearly twice the performance of an A10), but it has no ECC memory, no vGPU virtualization support, and no datacenter-grade cooling or stability certification, and its license terms prohibit datacenter deployment. Fine for experiments and development—for production, buy professional cards.
Rules of thumb (models < 20B, online inference):
text
FP16 model fits on one card (≤ 24GB) and QPS is modest → L4 / A10
FP16 model fits on one card but QPS is high (>200) → L40S / A100
Model doesn't fit on one card (≥ 70B) → multi-GPU tensor parallelism (H100/A100×N)4. CPU Inference: The Underrated Second Front
CPU inference has come back into fashion in the LLM era (llama.cpp runs 7B quantized models at 5–10 tokens/s on an ordinary laptop). Its optimization levers:
- Instruction sets: AVX2 → AVX-512 → AMX (Intel)—up to a 3–5× difference;
- Libraries: oneDNN (formerly MKL-DNN), OpenBLAS, and llama.cpp's own ggml;
- Cache friendliness: repack weights into CPU-friendly memory layouts to avoid cache misses.
When to pick CPU: low QPS + small models + budget constraints, plus edge environments with no GPU. The decision rule is in Section 4 of Inference: From Forward Pass to Inference Engines.
5. NPUs, TPUs, and Custom Silicon
| Chip | Examples | Positioning | Inference performance |
|---|---|---|---|
| Edge NPUs | Huawei Ascend Atlas, RK3588 NPU, Apple Neural Engine | On-device/edge AI | Excellent energy efficiency, but narrow operator support |
| Cloud TPUs | Google TPU v5e | Cloud LLM training and inference | Strong compute, tied to the TensorFlow/JAX ecosystem |
| Other ASICs | AWS Inferentia, Groq LPU | Dedicated inference | Inferentia offers good value; the ecosystem constraints are real |
Selection logic: the allure of NPUs/ASICs is cost; the risk is ecosystem lock-in. Once a model's operators aren't supported, you either rewrite or walk away. Most teams "run GPU as the baseline and pilot NPUs in specific scenarios" rather than migrating everything up front.
6. Multi-GPU Setups: The Three Parallelisms
When a model doesn't fit on one card, or one card can't deliver the throughput, go multi-GPU. The three roles (see LLM Inference Optimization and Paper: Parallel and Distributed Inference):
text
Data parallelism: full model on every card, different batches ── solves "not enough throughput"
Tensor parallelism: weights sliced by layer across cards, frequent communication ── solves "doesn't fit on one card"
Pipeline parallelism: layer groups across cards, activations passed between ── solves "doesn't fit + cheaper communication"Key point: every card you add to tensor parallelism adds communication overhead; 2–8 cards is usually the sweet spot; beyond 8, combine with pipeline parallelism. Multi-GPU setups require high-speed interconnects like NVLink/InfiniBand, or communication eats the gains.
7. Cloud Selection Cheat Sheet
Rough estimates by "model size × target QPS" (single instance, assuming dynamic batching is active):
| Model size (FP16 weights) | QPS < 10 | QPS 10–100 | QPS 100–1000 | QPS > 1000 |
|---|---|---|---|---|
| < 1 GB (small CNN/MLP) | 2–4 core CPU | 8-core CPU / T4 | T4×1–2 | L4×2+ |
| 1–7 GB (BERT-class) | T4 | T4 / L4 | L4×1–2 | L40S×2+ |
| 7–30 GB (7B–13B LLM) | L4 | L4 / L40S | L40S / A100 | A100×2+ (tensor parallel) |
| > 30 GB (70B+) | Not realistic | A100/H100×2 | A100×2–4 | H100×4+ (tensor + pipeline) |
The Table Is Only a Starting Point
Real QPS ceilings depend on the model, batch size, and latency budget—you must calibrate with load tests. Method: see Load Testing and Capacity Planning—first find one instance's max QPS, then derive the instance count from the target QPS, and keep 30% headroom for traffic spikes.
A Complete Worked Example
Goal: deploy a 13B model (~26GB in FP16) with P99 latency < 1s and 1 million calls per day (average QPS ≈ 12, peak ≈ 60).
text
1) VRAM: 26GB FP16 weights + activations and KV cache → need ≥ 48GB per card → L40S/A100
2) Bandwidth: at 60 QPS peak, continuous batching yields 1000+ token/s → one card's bandwidth suffices
3) Compute: LLM generation is bandwidth-bound; compute isn't the bottleneck
4) Verdict: start with 1× L40S (48GB), scale to 2× after load testingTrade-offs
| Decision point | Options | How to choose |
|---|---|---|
| Capacity first vs speed first | More VRAM vs more bandwidth | Satisfy capacity first, then trade off between bandwidth and compute |
| GPU vs CPU | GPU fast but pricey; CPU cheap but slow | High QPS or latency-sensitive → GPU; low QPS small models → CPU |
| Professional vs consumer cards | Stability vs price | Professional for production; consumer is fine for development |
| Single GPU vs multi-GPU | Simple vs capacity/throughput | Quantize to save VRAM first; go multi-GPU only if that's not enough |
| GPU vs NPU/ASIC | Ecosystem vs cost | GPU as the baseline; pilot ASICs in stable scenarios |
One-line summary: VRAM decides whether it runs, bandwidth decides how fast, compute decides how much it can scale—do the math, pick the card, then calibrate with load tests.
Further Reading
- Inference: From Forward Pass to Inference Engines — the Roofline analysis of bandwidth vs compute
- Quantization — the first-priority fix when VRAM runs short
- LLM Inference Optimization — the KV cache memory ledger and multi-GPU parallelism
- Performance Optimization and Capacity Planning — how load tests calibrate hardware estimates
- Load Testing and Capacity Planning — the full method from estimation to measurement