Skip to content

GPUs and Hardware Selection

At a glance Whether a model runs at all, and how fast, is half decided by the hardware. This article covers the key GPU specs (VRAM, bandwidth, compute), VRAM-to-model-size arithmetic, CPU/GPU/NPU/TPU comparisons, and how to reason about instance selection for inference workloads.

GPUs and Hardware Selection ​

The one-sentence definition: hardware selection is finding a machine that satisfies "model size × latency requirement × throughput demand × budget"—VRAM sets the capacity floor, bandwidth sets the speed ceiling, and compute sets peak throughput.

Industry insight: the vast majority of deployment incidents aren't code problems—they're hardware mismatches: cramming a 7B model into a 16GB T4 and OOM-restarting all day; requiring 5ms risk-control latency but picking a bare CPU to save budget. The opposite failure is just as common—over-provisioning: an A100 serving a model that gets 10,000 calls a day sits below 5% utilization, burning serious money in electricity every year. A deployment engineer who knows how to do the math spends half their time calculating VRAM and bandwidth, and the other half load-testing to verify their own estimates.

1. VRAM: The Hard Constraint—Do This Math First ​

1.1 The Model VRAM Estimation Formula ​

text
Model weight VRAM ≈ parameter count × bytes per parameter
Runtime VRAM      ≈ weight VRAM × 1.2 headroom (activations + intermediate tensors + engine overhead)

Bytes per parameter by precision: FP32=4, FP16/BF16=2, INT8=1, INT4=0.5.

Inference VRAM for common LLMs (weights only, excluding the KV cache):

ModelParametersFP16INT8INT4
Llama-7B7B~14 GB~7 GB~3.5 GB
Llama-13B13B~26 GB~13 GB~6.5 GB
Llama-70B70B~140 GB~70 GB~35 GB

Don't Forget the KV Cache and Activations

For LLM inference, the KV cache grows linearly with concurrency and can end up larger than the weights (a 7B model serving 128 concurrent requests can consume 40GB+ of KV cache alone). The complete "VRAM ledger" is covered in LLM Inference Optimization.

1.2 Three Ways Out When VRAM Falls Short (In Priority Order) ​

  1. Quantize: FP16 → INT8 saves half outright, INT4 saves 75%—see Quantization;
  2. Offload: keep inactive weights in CPU memory/on disk, trading latency for capacity;
  3. Bigger or more cards: tensor parallelism splits the weights across GPUs—see LLM Inference Optimization.

2. Key GPU Specs: Three to Check First, One to Check Last ​

MetricUnitMeaningImpact on inference
VRAM capacityGBHow large a model fitsHard constraint—fall short and you OOM
Memory bandwidthGB/sHow much data moves per secondSpeed of memory-bound operators (inference's main battleground)
ComputeTFLOPS (FP16/INT8)Operations per secondCeiling for compute-bound operators (training/large batches)
NVLink interconnectGB/sMulti-GPU communication bandwidthEfficiency of multi-GPU parallelism (tensor parallelism)

So what's the reading order? For online inference: VRAM capacity first, memory bandwidth second, compute last. Inference is bandwidth-intensive (see Inference: From Forward Pass to Inference Engines), so a mid-range card with high bandwidth often beats a compute-heavy flagship for small-batch online serving.

Inference Is Bandwidth-Intensive; Training Is Compute-Intensive ​

Through the Roofline lens: every operator has an "arithmetic intensity" (FLOP/byte). Small-batch inference has extremely low arithmetic intensity (a few operations per byte of weights read), so the bottleneck is bandwidth; large-batch training has high arithmetic intensity, so the bottleneck is compute. This explains two observations:

  • The gap between a T4 (320 GB/s) and an A100 (2 TB/s) on small-batch inference is far smaller than the gap in their compute specs (65 vs 312 TFLOPS);
  • INT8 quantization gives inference nearly a 2× speedup (half the bytes to move) but means little for training.

3. Comparing Mainstream Inference GPUs ​

Using models on the market as of 2024 (specs are approximations from public datasheets—defer to official whitepapers before purchasing):

GPUVRAMMemory bandwidthFP16 computeTypical use
T416 GB320 GB/s65 TFLOPSEntry-level online inference, batch small models
L424 GB300 GB/s121 TFLOPST4 upgrade path, video/vision
A1024 GB600 GB/s125 TFLOPSMid-size online inference, multiple small models
L40S48 GB864 GB/s362 TFLOPSBoth LLM inference and training
A100 40G/80G40/80 GB1.5–2 TB/s312 TFLOPSLLMs, tensor parallelism, training
H10080 GB3.35 TB/s989 TFLOPS (with sparsity)Flagship inference and training
RTX 4090 (consumer)24 GB1 TB/s330 TFLOPSDevelopment, internal tools, small services (no ECC/no vGPU)

Consumer vs Professional Cards

The RTX 4090 is exceptional value (nearly twice the performance of an A10), but it has no ECC memory, no vGPU virtualization support, and no datacenter-grade cooling or stability certification, and its license terms prohibit datacenter deployment. Fine for experiments and development—for production, buy professional cards.

Rules of thumb (models < 20B, online inference):

text
FP16 model fits on one card (≤ 24GB) and QPS is modest     → L4 / A10
FP16 model fits on one card but QPS is high (>200)         → L40S / A100
Model doesn't fit on one card (≥ 70B)                      → multi-GPU tensor parallelism (H100/A100×N)

4. CPU Inference: The Underrated Second Front ​

CPU inference has come back into fashion in the LLM era (llama.cpp runs 7B quantized models at 5–10 tokens/s on an ordinary laptop). Its optimization levers:

  • Instruction sets: AVX2 → AVX-512 → AMX (Intel)—up to a 3–5× difference;
  • Libraries: oneDNN (formerly MKL-DNN), OpenBLAS, and llama.cpp's own ggml;
  • Cache friendliness: repack weights into CPU-friendly memory layouts to avoid cache misses.

When to pick CPU: low QPS + small models + budget constraints, plus edge environments with no GPU. The decision rule is in Section 4 of Inference: From Forward Pass to Inference Engines.

5. NPUs, TPUs, and Custom Silicon ​

ChipExamplesPositioningInference performance
Edge NPUsHuawei Ascend Atlas, RK3588 NPU, Apple Neural EngineOn-device/edge AIExcellent energy efficiency, but narrow operator support
Cloud TPUsGoogle TPU v5eCloud LLM training and inferenceStrong compute, tied to the TensorFlow/JAX ecosystem
Other ASICsAWS Inferentia, Groq LPUDedicated inferenceInferentia offers good value; the ecosystem constraints are real

Selection logic: the allure of NPUs/ASICs is cost; the risk is ecosystem lock-in. Once a model's operators aren't supported, you either rewrite or walk away. Most teams "run GPU as the baseline and pilot NPUs in specific scenarios" rather than migrating everything up front.

6. Multi-GPU Setups: The Three Parallelisms ​

When a model doesn't fit on one card, or one card can't deliver the throughput, go multi-GPU. The three roles (see LLM Inference Optimization and Paper: Parallel and Distributed Inference):

text
Data parallelism:     full model on every card, different batches  ── solves "not enough throughput"
Tensor parallelism:   weights sliced by layer across cards, frequent communication ── solves "doesn't fit on one card"
Pipeline parallelism: layer groups across cards, activations passed between ── solves "doesn't fit + cheaper communication"

Key point: every card you add to tensor parallelism adds communication overhead; 2–8 cards is usually the sweet spot; beyond 8, combine with pipeline parallelism. Multi-GPU setups require high-speed interconnects like NVLink/InfiniBand, or communication eats the gains.

7. Cloud Selection Cheat Sheet ​

Rough estimates by "model size × target QPS" (single instance, assuming dynamic batching is active):

Model size (FP16 weights)QPS < 10QPS 10–100QPS 100–1000QPS > 1000
< 1 GB (small CNN/MLP)2–4 core CPU8-core CPU / T4T4×1–2L4×2+
1–7 GB (BERT-class)T4T4 / L4L4×1–2L40S×2+
7–30 GB (7B–13B LLM)L4L4 / L40SL40S / A100A100×2+ (tensor parallel)
> 30 GB (70B+)Not realisticA100/H100×2A100×2–4H100×4+ (tensor + pipeline)

The Table Is Only a Starting Point

Real QPS ceilings depend on the model, batch size, and latency budget—you must calibrate with load tests. Method: see Load Testing and Capacity Planning—first find one instance's max QPS, then derive the instance count from the target QPS, and keep 30% headroom for traffic spikes.

A Complete Worked Example ​

Goal: deploy a 13B model (~26GB in FP16) with P99 latency < 1s and 1 million calls per day (average QPS ≈ 12, peak ≈ 60).

text
1) VRAM:      26GB FP16 weights + activations and KV cache → need ≥ 48GB per card → L40S/A100
2) Bandwidth: at 60 QPS peak, continuous batching yields 1000+ token/s → one card's bandwidth suffices
3) Compute:   LLM generation is bandwidth-bound; compute isn't the bottleneck
4) Verdict:   start with 1× L40S (48GB), scale to 2× after load testing

Trade-offs ​

Decision pointOptionsHow to choose
Capacity first vs speed firstMore VRAM vs more bandwidthSatisfy capacity first, then trade off between bandwidth and compute
GPU vs CPUGPU fast but pricey; CPU cheap but slowHigh QPS or latency-sensitive → GPU; low QPS small models → CPU
Professional vs consumer cardsStability vs priceProfessional for production; consumer is fine for development
Single GPU vs multi-GPUSimple vs capacity/throughputQuantize to save VRAM first; go multi-GPU only if that's not enough
GPU vs NPU/ASICEcosystem vs costGPU as the baseline; pilot ASICs in stable scenarios

One-line summary: VRAM decides whether it runs, bandwidth decides how fast, compute decides how much it can scale—do the math, pick the card, then calibrate with load tests.

Further Reading ​

References ​