Appearance
Hardware Primer
Every software optimization for inference acceleration ultimately lands on concrete hardware: SM counts, Tensor Core throughput, HBM bandwidth, NVLink interconnects — these numbers set the theoretical ceiling. But hardware is far more than a spec sheet: it has microarchitecture features (Hopper's TMA, Blackwell's FP4 Tensor Cores), interconnect topologies (NVSwitch vs PCIe), and power/thermal constraints. This page is the "hardware reference base" for the site's chapters on GPU Architecture and Optimization, The GPU Memory Hierarchy and the Bandwidth Wall, The Roofline Model and Compute Analysis, and Distributed Inference (TP/PP), cross-validated with Benchmark Data & Tool Profiles.
How to use
This page is not a procurement guide — it's a quick reference of "hardware parameters + microarchitecture intuition." Read Section 1 first to build a mental model of "what a GPU is," then look at the three-generation comparison in Section 2, and finally consult the other chips as needed. Parameters change as new products launch — defer to vendors' current official specs for real selection decisions; the numbers on this page are as of 2026-08.
1. GPU Architecture Basics
A GPU does massive parallel computation by packing arithmetic into hardware-level parallel execution units. Understanding inference optimization requires a mental model of GPU microarchitecture first. See GPU Architecture and Optimization.
1.1 SM / warp / SIMT
The SM (streaming multiprocessor) is the GPU's basic compute unit. A GPU consists of multiple SMs, each containing: several CUDA cores (scalar FP32/FP64 units), several Tensor Cores (dedicated matrix multiply-accumulate units), shared memory / L1 (fast on-chip SRAM), a register file, and thread schedulers.
The GPU execution model is called SIMT (single instruction, multiple threads): one instruction is issued to multiple threads executing simultaneously — like SIMD, but each thread has its own registers and program counter. Every 32 threads form a warp, the smallest scheduling unit of an SM — threads within one warp must execute in lockstep on the same instruction; when branches diverge (warp divergence), threads that don't take the branch idle.
1.2 Tensor Core and matrix multiplication
The Tensor Core is a dedicated matrix multiply-accumulate unit introduced with the Volta architecture (V100, 2017). In one clock cycle, a single Tensor Core completes one D = A×B + C 4×4×4 matmul (FP16 inputs, FP16/FP32 accumulation). A100's Tensor Cores support BF16, FP16, TF32, INT8, and INT4; Hopper added native FP8 (E4M3/E5M2); Blackwell added FP4 and FP6.
Nearly all compute cost of LLM inference sits in GEMM (general matrix multiply): prefill is large-batch GEMM (Q/K/V = X × W, compute-bound); decode is small-batch GEMM (x × W, memory-bound — see The GPU Memory Hierarchy and the Bandwidth Wall). Tensor Core throughput directly sets the prefill ceiling.
1.3 Memory hierarchy and bandwidth
A GPU has three levels of storage, with latency and bandwidth degrading level by level:
| Level | Capacity (A100 single card) | Bandwidth magnitude | Purpose |
|---|---|---|---|
| Registers | 256 KB / SM × 108 = 27 MB | ~30 TB/s | Per-thread private variables — fastest |
| Shared memory / L1 (SRAM) | 192 KB / SM | ~20 TB/s | Shared within warp/block — tiles inside a kernel |
| L2 cache | 40 MB (shared across SMs) | ~7 TB/s | Cross-SM coherence, HBM buffer |
| HBM (external memory) | 80 GB | 2.0 TB/s | Model weights, KV cache, activations |
The decode-phase bottleneck of LLM inference is exactly the step "read weights from HBM into the Tensor Core" — every generated token streams the entire model weight set once. See The Roofline Model and Compute Analysis. This is why HBM bandwidth is one of the most important hardware metrics for LLM inference.
1.4 cuDNN / cuBLAS / NCCL
NVIDIA provides a full stack of low-level libraries:
- cuBLAS: the BLAS matrix-multiplication library — the foundation of GEMM; the GEMM layers under TensorRT/PyTorch both call it;
- cuDNN: the deep-learning operator library — high-performance convolutions, pooling, normalization, and activations, with kernel auto-tuning built in;
- NCCL: the multi-GPU collective communication library (AllReduce, AllGather, Broadcast) — the layer under PyTorch DDP/FSDP and TP/PP.
These libraries are the default kit for "using a GPU well" — hand-written CUDA kernels usually can't beat them, so the optimization direction should be "tune them / call their higher-level wrappers," not rewrite them. See TensorRT and GPU Inference.
2. Three Generations of NVIDIA Datacenter GPUs
NVIDIA is the de facto mainstream of datacenter AI compute, and its three generations cover nearly every deployment scenario from 2020 to 2026. See Benchmark Data & Tool Profiles for per-generation measured data.
2.1 Microarchitectures and representative models
| Microarchitecture | Year | Representative datacenter cards | Key additions |
|---|---|---|---|
| Ampere | 2020 | A100 40GB / 80GB | TF32, 2:4 sparsity, MIG |
| Hopper | 2022 | H100 SXM / PCIe, H200 | Native FP8, Transformer Engine, TMA, multi-instance GPU |
| Blackwell | 2024 | B200, GB200 NVL72 | Native FP4, second-generation Transformer Engine, NVLink Switch 1.8 TB/s |
2.2 The big spec table
The table below compares the three generations side by side. All numbers are SXM versions (PCIe versions are slightly lower on bandwidth and power); defer to official launch specs.
| Spec | A100 80GB | H100 SXM 80GB | H200 141GB | B200 |
|---|---|---|---|---|
| Microarchitecture | Ampere | Hopper | Hopper | Blackwell |
| SM count | 108 | 132 | 132 | 148 |
| HBM type | HBM2e | HBM3 | HBM3e | HBM3e |
| HBM capacity | 80 GB | 80 GB | 141 GB | 192 GB |
| HBM bandwidth | 2.0 TB/s | 3.35 TB/s | 4.8 TB/s | 8.0 TB/s |
| L2 cache | 40 MB | 50 MB | 50 MB | 60 MB |
| FP16/BF16 dense compute | 312 TFLOPS | 989 TFLOPS | 989 TFLOPS | 2250 TFLOPS |
| FP16/BF16 sparse compute | 624 TFLOPS | 1979 TFLOPS | 1979 TFLOPS | 4500 TFLOPS |
| FP8 compute | — | 1979 TFLOPS (sparse) | 1979 TFLOPS (sparse) | 9000 TFLOPS (sparse) |
| FP4 compute | — | — | — | 18000 TFLOPS (sparse) |
| INT8 compute (sparse) | 1248 TOPS | 3958 TOPS | 3958 TOPS | ~9000 TOPS |
| NVLink bandwidth | 600 GB/s | 900 GB/s | 900 GB/s | 1800 GB/s |
| PCIe | Gen4 ×16 | Gen5 ×16 | Gen5 ×16 | Gen5 ×16 |
| TDP | 400 W | 700 W | 700 W | 1000 W |
| Process node | 7 nm | 4 nm (TSMC N4) | 4 nm | 4 nm (TSMC N4P) |
TFLOPS/TOPS units and sparsity
TFLOPS (trillion floating-point operations per second) is for floating point; TOPS (trillion operations per second) is for integers. The "sparse compute" figures in the table are theoretical peaks with NVIDIA's 2:4 structured sparsity enabled — every 4 elements containing 2 zeros doubles throughput, but only models meeting the sparsity pattern can benefit. LLM inference is generally evaluated on dense compute. See Pruning and Sparsification.
2.3 What changed across the three generations
Ampere A100 (2020 — see Benchmark Data & Tool Profiles): the FP16 training workhorse; TF32 was a clever "FP19 format compatible with FP32" that let FP32 training accelerate without code changes. Its HBM2e bandwidth of 2.0 TB/s was already the bottleneck for 70B-LLM decode. MIG partitions the card into 7 isolated instances — handy for multi-tenancy.
Hopper H100/H200 (2022/2024): the hardware Transformer Engine automatically selects FP8/FP16/BF16 precision per layer; the TMA (tensor memory accelerator) hardwires "asynchronously copy tiles from HBM into shared memory," making kernels both simpler and faster; native FP8 is the watershed of LLM quantization — near-FP16 accuracy with doubled compute. H200 upgrades to HBM3e at 4.8 TB/s and 141 GB, purpose-built against the LLM decode bandwidth wall and long-context capacity. See Weight-Only Quantization and Mixed Precision.
Blackwell B200 (2024 — see Benchmark Data & Tool Profiles): native FP4 Tensor Cores double theoretical compute over H100 FP8 again; 192 GB of HBM3e lets a single card hold a full 405B model; the NVLink Switch interconnects 72 B200s into "one giant GPU," making single-machine inference of trillion-parameter models feasible. The GB200 NVL72 rack delivers 720 PFLOPS of FP8 compute.
3. Consumer vs Datacenter GPUs
Consumer cards are not "neutered datacenter cards" — they follow a different design philosophy: small memory but respectable single-precision compute, weak interconnect (no NVLink), and prices an order of magnitude lower. Valuable for individual learning and small-scale inference. See llama.cpp and GGUF.
| Spec | RTX 4090 | RTX 5090 | A100 80GB | H100 SXM |
|---|---|---|---|---|
| Microarchitecture | Ada Lovelace | Blackwell (GB202) | Ampere | Hopper |
| Memory | 24 GB GDDR6X | 32 GB GDDR7 | 80 GB HBM2e | 80 GB HBM3 |
| Memory bandwidth | 1008 GB/s | 1792 GB/s | 2.0 TB/s | 3.35 TB/s |
| FP16 dense compute | 165 TFLOPS (no sparsity) | ~210 TFLOPS | 312 TFLOPS | 989 TFLOPS |
| FP8 compute | — | — | — | 1979 TFLOPS (sparse) |
| NVLink | ✗ (PCIe only) | ✗ | ✓ 600 GB/s | ✓ 900 GB/s |
| TDP | 450 W | 575 W | 400 W | 700 W |
| Suggested use | LLM inference up to 7B | LLM inference 13B–30B | Training / 70B LLM inference | Large-scale training and inference |
Two pitfalls of consumer cards
Memory is the ceiling: an RTX 4090's 24 GB just fits Llama-3-8B in FP16 (16 GB); 70B requires INT4 quantization plus offload to barely run. Weak interconnect: with no NVLink, multi-GPU can only use PCIe, and TP inference hits a severe bandwidth wall — 2×4090 on 70B loses to a single H100. See Distributed Inference (TP/PP).
4. AMD Instinct MI300X
AMD's datacenter AI accelerator, benchmarking against H100. MI300X (2023) uses the CDNA 3 architecture with 192 GB of HBM3 at 5.3 TB/s — 2.4× the memory of an H100 80GB per card, fitting a full Llama-3-70B or even a quantized 405B. FP16 dense compute is 1307 TFLOPS and FP8 2615 TFLOPS — both slightly above H100.
The software stack is ROCm (the counterpart of CUDA), with native PyTorch and vLLM support. The shortfall is ecosystem maturity and community material — debugging and kernel-optimization toolchains are still catching up. See Inference Engine Comparison.
5. Apple Silicon
The Neural Engine (NPU) of Apple M2/M3 Ultra targets on-device inference, and the unified memory architecture lets "the GPU access main memory directly" without copies — especially friendly to LLM inference: the M3 Ultra's 192 GB of unified memory can run a quantized Llama-3-70B, and llama.cpp supports the Metal backend natively. See llama.cpp and GGUF and Mobile Deployment.
| Model | Unified memory | Memory bandwidth | Neural Engine | Metal acceleration |
|---|---|---|---|---|
| M2 Ultra | 192 GB | 800 GB/s | 32 TOPS (INT8) | ✓ |
| M3 Ultra | 192 GB | 800 GB/s | 36 TOPS (INT8) | ✓ |
Where a Mac fits for LLMs
A Mac suits personal local inference and on-device prototyping: unified memory makes "large memory" nearly free (M3 Ultra 192GB ≈ the price of one H100 with 2.4× the memory), but compute is far below datacenter cards and throughput differs by orders of magnitude. Batch serving still needs GPUs. See Mobile Deployment.
6. Huawei Ascend 910B
The representative of domestic Chinese AI accelerators — Da Vinci architecture, CANN software stack (the counterpart of CUDA). The 910B delivers ~320 TFLOPS of FP16 compute per card with 64 GB HBM at 1.6 TB/s — benchmarking against A100 with slightly less memory. The companion MindIE inference engine and MindSpore training framework are the compliance-oriented alternatives in China; PyTorch adapts via torch_npu. The shortfall is ecosystem and open-source toolchain maturity — community material still trails CUDA. See Inference Engine Comparison.
7. Intel CPU and AMX
AMX (Advanced Matrix Extensions), introduced with Intel Sapphire Rapids (2023), adds dedicated matrix-multiply units to the CPU — essentially Tile registers plus TMUL (tile matrix multiply) instructions inside x86. INT8/BF16 matmul throughput multiplies versus AVX-512. AMX turns the CPU from "barely runs" into a realistic option for small-model inference: BERT-Base INT8 reaches ~1000 tokens/s on a 60-core Sapphire Rapids. See OpenVINO and CPU Inference.
CPU inference's advantage is not throughput but ease of deployment — every server has a CPU: no GPU to install, no PCIe, no drivers. Common for low-latency small-batch scenarios or edge-node inference in distributed systems. Pair it with OpenVINO and ONNX Runtime's graph optimizations for further gains.
8. NPUs and Edge AI Chips
NPU (neural processing unit) is a generic term for ASICs dedicated to neural-network inference — broadly including TPUs, Ascend, Apple Neural Engine, Qualcomm Hexagon, and more. Representative edge chips:
| Chip | Vendor | Compute | Use case |
|---|---|---|---|
| Hexagon NPU | Qualcomm | 45 TOPS from Snapdragon 8 Gen3 | On-phone LLMs, Stable Diffusion |
| APU (fusion) | MediaTek | INT8 compute from Dimensity 9300 | On-phone generative AI |
| Edge TPU | 4 TOPS (INT8) | Coral devices, lightweight classification/detection | |
| Jetson Orin | NVIDIA | 275 TOPS (sparse INT8) | Robotics, autonomous-driving edge |
| Rockchip RK3588 | Rockchip | 6 TOPS (NPU) | Edge gateways, lightweight vision |
The core constraints of edge inference are power (phones at 5W, IoT at 1W scale) and memory (phones at 8–16 GB shared with the OS). LLM on-device deployment depends on 1–4B small models + INT4 quantization + inference-engine optimization pushed to the extreme. See Mobile Deployment.
9. Quick-Reference Spec Table
Key specs of the hardware above, consolidated for horizontal comparison. Compute and bandwidth follow official launch specifications.
| Hardware | Type | Memory | Bandwidth | FP16 compute | FP8 compute | TDP |
|---|---|---|---|---|---|---|
| A100 80GB SXM | GPU | 80 GB HBM2e | 2.0 TB/s | 312 TFLOPS | — | 400 W |
| H100 SXM 80GB | GPU | 80 GB HBM3 | 3.35 TB/s | 989 TFLOPS | 1979 TFLOPS (sparse) | 700 W |
| H200 141GB | GPU | 141 GB HBM3e | 4.8 TB/s | 989 TFLOPS | 1979 TFLOPS (sparse) | 700 W |
| B200 | GPU | 192 GB HBM3e | 8.0 TB/s | 2250 TFLOPS | 9000 TFLOPS (sparse) | 1000 W |
| RTX 4090 | GPU | 24 GB GDDR6X | 1008 GB/s | 165 TFLOPS | — | 450 W |
| RTX 5090 | GPU | 32 GB GDDR7 | 1792 GB/s | ~210 TFLOPS | — | 575 W |
| MI300X | GPU | 192 GB HBM3 | 5.3 TB/s | 1307 TFLOPS | 2615 TFLOPS (sparse) | 750 W |
| M3 Ultra | SoC+NPU | 192 GB unified | 800 GB/s | 27 TFLOPS (GPU) | — | ~370 W (whole machine) |
| Ascend 910B | NPU | 64 GB HBM | 1.6 TB/s | ~320 TFLOPS | — | 310 W |
| Sapphire Rapids Xeon | CPU | DDR5 8-channel | ~460 GB/s | AMX BF16 ~50 TFLOPS | — | 350 W |
| Snapdragon 8 Gen3 | SoC+NPU | LPDDR5X shared | ~77 GB/s | — | — | ~5 W (NPU) |
10. Selection Logic
Hardware selection concerns completely different things per scenario — first separate training vs inference, and datacenter vs workstation vs edge.
10.1 Training vs Inference
| Dimension | Training | Inference |
|---|---|---|
| Key bottleneck | Compute + interconnect (TP/PP multi-GPU) | Bandwidth + memory capacity (LLM decode) |
| Memory needs | Activations + optimizer states — wants capacity | Model weights + KV cache |
| Precision | FP16/BF16 mostly, some FP32 | INT8/INT4/FP8 mostly, FP16 as fallback |
| Interconnect | TP requires NVLink | TP requires NVLink; single card needs none |
| Card selection | H100/H200/B200 — high compute, strong interconnect | Depends on batch: large batch wants compute, small batch wants bandwidth |
The special trait of LLM inference is that decode is memory-intensive — single-card throughput ≈ bandwidth / model bytes. So HBM bandwidth is one of the most important card-selection metrics for LLM inference, arguably more important than compute. See The Roofline Model and Compute Analysis and The GPU Memory Hierarchy and the Bandwidth Wall.
10.2 Datacenter vs Workstation vs Edge
Datacenter: A100/H100/H200/B200 + NVLink + NVSwitch. Throughput and multi-GPU interconnect are the lifeblood; the 8-GPU-per-node NVSwitch topology is standard equipment for 70B+ TP inference; B200 NVL72 wires 72 GPUs into "one giant GPU," making single-machine inference of trillion-parameter models feasible. Power and cooling (liquid cooling) are the deployment challenge. See Distributed Inference (TP/PP).
Workstation: RTX 4090/5090 or Apple M Ultra. Small memory but cheap; no NVLink but PCIe is enough for single-card work. Suits personal learning, prototype validation, and small-batch serving. See llama.cpp and GGUF.
Edge: phone NPUs, Jetson, Edge TPU. Small compute, small memory, extremely power-sensitive. INT4 quantization + small models are mandatory, paired with llama.cpp or ONNX Runtime Mobile. See Mobile Deployment.
Three rules of thumb for selection
- Fix the batch size before picking the card: single-user chat at batch≈1 wants bandwidth; offline batch processing at large batch wants compute;
- Memory capacity is a hard wall: model + KV cache must fit before anything else — if it doesn't fit, quantize or use TP;
- Interconnect sets the multi-GPU ceiling: NVLink is a hard requirement for TP inference; PCIe multi-GPU only suits independent services that don't communicate.
Further Reading
- GPU Architecture and Optimization — the optimization principles behind the hardware
- The GPU Memory Hierarchy and the Bandwidth Wall — detailed analysis of HBM/SRAM/L2
- The Roofline Model and Compute Analysis — one model unifying compute and bandwidth
- Distributed Inference (TP/PP) — multi-GPU interconnect and parallelism strategies
- Weight-Only Quantization and Mixed Precision — hardware support for FP8/INT8 quantization
- Benchmark Data & Tool Profiles — measured throughput data per hardware
- Glossary — precise definitions of the terms used here
- llama.cpp and GGUF — consumer-card/edge deployment practice
- Mobile Deployment — NPU and edge inference