Skip to content

Hardware Primer

At a glance Hardware architecture for inference acceleration — GPU/SM/Tensor Core, a three-generation comparison of NVIDIA datacenter cards (A100/H100/H200/B200) on memory, bandwidth, compute, and power, consumer GPUs, AMD MI300X, Apple Silicon, Ascend 910B, Intel AMX, NPUs, and the training-vs-inference selection logic.

Hardware Primer ​

Every software optimization for inference acceleration ultimately lands on concrete hardware: SM counts, Tensor Core throughput, HBM bandwidth, NVLink interconnects — these numbers set the theoretical ceiling. But hardware is far more than a spec sheet: it has microarchitecture features (Hopper's TMA, Blackwell's FP4 Tensor Cores), interconnect topologies (NVSwitch vs PCIe), and power/thermal constraints. This page is the "hardware reference base" for the site's chapters on GPU Architecture and Optimization, The GPU Memory Hierarchy and the Bandwidth Wall, The Roofline Model and Compute Analysis, and Distributed Inference (TP/PP), cross-validated with Benchmark Data & Tool Profiles.

How to use

This page is not a procurement guide — it's a quick reference of "hardware parameters + microarchitecture intuition." Read Section 1 first to build a mental model of "what a GPU is," then look at the three-generation comparison in Section 2, and finally consult the other chips as needed. Parameters change as new products launch — defer to vendors' current official specs for real selection decisions; the numbers on this page are as of 2026-08.

1. GPU Architecture Basics ​

A GPU does massive parallel computation by packing arithmetic into hardware-level parallel execution units. Understanding inference optimization requires a mental model of GPU microarchitecture first. See GPU Architecture and Optimization.

1.1 SM / warp / SIMT ​

The SM (streaming multiprocessor) is the GPU's basic compute unit. A GPU consists of multiple SMs, each containing: several CUDA cores (scalar FP32/FP64 units), several Tensor Cores (dedicated matrix multiply-accumulate units), shared memory / L1 (fast on-chip SRAM), a register file, and thread schedulers.

The GPU execution model is called SIMT (single instruction, multiple threads): one instruction is issued to multiple threads executing simultaneously — like SIMD, but each thread has its own registers and program counter. Every 32 threads form a warp, the smallest scheduling unit of an SM — threads within one warp must execute in lockstep on the same instruction; when branches diverge (warp divergence), threads that don't take the branch idle.

1.2 Tensor Core and matrix multiplication ​

The Tensor Core is a dedicated matrix multiply-accumulate unit introduced with the Volta architecture (V100, 2017). In one clock cycle, a single Tensor Core completes one D = A×B + C 4×4×4 matmul (FP16 inputs, FP16/FP32 accumulation). A100's Tensor Cores support BF16, FP16, TF32, INT8, and INT4; Hopper added native FP8 (E4M3/E5M2); Blackwell added FP4 and FP6.

Nearly all compute cost of LLM inference sits in GEMM (general matrix multiply): prefill is large-batch GEMM (Q/K/V = X × W, compute-bound); decode is small-batch GEMM (x × W, memory-bound — see The GPU Memory Hierarchy and the Bandwidth Wall). Tensor Core throughput directly sets the prefill ceiling.

1.3 Memory hierarchy and bandwidth ​

A GPU has three levels of storage, with latency and bandwidth degrading level by level:

LevelCapacity (A100 single card)Bandwidth magnitudePurpose
Registers256 KB / SM × 108 = 27 MB~30 TB/sPer-thread private variables — fastest
Shared memory / L1 (SRAM)192 KB / SM~20 TB/sShared within warp/block — tiles inside a kernel
L2 cache40 MB (shared across SMs)~7 TB/sCross-SM coherence, HBM buffer
HBM (external memory)80 GB2.0 TB/sModel weights, KV cache, activations

The decode-phase bottleneck of LLM inference is exactly the step "read weights from HBM into the Tensor Core" — every generated token streams the entire model weight set once. See The Roofline Model and Compute Analysis. This is why HBM bandwidth is one of the most important hardware metrics for LLM inference.

1.4 cuDNN / cuBLAS / NCCL ​

NVIDIA provides a full stack of low-level libraries:

  • cuBLAS: the BLAS matrix-multiplication library — the foundation of GEMM; the GEMM layers under TensorRT/PyTorch both call it;
  • cuDNN: the deep-learning operator library — high-performance convolutions, pooling, normalization, and activations, with kernel auto-tuning built in;
  • NCCL: the multi-GPU collective communication library (AllReduce, AllGather, Broadcast) — the layer under PyTorch DDP/FSDP and TP/PP.

These libraries are the default kit for "using a GPU well" — hand-written CUDA kernels usually can't beat them, so the optimization direction should be "tune them / call their higher-level wrappers," not rewrite them. See TensorRT and GPU Inference.

2. Three Generations of NVIDIA Datacenter GPUs ​

NVIDIA is the de facto mainstream of datacenter AI compute, and its three generations cover nearly every deployment scenario from 2020 to 2026. See Benchmark Data & Tool Profiles for per-generation measured data.

2.1 Microarchitectures and representative models ​

MicroarchitectureYearRepresentative datacenter cardsKey additions
Ampere2020A100 40GB / 80GBTF32, 2:4 sparsity, MIG
Hopper2022H100 SXM / PCIe, H200Native FP8, Transformer Engine, TMA, multi-instance GPU
Blackwell2024B200, GB200 NVL72Native FP4, second-generation Transformer Engine, NVLink Switch 1.8 TB/s

2.2 The big spec table ​

The table below compares the three generations side by side. All numbers are SXM versions (PCIe versions are slightly lower on bandwidth and power); defer to official launch specs.

SpecA100 80GBH100 SXM 80GBH200 141GBB200
MicroarchitectureAmpereHopperHopperBlackwell
SM count108132132148
HBM typeHBM2eHBM3HBM3eHBM3e
HBM capacity80 GB80 GB141 GB192 GB
HBM bandwidth2.0 TB/s3.35 TB/s4.8 TB/s8.0 TB/s
L2 cache40 MB50 MB50 MB60 MB
FP16/BF16 dense compute312 TFLOPS989 TFLOPS989 TFLOPS2250 TFLOPS
FP16/BF16 sparse compute624 TFLOPS1979 TFLOPS1979 TFLOPS4500 TFLOPS
FP8 compute—1979 TFLOPS (sparse)1979 TFLOPS (sparse)9000 TFLOPS (sparse)
FP4 compute———18000 TFLOPS (sparse)
INT8 compute (sparse)1248 TOPS3958 TOPS3958 TOPS~9000 TOPS
NVLink bandwidth600 GB/s900 GB/s900 GB/s1800 GB/s
PCIeGen4 ×16Gen5 ×16Gen5 ×16Gen5 ×16
TDP400 W700 W700 W1000 W
Process node7 nm4 nm (TSMC N4)4 nm4 nm (TSMC N4P)

TFLOPS/TOPS units and sparsity

TFLOPS (trillion floating-point operations per second) is for floating point; TOPS (trillion operations per second) is for integers. The "sparse compute" figures in the table are theoretical peaks with NVIDIA's 2:4 structured sparsity enabled — every 4 elements containing 2 zeros doubles throughput, but only models meeting the sparsity pattern can benefit. LLM inference is generally evaluated on dense compute. See Pruning and Sparsification.

2.3 What changed across the three generations ​

Ampere A100 (2020 — see Benchmark Data & Tool Profiles): the FP16 training workhorse; TF32 was a clever "FP19 format compatible with FP32" that let FP32 training accelerate without code changes. Its HBM2e bandwidth of 2.0 TB/s was already the bottleneck for 70B-LLM decode. MIG partitions the card into 7 isolated instances — handy for multi-tenancy.

Hopper H100/H200 (2022/2024): the hardware Transformer Engine automatically selects FP8/FP16/BF16 precision per layer; the TMA (tensor memory accelerator) hardwires "asynchronously copy tiles from HBM into shared memory," making kernels both simpler and faster; native FP8 is the watershed of LLM quantization — near-FP16 accuracy with doubled compute. H200 upgrades to HBM3e at 4.8 TB/s and 141 GB, purpose-built against the LLM decode bandwidth wall and long-context capacity. See Weight-Only Quantization and Mixed Precision.

Blackwell B200 (2024 — see Benchmark Data & Tool Profiles): native FP4 Tensor Cores double theoretical compute over H100 FP8 again; 192 GB of HBM3e lets a single card hold a full 405B model; the NVLink Switch interconnects 72 B200s into "one giant GPU," making single-machine inference of trillion-parameter models feasible. The GB200 NVL72 rack delivers 720 PFLOPS of FP8 compute.

3. Consumer vs Datacenter GPUs ​

Consumer cards are not "neutered datacenter cards" — they follow a different design philosophy: small memory but respectable single-precision compute, weak interconnect (no NVLink), and prices an order of magnitude lower. Valuable for individual learning and small-scale inference. See llama.cpp and GGUF.

SpecRTX 4090RTX 5090A100 80GBH100 SXM
MicroarchitectureAda LovelaceBlackwell (GB202)AmpereHopper
Memory24 GB GDDR6X32 GB GDDR780 GB HBM2e80 GB HBM3
Memory bandwidth1008 GB/s1792 GB/s2.0 TB/s3.35 TB/s
FP16 dense compute165 TFLOPS (no sparsity)~210 TFLOPS312 TFLOPS989 TFLOPS
FP8 compute———1979 TFLOPS (sparse)
NVLink✗ (PCIe only)✗✓ 600 GB/s✓ 900 GB/s
TDP450 W575 W400 W700 W
Suggested useLLM inference up to 7BLLM inference 13B–30BTraining / 70B LLM inferenceLarge-scale training and inference

Two pitfalls of consumer cards

Memory is the ceiling: an RTX 4090's 24 GB just fits Llama-3-8B in FP16 (16 GB); 70B requires INT4 quantization plus offload to barely run. Weak interconnect: with no NVLink, multi-GPU can only use PCIe, and TP inference hits a severe bandwidth wall — 2×4090 on 70B loses to a single H100. See Distributed Inference (TP/PP).

4. AMD Instinct MI300X ​

AMD's datacenter AI accelerator, benchmarking against H100. MI300X (2023) uses the CDNA 3 architecture with 192 GB of HBM3 at 5.3 TB/s — 2.4× the memory of an H100 80GB per card, fitting a full Llama-3-70B or even a quantized 405B. FP16 dense compute is 1307 TFLOPS and FP8 2615 TFLOPS — both slightly above H100.

The software stack is ROCm (the counterpart of CUDA), with native PyTorch and vLLM support. The shortfall is ecosystem maturity and community material — debugging and kernel-optimization toolchains are still catching up. See Inference Engine Comparison.

5. Apple Silicon ​

The Neural Engine (NPU) of Apple M2/M3 Ultra targets on-device inference, and the unified memory architecture lets "the GPU access main memory directly" without copies — especially friendly to LLM inference: the M3 Ultra's 192 GB of unified memory can run a quantized Llama-3-70B, and llama.cpp supports the Metal backend natively. See llama.cpp and GGUF and Mobile Deployment.

ModelUnified memoryMemory bandwidthNeural EngineMetal acceleration
M2 Ultra192 GB800 GB/s32 TOPS (INT8)✓
M3 Ultra192 GB800 GB/s36 TOPS (INT8)✓

Where a Mac fits for LLMs

A Mac suits personal local inference and on-device prototyping: unified memory makes "large memory" nearly free (M3 Ultra 192GB ≈ the price of one H100 with 2.4× the memory), but compute is far below datacenter cards and throughput differs by orders of magnitude. Batch serving still needs GPUs. See Mobile Deployment.

6. Huawei Ascend 910B ​

The representative of domestic Chinese AI accelerators — Da Vinci architecture, CANN software stack (the counterpart of CUDA). The 910B delivers ~320 TFLOPS of FP16 compute per card with 64 GB HBM at 1.6 TB/s — benchmarking against A100 with slightly less memory. The companion MindIE inference engine and MindSpore training framework are the compliance-oriented alternatives in China; PyTorch adapts via torch_npu. The shortfall is ecosystem and open-source toolchain maturity — community material still trails CUDA. See Inference Engine Comparison.

7. Intel CPU and AMX ​

AMX (Advanced Matrix Extensions), introduced with Intel Sapphire Rapids (2023), adds dedicated matrix-multiply units to the CPU — essentially Tile registers plus TMUL (tile matrix multiply) instructions inside x86. INT8/BF16 matmul throughput multiplies versus AVX-512. AMX turns the CPU from "barely runs" into a realistic option for small-model inference: BERT-Base INT8 reaches ~1000 tokens/s on a 60-core Sapphire Rapids. See OpenVINO and CPU Inference.

CPU inference's advantage is not throughput but ease of deployment — every server has a CPU: no GPU to install, no PCIe, no drivers. Common for low-latency small-batch scenarios or edge-node inference in distributed systems. Pair it with OpenVINO and ONNX Runtime's graph optimizations for further gains.

8. NPUs and Edge AI Chips ​

NPU (neural processing unit) is a generic term for ASICs dedicated to neural-network inference — broadly including TPUs, Ascend, Apple Neural Engine, Qualcomm Hexagon, and more. Representative edge chips:

ChipVendorComputeUse case
Hexagon NPUQualcomm45 TOPS from Snapdragon 8 Gen3On-phone LLMs, Stable Diffusion
APU (fusion)MediaTekINT8 compute from Dimensity 9300On-phone generative AI
Edge TPUGoogle4 TOPS (INT8)Coral devices, lightweight classification/detection
Jetson OrinNVIDIA275 TOPS (sparse INT8)Robotics, autonomous-driving edge
Rockchip RK3588Rockchip6 TOPS (NPU)Edge gateways, lightweight vision

The core constraints of edge inference are power (phones at 5W, IoT at 1W scale) and memory (phones at 8–16 GB shared with the OS). LLM on-device deployment depends on 1–4B small models + INT4 quantization + inference-engine optimization pushed to the extreme. See Mobile Deployment.

9. Quick-Reference Spec Table ​

Key specs of the hardware above, consolidated for horizontal comparison. Compute and bandwidth follow official launch specifications.

HardwareTypeMemoryBandwidthFP16 computeFP8 computeTDP
A100 80GB SXMGPU80 GB HBM2e2.0 TB/s312 TFLOPS—400 W
H100 SXM 80GBGPU80 GB HBM33.35 TB/s989 TFLOPS1979 TFLOPS (sparse)700 W
H200 141GBGPU141 GB HBM3e4.8 TB/s989 TFLOPS1979 TFLOPS (sparse)700 W
B200GPU192 GB HBM3e8.0 TB/s2250 TFLOPS9000 TFLOPS (sparse)1000 W
RTX 4090GPU24 GB GDDR6X1008 GB/s165 TFLOPS—450 W
RTX 5090GPU32 GB GDDR71792 GB/s~210 TFLOPS—575 W
MI300XGPU192 GB HBM35.3 TB/s1307 TFLOPS2615 TFLOPS (sparse)750 W
M3 UltraSoC+NPU192 GB unified800 GB/s27 TFLOPS (GPU)—~370 W (whole machine)
Ascend 910BNPU64 GB HBM1.6 TB/s~320 TFLOPS—310 W
Sapphire Rapids XeonCPUDDR5 8-channel~460 GB/sAMX BF16 ~50 TFLOPS—350 W
Snapdragon 8 Gen3SoC+NPULPDDR5X shared~77 GB/s——~5 W (NPU)

10. Selection Logic ​

Hardware selection concerns completely different things per scenario — first separate training vs inference, and datacenter vs workstation vs edge.

10.1 Training vs Inference ​

DimensionTrainingInference
Key bottleneckCompute + interconnect (TP/PP multi-GPU)Bandwidth + memory capacity (LLM decode)
Memory needsActivations + optimizer states — wants capacityModel weights + KV cache
PrecisionFP16/BF16 mostly, some FP32INT8/INT4/FP8 mostly, FP16 as fallback
InterconnectTP requires NVLinkTP requires NVLink; single card needs none
Card selectionH100/H200/B200 — high compute, strong interconnectDepends on batch: large batch wants compute, small batch wants bandwidth

The special trait of LLM inference is that decode is memory-intensive — single-card throughput ≈ bandwidth / model bytes. So HBM bandwidth is one of the most important card-selection metrics for LLM inference, arguably more important than compute. See The Roofline Model and Compute Analysis and The GPU Memory Hierarchy and the Bandwidth Wall.

10.2 Datacenter vs Workstation vs Edge ​

Datacenter: A100/H100/H200/B200 + NVLink + NVSwitch. Throughput and multi-GPU interconnect are the lifeblood; the 8-GPU-per-node NVSwitch topology is standard equipment for 70B+ TP inference; B200 NVL72 wires 72 GPUs into "one giant GPU," making single-machine inference of trillion-parameter models feasible. Power and cooling (liquid cooling) are the deployment challenge. See Distributed Inference (TP/PP).

Workstation: RTX 4090/5090 or Apple M Ultra. Small memory but cheap; no NVLink but PCIe is enough for single-card work. Suits personal learning, prototype validation, and small-batch serving. See llama.cpp and GGUF.

Edge: phone NPUs, Jetson, Edge TPU. Small compute, small memory, extremely power-sensitive. INT4 quantization + small models are mandatory, paired with llama.cpp or ONNX Runtime Mobile. See Mobile Deployment.

Three rules of thumb for selection

  1. Fix the batch size before picking the card: single-user chat at batch≈1 wants bandwidth; offline batch processing at large batch wants compute;
  2. Memory capacity is a hard wall: model + KV cache must fit before anything else — if it doesn't fit, quantize or use TP;
  3. Interconnect sets the multi-GPU ceiling: NVLink is a hard requirement for TP inference; PCIe multi-GPU only suits independent services that don't communicate.

Further Reading ​