Skip to content

Systems Fundamentals Cheat Sheet

At a glance A systems-knowledge cheat sheet every model deployment engineer needs: Linux, containers and Docker, Kubernetes, networking and process models, GPU environments, and performance tools—at your fingertips whenever you're troubleshooting a deployment.

Systems Fundamentals Cheat Sheet ​

How to Use This Page

This page is a troubleshooting cheat card: each section is organized as scenario → command → how to read the output. When you hit problems like "the service won't start", "the GPU won't show up", or "latency is mysteriously high", jump to the relevant section first, then follow the links in the main text to fill in the underlying systems knowledge.

Linux Commands Quick Reference ​

ScenarioCommandHow to Read the Output
Check whether a process exists and what it's usingps aux | grep pythonColumn 3 is CPU, column 4 is RSS (memory); note the PID for a later kill or top -p.
Watch resource usage in real timetop (or htop)Look for processes with high %CPU/%MEM; a load average above the core count means jobs are queuing. top -p PID watches a single process.
Check how much memory is leftfree -havailable is what you can actually use; buff/cache is reclaimable. A surging swap = memory pressure.
Check how much disk is leftdf -hA Use% near 100% causes write failures; pay special attention to model repository mounts.
See who is listening on portsss -tlnp (netstat -tlnp on older systems)Only LISTEN lines mean the service is ready; the PID/Program column confirms whether it's your process.
Find who occupies a specific portlsof -i :8000Pinpoints the process when troubleshooting "port already in use"; then decide whether to kill it or switch ports.
See which files a process has openlsof -p PIDLocates disk space that isn't being freed because "the file was deleted but the process still holds the handle".
Tail a logtail -f /path/to/logFollow errors in real time; grep -i error filters for keywords.
Run a service in the backgroundnohup cmd > log 2>&1 &Keeps running after you disconnect SSH; redirecting logs keeps the terminal free.

Containers and Docker ​

Image build best practices (they directly affect image size, build speed, and security):

  • Layers and caching: every RUN/COPY line in a Dockerfile creates a layer, and changing an early layer invalidates every later cache. Put rarely-changing dependency installs first (e.g., apt-get install, pip install) and frequently-changing source COPY last.
  • Merge instructions: combine multiple RUNs with && to reduce layer count (e.g., RUN apt-get update && apt-get install -y curl).
  • Run as non-root: USER nonroot (or create a dedicated user). Never run production containers as root—it's the number-one finding in image security scans.
  • Multi-stage builds: compile in FROM ... AS builder, then copy only the artifacts into the final FROM ..., leaving the compiler toolchain in the build stage.
  • Pin the base image tag (python:3.11-slim); never use latest.

Common docker run flags:

FlagPurposeExample
-dRun in the backgrounddocker run -d -p 8000:8000 my-svc
-p host:containerPort mapping-p 8000:8000
-v host:containerMount directories (model weights, logs)-v /models:/models
--gpus allPass through all GPUs (requires nvidia-container-toolkit)--gpus '"device=0,1"' selects specific cards
--nameName the container--name triton-server
--restart=unless-stoppedAuto-restart on abnormal exitA must for production services
--shm-sizeIncrease shared memory (needed by multi-process/NCCL)--shm-size 1g
--network hostUse the host network (high-performance/debug scenarios)Watch out for port conflicts

docker compose (docker-compose.yml): declare multi-container orchestration of "services + networks + volumes" in YAML. Common commands: docker compose up -d (start), docker compose logs -f (logs), docker compose down (stop). Standing up an inference service + Prometheus + Grafana together in one compose file is the fastest local simulation environment.

Troubleshooting scenario: no GPU visible inside the container Symptom: nvidia-smi works on the host, but inside the container you get CUDA driver version is insufficient or the device can't be found. Debug chain: nvidia-smi (is the driver OK?) → is nvidia-container-toolkit installed on the host, followed by sudo nvidia-ctk runtime configure --runtime=docker and systemctl restart docker? → verify with docker run --gpus all nvidia/cuda:... nvidia-smi. See the official NVIDIA installation docs.

Kubernetes Core Concepts ​

ObjectOne-Line DefinitionRelation to Inference Services
PodThe smallest schedulable unit; one or more containers share a network/IPThe instance of a model service; usually "one inference process" per Pod
DeploymentDeclaratively manages Pod replica counts (ReplicaSet) and rolling updatesThe standard carrier for stateless inference replicas
ServiceA stable in-cluster access entry point (ClusterIP/DNS)Load balances across Pods; pair with kubectl port-forward for debugging
IngressThe cluster's external HTTP(S) entry routingExposes APIs like /v1/models externally; can attach TLS
HPA (Horizontal Pod Autoscaler)Automatically scales replicas on CPU/custom metricsScale inference replicas on QPS or GPU utilization
ConfigMap / SecretConfiguration and secret injectionModel paths, HF tokens, environment variables
PVCPersistent storage claimMount model weights to avoid re-downloading on every cold start

Resource requests and limits (resources.requests/limits)—the pitfall inference services hit most often:

yaml
resources:
  requests: { cpu: "4", memory: 16Gi, nvidia.com/gpu: 1 }
  limits:   { cpu: "4", memory: 24Gi, nvidia.com/gpu: 1 }
  • requests determine scheduling and QoS (a Pod without requests is BestEffort and gets evicted first); limits keep a process from eating all node memory and triggering OOMKill.
  • Give inference containers limits.cpu equal to requests.cpu (Guaranteed QoS) to avoid preemption, at the cost of reserving those resources.
  • Declare GPUs with nvidia.com/gpu (device plugin); a GPU's requests must equal its limits.

Health probes—they decide whether "traffic gets routed to it or cut from it":

ProbePurposeRecommendation for Inference
livenessProbeRestart the container when the process is deadlockedProbe /healthz; auto-restart on failure
readinessProbeOnly receive traffic once readyShould fail until the model finishes loading, keeping traffic away from replicas where "the model isn't up yet"
startupProbeProtection for slow startupsLarge models take tens of seconds to load; relax liveness via startup to avoid false kills

Troubleshooting scenario: service is up but requests return 503/Connection Refused Start with kubectl get pods for status → kubectl logs to check whether the model finished loading → verify the readinessProbe path and return codes; for repeated CrashLoops, run kubectl describe pod and look for OOMKilled/ImagePullBackOff in Events.

Networking Basics ​

HTTP status code cheatsheet (your first instinct when troubleshooting):

CodeMeaningCommon Triggers
200SuccessA normal inference response
4xxClient error400 bad parameters, 401/403 auth, 404 path, 429 rate limited
5xxServer error500 inference exception, 502 gateway can't reach backend, 503 service not ready, 504 timeout
  • 502/503 in inference: first distinguish between "the load balancer can't reach the Pod" (check Ingress/gateway logs) and "the in-Pod service isn't ready" (check probes).
  • 429 (rate limiting): common when dynamic batching backs up or replicas are insufficient; retry while honoring the Retry-After header.

TCP three-way handshake and timeouts:

  • The handshake (SYN → SYN-ACK → ACK) establishes the connection; capture packets with tcpdump -i any port 8000 to see whether it completes.
  • Common timeout parameters: client connection timeout connect_timeout, server processing timeout read_timeout, overall total_timeout. For inference services, set them separately: 5s for connections, 60s+ for first byte (model inference can be slow), and scale the total timeout to your maximum single-request duration.
  • Connection pools and keep-alive: inference is a high-frequency, small-request workload, so reuse connections (HTTP keep-alive, gRPC channel reuse)—otherwise every request pays the handshake cost. The gap between load tests with ab -k (keep-alive) and without it often is the performance bottleneck.

gRPC: a high-performance RPC built on HTTP/2, suited to throughput-first inference (Triton's gRPC port is the classic example). Characteristics: binary protocol, streaming support, long-lived connection reuse. Rule of thumb: "latency-sensitive + internal systems" → gRPC; "external + easy debugging" → REST.

Processes, Threads, and Async Models ​

ConceptDefinitionImpact on Inference Services
ProcessAn execution unit with isolated resources; each has its own memory spaceMulti-process deployment (e.g., Gunicorn/Uvicorn workers) uses multiple cores in parallel, but each copy loads its own model (VRAM doubles)
ThreadAn execution flow within a process, sharing memoryConstrained by the GIL in Python
GIL (Global Interpreter Lock)CPython allows only one thread to execute bytecode at a timeMulti-threading does not speed up CPU-bound Python code; but inference's heavy computation happens inside C/CUDA libraries, which release the GIL, so multi-threading still works in "Python + inference engine" setups
CoroutineUser-space cooperative scheduling within a single threadFastAPI's async def and the asyncio event-loop model—suited to IO-bound work (network waits, stream forwarding)
Event LoopA single thread managing IO event callbacksFor LLM streaming output, the event loop lets one process serve large numbers of in-flight streams simultaneously

Key takeaway (concurrency model for inference services)

  • CPU-heavy model computation → hand it to the inference engine (C/CUDA code that releases the GIL) and squeeze the hardware with threads/batching.
  • IO-heavy request handling → use asyncio/FastAPI async for concurrency; one worker carries large numbers of long-lived connections.
  • Go multi-process only when you need multi-core parallelism, but first check whether "process count × per-process model VRAM" actually fits.

GPU Environment ​

Reading nvidia-smi output:

+-----------------------------------------------------------------------------+
| NVIDIA-SMI 550.54.15    Driver Version: 550.54.15    CUDA Version: 12.4     |
|-------------------------------+----------------------+----------------------+
| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |
| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |
|===============================+======================+======================|
|   0  A100-SXM4-40GB       On  | 00000000:00:04.0 On  |                    0 |
|  0%   47C    P0    68W / 400W |   3640MiB / 40960MiB |     44%      Default |
+-----------------------------------------------------------------------------+
ColumnMeaningWhat to Watch
Driver Version / CUDA VersionDriver version / the highest CUDA the driver supportsThe driver "supporting" 12.4 ≠ the runtime having CUDA 12.4 installed; check this layer first when containers report libcudart errors
Memory-UsageUsed / total VRAMLow GPU-Util with full memory = the model is stuck loading or there's a leak
GPU-UtilSM utilizationNear 0 while the service is receiving requests = CPU/bandwidth bottleneck or the GPU isn't actually being used
Temp / PwrTemperature / current power drawHitting the power or thermal wall causes sudden performance drops
Ecc ErrorsVRAM ECC error countSustained growth is a hardware failure signal

CUDA / cuDNN version matching (the number-one deployment pitfall):

  • Three layers must match: NVIDIA driver → CUDA runtime → the CUDA version the inference framework (PyTorch/TensorRT, etc.) was compiled with.
  • Drivers are backward compatible: the driver just needs to be ≥ the CUDA version the framework requires. To check: nvidia-smi shows the CUDA version the driver supports; python -c "import torch; print(torch.version.cuda)" shows the version PyTorch was built against.
  • Prefer containers: use official images directly (nvidia/cuda:12.4-runtime-ubuntu22.04, NGC's PyTorch/Triton images)—framework and CUDA versions are already paired, so don't install them by hand.

Drivers and containers (nvidia-container-toolkit): containers need the toolkit installed and the runtime configured for Docker/containerd before they can access GPUs—see the container section link above.

VRAM leak troubleshooting:

ScenarioCommand/MethodHow to Judge
Memory fully consumed and never freedWatch Memory-Usage continuously in nvidia-smiVRAM still climbing while the service is idle = a leak (often an inference engine/framework bug)
Find who is holding VRAMnvidia-smi --query-compute-apps=pid,used_memory --format=csvLook for leftover zombie Python processes and kill them
Per-process usage not visible inside the containernvidia-smi only shows the current containerUse --display or run fuser -v /dev/nvidia* directly on the host to investigate
False positives from PyTorch's cachetorch.cuda.empty_cache()VRAM held by PyTorch's caching pool is normal; only load testing reveals the true peak

Performance Investigation Commands ​

ToolWhat It ShowsHow to Use It for Inference Services
vmstat 1Process/memory/IO/CPU overviewr (run queue) persistently above the core count = CPU queuing; non-zero si/so = memory pressure and swapping
iostat -x 1Disk throughput and utilizationCheck %util and await when model cold starts/weight loading are slow; high await = slow disk
pidstat 1Per-process CPU/memoryPinpoint which process is eating CPU (e.g., tokenizer threads vs. inference threads)
perf top / perf recordHot-function samplingFor Python services, start with py-spy top --pid for the Python stack, then reach for perf for lower-level hotspots
flamegraphCall-stack time visualizationGenerate from perf/py-spy data; see at a glance "which call stack the time went to"
sar -n DEV 1NIC traffic and errorsJudge whether the network has become the throughput bottleneck
nload / iftopReal-time traffic and connectionsInvestigate bandwidth usage from LLM streaming output

Troubleshooting approach: when latency is high, split it into layers Client-perceived time = network round trip + server-side queuing + inference compute + post-processing. Use time curl to isolate the network; use server-side metrics (e.g., Triton/vLLM request-level metrics) to see the "queuing vs. execution" split; use flamegraphs to find hotspots inside execution. Layered localization beats blind guessing at a single layer by a wide margin.

Further Reading ​