Appearance
Systems Fundamentals Cheat Sheet
How to Use This Page
This page is a troubleshooting cheat card: each section is organized as scenario → command → how to read the output. When you hit problems like "the service won't start", "the GPU won't show up", or "latency is mysteriously high", jump to the relevant section first, then follow the links in the main text to fill in the underlying systems knowledge.
Linux Commands Quick Reference
| Scenario | Command | How to Read the Output |
|---|---|---|
| Check whether a process exists and what it's using | ps aux | grep python | Column 3 is CPU, column 4 is RSS (memory); note the PID for a later kill or top -p. |
| Watch resource usage in real time | top (or htop) | Look for processes with high %CPU/%MEM; a load average above the core count means jobs are queuing. top -p PID watches a single process. |
| Check how much memory is left | free -h | available is what you can actually use; buff/cache is reclaimable. A surging swap = memory pressure. |
| Check how much disk is left | df -h | A Use% near 100% causes write failures; pay special attention to model repository mounts. |
| See who is listening on ports | ss -tlnp (netstat -tlnp on older systems) | Only LISTEN lines mean the service is ready; the PID/Program column confirms whether it's your process. |
| Find who occupies a specific port | lsof -i :8000 | Pinpoints the process when troubleshooting "port already in use"; then decide whether to kill it or switch ports. |
| See which files a process has open | lsof -p PID | Locates disk space that isn't being freed because "the file was deleted but the process still holds the handle". |
| Tail a log | tail -f /path/to/log | Follow errors in real time; grep -i error filters for keywords. |
| Run a service in the background | nohup cmd > log 2>&1 & | Keeps running after you disconnect SSH; redirecting logs keeps the terminal free. |
Containers and Docker
Image build best practices (they directly affect image size, build speed, and security):
- Layers and caching: every
RUN/COPYline in aDockerfilecreates a layer, and changing an early layer invalidates every later cache. Put rarely-changing dependency installs first (e.g.,apt-get install,pip install) and frequently-changing sourceCOPYlast. - Merge instructions: combine multiple
RUNs with&&to reduce layer count (e.g.,RUN apt-get update && apt-get install -y curl). - Run as non-root:
USER nonroot(or create a dedicated user). Never run production containers as root—it's the number-one finding in image security scans. - Multi-stage builds: compile in
FROM ... AS builder, then copy only the artifacts into the finalFROM ..., leaving the compiler toolchain in the build stage. - Pin the base image tag (
python:3.11-slim); never uselatest.
Common docker run flags:
| Flag | Purpose | Example |
|---|---|---|
-d | Run in the background | docker run -d -p 8000:8000 my-svc |
-p host:container | Port mapping | -p 8000:8000 |
-v host:container | Mount directories (model weights, logs) | -v /models:/models |
--gpus all | Pass through all GPUs (requires nvidia-container-toolkit) | --gpus '"device=0,1"' selects specific cards |
--name | Name the container | --name triton-server |
--restart=unless-stopped | Auto-restart on abnormal exit | A must for production services |
--shm-size | Increase shared memory (needed by multi-process/NCCL) | --shm-size 1g |
--network host | Use the host network (high-performance/debug scenarios) | Watch out for port conflicts |
docker compose (docker-compose.yml): declare multi-container orchestration of "services + networks + volumes" in YAML. Common commands: docker compose up -d (start), docker compose logs -f (logs), docker compose down (stop). Standing up an inference service + Prometheus + Grafana together in one compose file is the fastest local simulation environment.
Troubleshooting scenario: no GPU visible inside the container Symptom:
nvidia-smiworks on the host, but inside the container you getCUDA driver version is insufficientor the device can't be found. Debug chain:nvidia-smi(is the driver OK?) → isnvidia-container-toolkitinstalled on the host, followed bysudo nvidia-ctk runtime configure --runtime=dockerandsystemctl restart docker? → verify withdocker run --gpus all nvidia/cuda:... nvidia-smi. See the official NVIDIA installation docs.
Kubernetes Core Concepts
| Object | One-Line Definition | Relation to Inference Services |
|---|---|---|
| Pod | The smallest schedulable unit; one or more containers share a network/IP | The instance of a model service; usually "one inference process" per Pod |
| Deployment | Declaratively manages Pod replica counts (ReplicaSet) and rolling updates | The standard carrier for stateless inference replicas |
| Service | A stable in-cluster access entry point (ClusterIP/DNS) | Load balances across Pods; pair with kubectl port-forward for debugging |
| Ingress | The cluster's external HTTP(S) entry routing | Exposes APIs like /v1/models externally; can attach TLS |
| HPA (Horizontal Pod Autoscaler) | Automatically scales replicas on CPU/custom metrics | Scale inference replicas on QPS or GPU utilization |
| ConfigMap / Secret | Configuration and secret injection | Model paths, HF tokens, environment variables |
| PVC | Persistent storage claim | Mount model weights to avoid re-downloading on every cold start |
Resource requests and limits (resources.requests/limits)—the pitfall inference services hit most often:
yaml
resources:
requests: { cpu: "4", memory: 16Gi, nvidia.com/gpu: 1 }
limits: { cpu: "4", memory: 24Gi, nvidia.com/gpu: 1 }requestsdetermine scheduling and QoS (a Pod without requests is BestEffort and gets evicted first);limitskeep a process from eating all node memory and triggering OOMKill.- Give inference containers
limits.cpuequal torequests.cpu(Guaranteed QoS) to avoid preemption, at the cost of reserving those resources. - Declare GPUs with
nvidia.com/gpu(device plugin); a GPU's requests must equal its limits.
Health probes—they decide whether "traffic gets routed to it or cut from it":
| Probe | Purpose | Recommendation for Inference |
|---|---|---|
livenessProbe | Restart the container when the process is deadlocked | Probe /healthz; auto-restart on failure |
readinessProbe | Only receive traffic once ready | Should fail until the model finishes loading, keeping traffic away from replicas where "the model isn't up yet" |
startupProbe | Protection for slow startups | Large models take tens of seconds to load; relax liveness via startup to avoid false kills |
Troubleshooting scenario: service is up but requests return 503/Connection Refused Start with
kubectl get podsfor status →kubectl logsto check whether the model finished loading → verify the readinessProbe path and return codes; for repeated CrashLoops, runkubectl describe podand look for OOMKilled/ImagePullBackOff in Events.
Networking Basics
HTTP status code cheatsheet (your first instinct when troubleshooting):
| Code | Meaning | Common Triggers |
|---|---|---|
| 200 | Success | A normal inference response |
| 4xx | Client error | 400 bad parameters, 401/403 auth, 404 path, 429 rate limited |
| 5xx | Server error | 500 inference exception, 502 gateway can't reach backend, 503 service not ready, 504 timeout |
- 502/503 in inference: first distinguish between "the load balancer can't reach the Pod" (check Ingress/gateway logs) and "the in-Pod service isn't ready" (check probes).
- 429 (rate limiting): common when dynamic batching backs up or replicas are insufficient; retry while honoring the
Retry-Afterheader.
TCP three-way handshake and timeouts:
- The handshake (SYN → SYN-ACK → ACK) establishes the connection; capture packets with
tcpdump -i any port 8000to see whether it completes. - Common timeout parameters: client connection timeout
connect_timeout, server processing timeoutread_timeout, overalltotal_timeout. For inference services, set them separately: 5s for connections, 60s+ for first byte (model inference can be slow), and scale the total timeout to your maximum single-request duration. - Connection pools and keep-alive: inference is a high-frequency, small-request workload, so reuse connections (HTTP keep-alive, gRPC channel reuse)—otherwise every request pays the handshake cost. The gap between load tests with
ab -k(keep-alive) and without it often is the performance bottleneck.
gRPC: a high-performance RPC built on HTTP/2, suited to throughput-first inference (Triton's gRPC port is the classic example). Characteristics: binary protocol, streaming support, long-lived connection reuse. Rule of thumb: "latency-sensitive + internal systems" → gRPC; "external + easy debugging" → REST.
Processes, Threads, and Async Models
| Concept | Definition | Impact on Inference Services |
|---|---|---|
| Process | An execution unit with isolated resources; each has its own memory space | Multi-process deployment (e.g., Gunicorn/Uvicorn workers) uses multiple cores in parallel, but each copy loads its own model (VRAM doubles) |
| Thread | An execution flow within a process, sharing memory | Constrained by the GIL in Python |
| GIL (Global Interpreter Lock) | CPython allows only one thread to execute bytecode at a time | Multi-threading does not speed up CPU-bound Python code; but inference's heavy computation happens inside C/CUDA libraries, which release the GIL, so multi-threading still works in "Python + inference engine" setups |
| Coroutine | User-space cooperative scheduling within a single thread | FastAPI's async def and the asyncio event-loop model—suited to IO-bound work (network waits, stream forwarding) |
| Event Loop | A single thread managing IO event callbacks | For LLM streaming output, the event loop lets one process serve large numbers of in-flight streams simultaneously |
Key takeaway (concurrency model for inference services)
- CPU-heavy model computation → hand it to the inference engine (C/CUDA code that releases the GIL) and squeeze the hardware with threads/batching.
- IO-heavy request handling → use asyncio/FastAPI async for concurrency; one worker carries large numbers of long-lived connections.
- Go multi-process only when you need multi-core parallelism, but first check whether "process count × per-process model VRAM" actually fits.
GPU Environment
Reading nvidia-smi output:
+-----------------------------------------------------------------------------+
| NVIDIA-SMI 550.54.15 Driver Version: 550.54.15 CUDA Version: 12.4 |
|-------------------------------+----------------------+----------------------+
| GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. |
|===============================+======================+======================|
| 0 A100-SXM4-40GB On | 00000000:00:04.0 On | 0 |
| 0% 47C P0 68W / 400W | 3640MiB / 40960MiB | 44% Default |
+-----------------------------------------------------------------------------+| Column | Meaning | What to Watch |
|---|---|---|
Driver Version / CUDA Version | Driver version / the highest CUDA the driver supports | The driver "supporting" 12.4 ≠ the runtime having CUDA 12.4 installed; check this layer first when containers report libcudart errors |
Memory-Usage | Used / total VRAM | Low GPU-Util with full memory = the model is stuck loading or there's a leak |
GPU-Util | SM utilization | Near 0 while the service is receiving requests = CPU/bandwidth bottleneck or the GPU isn't actually being used |
Temp / Pwr | Temperature / current power draw | Hitting the power or thermal wall causes sudden performance drops |
Ecc Errors | VRAM ECC error count | Sustained growth is a hardware failure signal |
CUDA / cuDNN version matching (the number-one deployment pitfall):
- Three layers must match: NVIDIA driver → CUDA runtime → the CUDA version the inference framework (PyTorch/TensorRT, etc.) was compiled with.
- Drivers are backward compatible: the driver just needs to be ≥ the CUDA version the framework requires. To check:
nvidia-smishows the CUDA version the driver supports;python -c "import torch; print(torch.version.cuda)"shows the version PyTorch was built against. - Prefer containers: use official images directly (
nvidia/cuda:12.4-runtime-ubuntu22.04, NGC's PyTorch/Triton images)—framework and CUDA versions are already paired, so don't install them by hand.
Drivers and containers (nvidia-container-toolkit): containers need the toolkit installed and the runtime configured for Docker/containerd before they can access GPUs—see the container section link above.
VRAM leak troubleshooting:
| Scenario | Command/Method | How to Judge |
|---|---|---|
| Memory fully consumed and never freed | Watch Memory-Usage continuously in nvidia-smi | VRAM still climbing while the service is idle = a leak (often an inference engine/framework bug) |
| Find who is holding VRAM | nvidia-smi --query-compute-apps=pid,used_memory --format=csv | Look for leftover zombie Python processes and kill them |
| Per-process usage not visible inside the container | nvidia-smi only shows the current container | Use --display or run fuser -v /dev/nvidia* directly on the host to investigate |
| False positives from PyTorch's cache | torch.cuda.empty_cache() | VRAM held by PyTorch's caching pool is normal; only load testing reveals the true peak |
Performance Investigation Commands
| Tool | What It Shows | How to Use It for Inference Services |
|---|---|---|
vmstat 1 | Process/memory/IO/CPU overview | r (run queue) persistently above the core count = CPU queuing; non-zero si/so = memory pressure and swapping |
iostat -x 1 | Disk throughput and utilization | Check %util and await when model cold starts/weight loading are slow; high await = slow disk |
pidstat 1 | Per-process CPU/memory | Pinpoint which process is eating CPU (e.g., tokenizer threads vs. inference threads) |
perf top / perf record | Hot-function sampling | For Python services, start with py-spy top --pid for the Python stack, then reach for perf for lower-level hotspots |
flamegraph | Call-stack time visualization | Generate from perf/py-spy data; see at a glance "which call stack the time went to" |
sar -n DEV 1 | NIC traffic and errors | Judge whether the network has become the throughput bottleneck |
nload / iftop | Real-time traffic and connections | Investigate bandwidth usage from LLM streaming output |
Troubleshooting approach: when latency is high, split it into layers Client-perceived time = network round trip + server-side queuing + inference compute + post-processing. Use
time curlto isolate the network; use server-side metrics (e.g., Triton/vLLM request-level metrics) to see the "queuing vs. execution" split; use flamegraphs to find hotspots inside execution. Layered localization beats blind guessing at a single layer by a wide margin.