Appearance
Putting observability into practice means turning "how is the service actually doing" into an engineering capability you can answer at any moment: the three pillars — metrics, logs, and traces — all in place, with alerts that are precise and not noisy. Why bother? An inference service without observability is like driving without a dashboard — failures such as model drift, latency degradation, and GPU memory leaks never crash loudly; they just let users feel "it's slower and dumber" before you do. This guide walks through every step from zero to production-grade observability, with working code and configs: metric instrumentation, scraping, GPU collection, Grafana dashboards, SLO burn-rate alerts, structured logging, and distributed tracing — ending with a 30-day post-launch observation checklist. For the theory and SLO definitions, see Monitoring and Observability.
The Goal in One Sentence
No missing metrics, searchable logs, quiet alerts — the four golden signals (RPS, latency, error rate, saturation) all covered, and alerts fire only when "a human needs to act".
1. Instrumenting Metrics (Prometheus Client)
Use prometheus-client to instrument the three golden metrics for an inference service: request counters, latency histograms, and error counters. This is the full version of the metrics module from Deploying a Model from Scratch.
python
# app/metrics.py — all metric definitions in one file, referenced across the project
from prometheus_client import Counter, Histogram, Gauge
# Request counter: labels distinguish model and version for easy per-version comparison
REQUESTS = Counter("predict_requests_total", "Total inference requests",
["model", "version"])
ERRORS = Counter("predict_errors_total", "Failed inference requests",
["model", "version", "error_type"])
# Latency histogram: buckets sized to the service's latency range, covering the P99 target
LATENCY = Histogram("predict_latency_seconds", "Inference latency",
["model", "version"],
buckets=(0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1.0, 2.5))
# Queue depth: saturation signal; a Gauge sampled every 5 seconds is enough
QUEUE = Gauge("predict_queue_depth", "Requests currently queued")Three rules for label design (labels are a high-risk design point — get them wrong and everything downstream suffers):
- Cardinality must stay low: label value combinations × time series = storage explosion. High-cardinality values like
user_idorrequest_idmust never become labels — they belong in logs only; - Only include dimensions you will actually query by:
model,version,error_type,regionare reasonable; temporary debugging fields stay out of production metrics; - Keep label sets globally consistent: all metrics share the same
model/versionlabels so alerts and dashboards can join across metrics.
Instrument at the service's entrance and exit (see main.py in Deploying a Model from Scratch); use a middleware or decorator to handle it uniformly instead of sprinkling inc() calls through business code:
python
# app/middleware.py — unified instrumentation
import time
from starlette.middleware.base import BaseHTTPMiddleware
from app.metrics import REQUESTS, ERRORS, LATENCY
class MetricsMiddleware(BaseHTTPMiddleware):
async def dispatch(self, request, call_next):
t0 = time.perf_counter()
try:
resp = await call_next(request)
except Exception:
ERRORS.labels(model="mobilenetv2", version="v1",
error_type="exception").inc()
raise
LATENCY.labels(model="mobilenetv2", version="v1").observe(
time.perf_counter() - t0)
REQUESTS.labels(model="mobilenetv2", version="v1").inc()
return resppython
# app/main.py — expose /metrics for Prometheus to scrape
from prometheus_client import generate_latest, CONTENT_TYPE_LATEST
from fastapi import Response
@app.get("/metrics")
def metrics():
return Response(generate_latest(), media_type=CONTENT_TYPE_LATEST)2. Service Discovery and Scraping
2.1 Simple setup: static_configs
yaml
# prometheus.yml
global:
scrape_interval: 15s # scrape interval; drop to 10s for latency-sensitive services
evaluation_interval: 30s # alert rule evaluation interval
scrape_configs:
- job_name: inference-api
static_configs:
- targets: ["api:8000"] # use the service name inside compose/Docker networks
labels: { env: prod }
- job_name: dcgm # GPU metrics, see Step 3
static_configs:
- targets: ["dcgm-exporter:9400"]2.2 Kubernetes: ServiceMonitor
In K8s, service instances come and go, so use declarative discovery with a ServiceMonitor (paired with Prometheus Operator / kube-prometheus-stack). Once the service carries the prometheus.io/scrape: "true" annotation it is scraped automatically:
yaml
# ServiceMonitor.yaml
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: inference-api
spec:
selector:
matchLabels:
app: inference-api # matches the Service's selector
endpoints:
- port: http # the port name declared on the Service
path: /metrics
interval: 15s
namespaceSelector:
any: trueFailed Scrapes Are the Sneakiest Monitoring Incident
When Prometheus cannot reach a target it usually doesn't error out — it just drops data silently: dashboards slowly empty out and alerts quietly stop working. For the first week after launch, check the up metric daily (up{job="inference-api"} == 1) — it is the only evidence that "the monitoring itself is alive".
3. Collecting GPU Metrics (dcgm-exporter)
When the inference service runs on GPUs, GPU utilization and memory watermark are the core signals for capacity and fault analysis (paired with the "GPU pegged vs CPU pegged" analysis in Load Testing and Capacity Planning). The standard approach is the NVIDIA DCGM exporter:
bash
# Start the DCGM exporter, exposing :9400/metrics (GPU utilization, memory, temperature, power)
docker run -d --gpus all --rm --cap-add SYS_ADMIN \
--name dcgm-exporter \
nvcr.io/nvidia/dcgm-exporter:3.3.5-ubuntu22.04Key metrics (DCGM's own names):
text
# GPU utilization (%), grouped by device
DCGM_FI_DEV_GPU_UTIL
# GPU memory used (bytes)
DCGM_FI_DEV_FB_USED
# GPU temperature and power
DCGM_FI_DEV_GPU_TEMP
DCGM_FI_DEV_POWER_USAGEWithout containers, nvidia-smi --query-gpu can output similar data, but not in Prometheus format — you would have to bridge it yourself through node_exporter's textfile collector, or use a community exporter such as nvidia_gpu_exporter. Debugging GPU memory problems also ties into pitfall #4 of Memory Leaks and OOM.
4. Grafana Dashboards
After connecting Grafana to the Prometheus data source, build an "Inference Service Overview" dashboard with five panels laid out in order:
| Panel | PromQL | What to Look For |
|---|---|---|
| RPS | sum(rate(predict_requests_total[1m])) by (model) | Traffic trend; watch traffic shifts around releases |
| P50/P99 latency | histogram_quantile(0.99, sum by (le, model) (rate(predict_latency_seconds_bucket[5m]))) | Whether P99 stays within the SLO |
| Error rate | sum(rate(predict_errors_total[5m])) / sum(rate(predict_requests_total[5m])) | Share of 5xx/inference errors |
| GPU utilization | avg(DCGM_FI_DEV_GPU_UTIL) | GPU pegged vs idle |
| Queue depth | predict_queue_depth | Saturation; a growing queue precedes an avalanche |
histogram_quantile's P99 estimate carries statistical error — denser buckets are more accurate, at the cost of storage. Design rule: buckets should span the range from "P50 up to 10× the P99 target"; for a P99 target of 100ms, buckets run from 5ms to 2.5s, growing roughly 2x per step (like the buckets array above).
5. Alert Rules (SLO Burn Rates + Severity Tiers)
5.1 Burn-rate alerts — the key to quiet alerting
"Threshold alerts" (fire when latency > 500ms) scream during transient spikes and stay silent during slow degradation — the root of alert storms. SLO burn rate alerts fire based on how fast the SLO error budget is being consumed: they trigger only when the budget burns fast enough over a given window. See Monitoring and Observability for the underlying theory.
yaml
# prometheus-alerts.yml
groups:
- name: inference-slo
rules:
# 30-day SLO of 99.9% (0.1% errors allowed)
# Burn rate 14.4 (5-minute window): burns 2% of the budget in about 2 hours → page
- alert: HighErrorRatePage
expr: |
sum(rate(predict_errors_total[5m]))
/ sum(rate(predict_requests_total[5m]))
> 0.1 * 14.4
labels:
severity: page
annotations:
summary: "Error burn rate above 14.4; 2% of budget gone within 2 hours"
# Burn rate 1.0 (1-hour window): budget exhausted within 30 days → warn, on-call queue
- alert: ErrorBudgetWarn
expr: |
sum(rate(predict_errors_total[1h]))
/ sum(rate(predict_requests_total[1h]))
> 0.1
labels:
severity: warn
annotations:
summary: "1-hour error rate exceeds SLO; error budget being consumed"5.2 Severity tiers and routing
| Level | Trigger | Action | Goal |
|---|---|---|---|
| page | Burn rate ≥ 14.4 (high rate, fast burn) | Phone/IM, pull someone in immediately | Human engages within 2 hours of an ongoing failure |
| warn | Burn rate ≥ 1 (slow burn) | Join the on-call queue, respond within 30 minutes | Keep the budget from eroding quietly |
| info | Edge cases (high GPU temperature, queue growth) | Log only, don't disturb | Used for retrospectives |
The standard for quiet alerting: at most 1 page per person per week; warns are acceptable but need an "acknowledge-and-close" workflow. If pages fire three weeks in a row, the thresholds or the capacity plan are wrong — that is not something on-call should just "tough out".
6. Logs and Tracing
6.1 Structured logging
Turn logs from "prose for humans" into "JSON for machines" so a log platform (Loki/ELK) can index, filter, and join them with metrics:
python
# app/logging_conf.py — JSON structured logging
import json, logging, sys
class JsonFormatter(logging.Formatter):
def format(self, record):
return json.dumps({
"ts": self.formatTime(record),
"level": record.levelname,
"logger": record.name,
"msg": record.getMessage(),
"request_id": getattr(record, "request_id", None), # correlate with traces
"model": getattr(record, "model", None),
}, ensure_ascii=False)
logging.basicConfig(stream=sys.stdout, level=logging.INFO)
logging.getLogger().handlers[0].setFormatter(JsonFormatter())The hidden value of structured logs is request_id: when logs, metrics, and traces share the same request_id, the three pillars line up when you investigate "which requests are dragging P99 down".
6.2 OpenTelemetry distributed tracing
Inference requests often span multiple hops: gateway → service → database/feature store → inference engine. OpenTelemetry (OTel) captures cross-service call chains in a uniform way:
python
# Minimal setup: auto-instrument FastAPI + a manual span for the inference stage
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor
trace.set_tracer_provider(TracerProvider())
trace.get_tracer_provider().add_span_processor(
BatchSpanProcessor(OTLPSpanExporter(endpoint="http://otel-collector:4318")))
FastAPIInstrumentor.instrument_app(app)
# Add a custom span inside the inference function to quantify "time spent in the engine"
tracer = trace.get_tracer("inference")
def predict_with_trace(images):
with tracer.start_as_current_span("model_inference") as span:
result = engine.run(images)
span.set_attribute("engine", "onnxruntime")
span.set_attribute("batch_size", images.shape[0])
return resultHooked up to Jaeger or Grafana Tempo, OTel lets you visualize "which hop slowed down the requests that hurt P99". A complete observability stack (Collector, Loki, Tempo) is its own engineering effort — scale it to your means. Small teams should get metrics + logs first, and add tracing once cross-service calls appear.
7. The 30-Day Post-Launch Observation Checklist
Getting the metrics in place is only the beginning — the first month after launch is the calibration period:
| Time | What to Watch | Action |
|---|---|---|
| Week 1 | up metric steady at 1; no empty dashboard panels | Fix scrape issues; align numbers with load test results |
| Week 2 | Baseline vs load test: is production P99 ≈ the load-test P99 | Gap > 2x → investigate environment differences |
| Week 2 | How often alerts actually fire | Too noisy → tune burn rates/thresholds; too quiet → inject a fault to verify |
| Week 4 | Error budget consumption rate | Budget > 80% remaining at month end → the 99.9% SLO was a sensible choice |
| Week 4 | Capacity watermark (GPU/queue-depth trends) | Approaching headroom → trigger the scaling plan |
Deliberately Inject One Failure (Game Day)
Within 30 days of launch, run one drill on purpose: kill a worker, inject 5% errors, and see whether alerts fire on time and whether the on-call person knows what to do. An alert you have never verified isn't an alert — it's a placebo.
8. Common Pitfalls
- Inconsistent metric definitions: RPS counted per second in one place and per minute in another; latency with queuing included in one panel and excluded in another — dashboard numbers contradict each other. Fix: write a comment defining each metric's semantics, and put units in panel titles.
- Bad histogram bucket design: buckets stop at 200ms, the P99 target is 100ms, and a degradation to 2s goes unnoticed — the bucket ceiling must cover 10× the P99 target.
- Alert storms: threshold alerts scream on every spike. Fix: switch to burn-rate alerts and add silence rules (maintenance windows).
- Not monitoring the monitoring: Prometheus dies and nobody notices. Fix: add alerts on
upand on Prometheus's own resources. - Prose logs: impossible to search or aggregate. Fix: JSON structure + request_id.
Checklist
- [ ] Four golden signals instrumented; labels low-cardinality and globally consistent;
- [ ]
/metricsscrapeable, with an alert onup == 1; - [ ] GPU metrics (utilization/memory) visible;
- [ ] Five Grafana panels in place; P99 computed from histograms, not gauges;
- [ ] SLO burn-rate alerts tiered (page/warn); no raw threshold spike alerts;
- [ ] Logs in JSON with request_id; OTel wired up for multi-service setups;
- [ ] 30-day observation checklist scheduled, including one game day.
Further Reading
- Monitoring and Observability — the theory behind golden signals, SLOs, and burn rates
- Load Testing and Capacity Planning — how to collect server-side metrics during load tests and set capacity
- Release Strategies: Canary and Rollback — what to monitor during canary releases and auto-rollback conditions
- Deploying a Model from Scratch — the monitoring step in the minimal deployment loop
- Common Pitfalls and Anti-Patterns — incidents caused by running with no monitoring at all
- Performance Metrics and Tuning — latency/throughput metric definitions and tuning
References
- prometheus/client_python: https://github.com/prometheus/client_python
- Prometheus query language docs: https://prometheus.io/docs/prometheus/latest/querying/basics/
- Prometheus Alertmanager: https://prometheus.io/docs/alerting/latest/alertmanager/
- NVIDIA DCGM exporter: https://github.com/NVIDIA/dcgm-exporter
- Grafana documentation: https://grafana.com/docs/grafana/latest/dashboards/
- OpenTelemetry Python docs: https://opentelemetry.io/docs/languages/python/