Skip to content

Putting Observability into Practice

At a glance A complete walkthrough for wiring an inference service into Prometheus + Grafana: custom metrics (request counts/latency histograms/error rates), GPU metric collection, SLO burn-rate alert rules, and logs with tracing, plus a 30-day post-launch observation checklist.

Putting observability into practice means turning "how is the service actually doing" into an engineering capability you can answer at any moment: the three pillars — metrics, logs, and traces — all in place, with alerts that are precise and not noisy. Why bother? An inference service without observability is like driving without a dashboard — failures such as model drift, latency degradation, and GPU memory leaks never crash loudly; they just let users feel "it's slower and dumber" before you do. This guide walks through every step from zero to production-grade observability, with working code and configs: metric instrumentation, scraping, GPU collection, Grafana dashboards, SLO burn-rate alerts, structured logging, and distributed tracing — ending with a 30-day post-launch observation checklist. For the theory and SLO definitions, see Monitoring and Observability.

The Goal in One Sentence

No missing metrics, searchable logs, quiet alerts — the four golden signals (RPS, latency, error rate, saturation) all covered, and alerts fire only when "a human needs to act".

1. Instrumenting Metrics (Prometheus Client) ​

Use prometheus-client to instrument the three golden metrics for an inference service: request counters, latency histograms, and error counters. This is the full version of the metrics module from Deploying a Model from Scratch.

python
# app/metrics.py — all metric definitions in one file, referenced across the project
from prometheus_client import Counter, Histogram, Gauge

# Request counter: labels distinguish model and version for easy per-version comparison
REQUESTS = Counter("predict_requests_total", "Total inference requests",
                   ["model", "version"])
ERRORS = Counter("predict_errors_total", "Failed inference requests",
                 ["model", "version", "error_type"])
# Latency histogram: buckets sized to the service's latency range, covering the P99 target
LATENCY = Histogram("predict_latency_seconds", "Inference latency",
                    ["model", "version"],
                    buckets=(0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1.0, 2.5))
# Queue depth: saturation signal; a Gauge sampled every 5 seconds is enough
QUEUE = Gauge("predict_queue_depth", "Requests currently queued")

Three rules for label design (labels are a high-risk design point — get them wrong and everything downstream suffers):

  1. Cardinality must stay low: label value combinations × time series = storage explosion. High-cardinality values like user_id or request_id must never become labels — they belong in logs only;
  2. Only include dimensions you will actually query by: model, version, error_type, region are reasonable; temporary debugging fields stay out of production metrics;
  3. Keep label sets globally consistent: all metrics share the same model/version labels so alerts and dashboards can join across metrics.

Instrument at the service's entrance and exit (see main.py in Deploying a Model from Scratch); use a middleware or decorator to handle it uniformly instead of sprinkling inc() calls through business code:

python
# app/middleware.py — unified instrumentation
import time
from starlette.middleware.base import BaseHTTPMiddleware
from app.metrics import REQUESTS, ERRORS, LATENCY

class MetricsMiddleware(BaseHTTPMiddleware):
    async def dispatch(self, request, call_next):
        t0 = time.perf_counter()
        try:
            resp = await call_next(request)
        except Exception:
            ERRORS.labels(model="mobilenetv2", version="v1",
                          error_type="exception").inc()
            raise
        LATENCY.labels(model="mobilenetv2", version="v1").observe(
            time.perf_counter() - t0)
        REQUESTS.labels(model="mobilenetv2", version="v1").inc()
        return resp
python
# app/main.py — expose /metrics for Prometheus to scrape
from prometheus_client import generate_latest, CONTENT_TYPE_LATEST
from fastapi import Response

@app.get("/metrics")
def metrics():
    return Response(generate_latest(), media_type=CONTENT_TYPE_LATEST)

2. Service Discovery and Scraping ​

2.1 Simple setup: static_configs ​

yaml
# prometheus.yml
global:
  scrape_interval: 15s          # scrape interval; drop to 10s for latency-sensitive services
  evaluation_interval: 30s      # alert rule evaluation interval

scrape_configs:
  - job_name: inference-api
    static_configs:
      - targets: ["api:8000"]   # use the service name inside compose/Docker networks
        labels: { env: prod }
  - job_name: dcgm              # GPU metrics, see Step 3
    static_configs:
      - targets: ["dcgm-exporter:9400"]

2.2 Kubernetes: ServiceMonitor ​

In K8s, service instances come and go, so use declarative discovery with a ServiceMonitor (paired with Prometheus Operator / kube-prometheus-stack). Once the service carries the prometheus.io/scrape: "true" annotation it is scraped automatically:

yaml
# ServiceMonitor.yaml
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: inference-api
spec:
  selector:
    matchLabels:
      app: inference-api          # matches the Service's selector
  endpoints:
    - port: http                 # the port name declared on the Service
      path: /metrics
      interval: 15s
  namespaceSelector:
    any: true

Failed Scrapes Are the Sneakiest Monitoring Incident

When Prometheus cannot reach a target it usually doesn't error out — it just drops data silently: dashboards slowly empty out and alerts quietly stop working. For the first week after launch, check the up metric daily (up{job="inference-api"} == 1) — it is the only evidence that "the monitoring itself is alive".

3. Collecting GPU Metrics (dcgm-exporter) ​

When the inference service runs on GPUs, GPU utilization and memory watermark are the core signals for capacity and fault analysis (paired with the "GPU pegged vs CPU pegged" analysis in Load Testing and Capacity Planning). The standard approach is the NVIDIA DCGM exporter:

bash
# Start the DCGM exporter, exposing :9400/metrics (GPU utilization, memory, temperature, power)
docker run -d --gpus all --rm --cap-add SYS_ADMIN \
  --name dcgm-exporter \
  nvcr.io/nvidia/dcgm-exporter:3.3.5-ubuntu22.04

Key metrics (DCGM's own names):

text
# GPU utilization (%), grouped by device
DCGM_FI_DEV_GPU_UTIL
# GPU memory used (bytes)
DCGM_FI_DEV_FB_USED
# GPU temperature and power
DCGM_FI_DEV_GPU_TEMP
DCGM_FI_DEV_POWER_USAGE

Without containers, nvidia-smi --query-gpu can output similar data, but not in Prometheus format — you would have to bridge it yourself through node_exporter's textfile collector, or use a community exporter such as nvidia_gpu_exporter. Debugging GPU memory problems also ties into pitfall #4 of Memory Leaks and OOM.

4. Grafana Dashboards ​

After connecting Grafana to the Prometheus data source, build an "Inference Service Overview" dashboard with five panels laid out in order:

PanelPromQLWhat to Look For
RPSsum(rate(predict_requests_total[1m])) by (model)Traffic trend; watch traffic shifts around releases
P50/P99 latencyhistogram_quantile(0.99, sum by (le, model) (rate(predict_latency_seconds_bucket[5m])))Whether P99 stays within the SLO
Error ratesum(rate(predict_errors_total[5m])) / sum(rate(predict_requests_total[5m]))Share of 5xx/inference errors
GPU utilizationavg(DCGM_FI_DEV_GPU_UTIL)GPU pegged vs idle
Queue depthpredict_queue_depthSaturation; a growing queue precedes an avalanche

histogram_quantile's P99 estimate carries statistical error — denser buckets are more accurate, at the cost of storage. Design rule: buckets should span the range from "P50 up to 10× the P99 target"; for a P99 target of 100ms, buckets run from 5ms to 2.5s, growing roughly 2x per step (like the buckets array above).

5. Alert Rules (SLO Burn Rates + Severity Tiers) ​

5.1 Burn-rate alerts — the key to quiet alerting ​

"Threshold alerts" (fire when latency > 500ms) scream during transient spikes and stay silent during slow degradation — the root of alert storms. SLO burn rate alerts fire based on how fast the SLO error budget is being consumed: they trigger only when the budget burns fast enough over a given window. See Monitoring and Observability for the underlying theory.

yaml
# prometheus-alerts.yml
groups:
  - name: inference-slo
    rules:
      # 30-day SLO of 99.9% (0.1% errors allowed)
      # Burn rate 14.4 (5-minute window): burns 2% of the budget in about 2 hours → page
      - alert: HighErrorRatePage
        expr: |
          sum(rate(predict_errors_total[5m]))
            / sum(rate(predict_requests_total[5m]))
              > 0.1 * 14.4
        labels:
          severity: page
        annotations:
          summary: "Error burn rate above 14.4; 2% of budget gone within 2 hours"
      # Burn rate 1.0 (1-hour window): budget exhausted within 30 days → warn, on-call queue
      - alert: ErrorBudgetWarn
        expr: |
          sum(rate(predict_errors_total[1h]))
            / sum(rate(predict_requests_total[1h]))
              > 0.1
        labels:
          severity: warn
        annotations:
          summary: "1-hour error rate exceeds SLO; error budget being consumed"

5.2 Severity tiers and routing ​

LevelTriggerActionGoal
pageBurn rate ≥ 14.4 (high rate, fast burn)Phone/IM, pull someone in immediatelyHuman engages within 2 hours of an ongoing failure
warnBurn rate ≥ 1 (slow burn)Join the on-call queue, respond within 30 minutesKeep the budget from eroding quietly
infoEdge cases (high GPU temperature, queue growth)Log only, don't disturbUsed for retrospectives

The standard for quiet alerting: at most 1 page per person per week; warns are acceptable but need an "acknowledge-and-close" workflow. If pages fire three weeks in a row, the thresholds or the capacity plan are wrong — that is not something on-call should just "tough out".

6. Logs and Tracing ​

6.1 Structured logging ​

Turn logs from "prose for humans" into "JSON for machines" so a log platform (Loki/ELK) can index, filter, and join them with metrics:

python
# app/logging_conf.py — JSON structured logging
import json, logging, sys

class JsonFormatter(logging.Formatter):
    def format(self, record):
        return json.dumps({
            "ts": self.formatTime(record),
            "level": record.levelname,
            "logger": record.name,
            "msg": record.getMessage(),
            "request_id": getattr(record, "request_id", None),  # correlate with traces
            "model": getattr(record, "model", None),
        }, ensure_ascii=False)

logging.basicConfig(stream=sys.stdout, level=logging.INFO)
logging.getLogger().handlers[0].setFormatter(JsonFormatter())

The hidden value of structured logs is request_id: when logs, metrics, and traces share the same request_id, the three pillars line up when you investigate "which requests are dragging P99 down".

6.2 OpenTelemetry distributed tracing ​

Inference requests often span multiple hops: gateway → service → database/feature store → inference engine. OpenTelemetry (OTel) captures cross-service call chains in a uniform way:

python
# Minimal setup: auto-instrument FastAPI + a manual span for the inference stage
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor

trace.set_tracer_provider(TracerProvider())
trace.get_tracer_provider().add_span_processor(
    BatchSpanProcessor(OTLPSpanExporter(endpoint="http://otel-collector:4318")))

FastAPIInstrumentor.instrument_app(app)

# Add a custom span inside the inference function to quantify "time spent in the engine"
tracer = trace.get_tracer("inference")
def predict_with_trace(images):
    with tracer.start_as_current_span("model_inference") as span:
        result = engine.run(images)
        span.set_attribute("engine", "onnxruntime")
        span.set_attribute("batch_size", images.shape[0])
        return result

Hooked up to Jaeger or Grafana Tempo, OTel lets you visualize "which hop slowed down the requests that hurt P99". A complete observability stack (Collector, Loki, Tempo) is its own engineering effort — scale it to your means. Small teams should get metrics + logs first, and add tracing once cross-service calls appear.

7. The 30-Day Post-Launch Observation Checklist ​

Getting the metrics in place is only the beginning — the first month after launch is the calibration period:

TimeWhat to WatchAction
Week 1up metric steady at 1; no empty dashboard panelsFix scrape issues; align numbers with load test results
Week 2Baseline vs load test: is production P99 ≈ the load-test P99Gap > 2x → investigate environment differences
Week 2How often alerts actually fireToo noisy → tune burn rates/thresholds; too quiet → inject a fault to verify
Week 4Error budget consumption rateBudget > 80% remaining at month end → the 99.9% SLO was a sensible choice
Week 4Capacity watermark (GPU/queue-depth trends)Approaching headroom → trigger the scaling plan

Deliberately Inject One Failure (Game Day)

Within 30 days of launch, run one drill on purpose: kill a worker, inject 5% errors, and see whether alerts fire on time and whether the on-call person knows what to do. An alert you have never verified isn't an alert — it's a placebo.

8. Common Pitfalls ​

  1. Inconsistent metric definitions: RPS counted per second in one place and per minute in another; latency with queuing included in one panel and excluded in another — dashboard numbers contradict each other. Fix: write a comment defining each metric's semantics, and put units in panel titles.
  2. Bad histogram bucket design: buckets stop at 200ms, the P99 target is 100ms, and a degradation to 2s goes unnoticed — the bucket ceiling must cover 10× the P99 target.
  3. Alert storms: threshold alerts scream on every spike. Fix: switch to burn-rate alerts and add silence rules (maintenance windows).
  4. Not monitoring the monitoring: Prometheus dies and nobody notices. Fix: add alerts on up and on Prometheus's own resources.
  5. Prose logs: impossible to search or aggregate. Fix: JSON structure + request_id.

Checklist ​

  • [ ] Four golden signals instrumented; labels low-cardinality and globally consistent;
  • [ ] /metrics scrapeable, with an alert on up == 1;
  • [ ] GPU metrics (utilization/memory) visible;
  • [ ] Five Grafana panels in place; P99 computed from histograms, not gauges;
  • [ ] SLO burn-rate alerts tiered (page/warn); no raw threshold spike alerts;
  • [ ] Logs in JSON with request_id; OTel wired up for multi-service setups;
  • [ ] 30-day observation checklist scheduled, including one game day.

Further Reading ​

References ​