Skip to content

Monitoring and Observability

At a glance A model deployment without monitoring is flying blind. This article covers the four golden signals for inference services, ML-specific data and concept drift detection, the three pillars of metrics, logs, and traces, and alerting and on-call practices.

Monitoring and Observability ​

One-sentence definition: monitoring means continuously collecting and alerting on metrics, logs, and traces from production inference services, so you find out the model is failing before it's too late — observability turns "is the model actually doing its job?" from a matter of faith into a matter of data.

Industry insight: monitoring a model is far harder than monitoring an ordinary service, because an ordinary service's failure signature is "the code doesn't change, so behavior doesn't change". A model's "code" (the weights) also doesn't change, but the data does — user habits shift, recommendation hot spots move, the style of incoming text drifts, and model quality silently decays. By the time the business side reports "recommendations have been getting worse lately", it has usually been slipping for a week or two. The industry consensus: inference service monitoring must watch two ledgers at once — "system health" and "model health". If the system goes down you need to know within minutes; if the model degrades you need to find out as early as possible (hours to days, via drift detection plus outcome recovery).

1. Why Model Monitoring Is Harder ​

text
Regular service:  code ──► behavior (fixed)     failure = broken code / infrastructure;
                                                restart or rollback restores it
Model service:    code + weights (fixed) ──► behavior (drifts with the data)
                          input distribution ──► quality (can degrade gradually)

Model monitoring must answer three classes of questions:

  1. System level: is the service alive? Fast enough? Does it still have resources?
  2. Data level: has the input distribution changed (data drift)? Has the target concept changed (concept drift)?
  3. Quality level: are the model's metrics (AUC / click-through rate / translation quality) declining?

2. The Four Golden Signals (USE Method) ​

The classic four golden signals (the RED/latency school) apply to inference services just as well:

CategoryExample metricsAlerting rules of thumb
LatencyP50/P99/P999, TTFT/TPOTAlert when P99 exceeds the SLO threshold for 5 consecutive minutes
TrafficQPS, concurrency, input bytesInvestigate any jump or drop of 5x or more
Errors5xx rate, timeout rate, OOM countError rate > 1% is a P1
SaturationGPU utilization, queue depth, VRAM watermarkQueue depth keeps growing / VRAM > 85%

Using the USE method on GPU services

For inference services, track saturation with two metrics: GPU utilization (SM%) and VRAM usage. But remember that inference is usually bandwidth-bound — a low SM% does not mean idle. What really matters is whether the queued-and-waiting count is growing. See Performance Optimization and Capacity Planning.

3. ML-Specific Monitoring: Drift Detection ​

1. Data drift (feature drift) vs concept drift ​

TypeDefinitionExamplesConsequence
Data driftThe input feature distribution changedA surge of new users, sudden weather shifts, changes in text styleThe mix of distributions the model has and hasn't seen becomes unbalanced
Concept driftInput distribution unchanged, but the "input → label" relationship changedUser tastes shift wholesale; exchange-rate policy changesThe patterns the model learned have gone stale

The distinction matters because the remedies differ: data drift → resample the training data, strengthen normalization; concept drift → retrain the model.

2. Detection methods ​

MethodApplies toThreshold rules of thumb
PSI (Population Stability Index)Distribution shift in a single featureAlert when PSI > 0.2 (0.1-0.2 watch, < 0.1 normal)
KS test (Kolmogorov-Smirnov)Whether two distributions differ significantlyp < 0.05 counts as drift
Embedding distribution distance (MMD, KL)High-dimensional semantic distributionsSet relative to a baseline
Prediction distribution monitoringChanges in the output probability distributionOften the earliest signal of concept drift
python
# Sketch of PSI computation (bin-count based)
def psi(expected, actual, bins=10):
    # Cut both distributions into 10 bins using the same quantiles
    # psi += (actual_ratio - expected_ratio) * ln(actual_ratio / expected_ratio)
    # Rule of thumb: > 0.2 counts as significant drift
    return psi_value

Two traps in drift detection

  1. What to monitor: monitor only "features the model truly depends on and that change in production" — monitoring everything just spams alerts;
  2. Refresh the baseline regularly: baselines only make sense relative to "last month"; if you keep the "training set" as a permanent baseline, normal product evolution will trigger alerts.

3. Outcome recovery: the complement to drift detection ​

Drift detection is a proxy signal; in the end it has to land on quality metrics. How to do it:

  • Online A/B tests or shadow scoring that continuously log model outputs against real feedback (e.g., click-through rate);
  • Sampled human labeling (e.g., translation quality, moderation accuracy);
  • Periodic offline evaluation (batch scoring), running an evaluation set against production data.

4. The Three Pillars of Observability ​

PillarToolsWhat matters for inference services
MetricsPrometheus + GrafanaGolden signals, drift metrics, GPU metrics
LogsStructured logs + Loki/ELKPer-request detail, error stacks, feature snapshots
TracesOpenTelemetry + Jaeger/TempoThe full path of one inference request across gateway / service / model

Practical essentials for landing the three pillars in inference services:

  1. Traces must carry the model version and an input summary: when debugging, "did this request hit the v2 model or v3?" is the single most common question;
  2. Structured logs: JSON with trace_id, model_version, and latency_ms fields, so they aggregate cleanly;
  3. Request sampling: full logging is expensive. A common combo: log every slow request (P99 outliers) + sample normal requests at 1%.

For complete implementation guidance see Observability in Practice.

5. Alert Design: SLO Burn Rates and Severity Tiers ​

1. SLO burn rate ​

Google SRE's alerting framework: alert based on how fast the error budget is being burned. If the SLO is "P99 < 100ms, met 99.9% of the time over 30 days":

  • Burn rate = actual error rate ÷ allowed error rate (0.1%);
  • Burn rate ≥ 14.4 sustained for 1 hour → page on-call (because 14.4 × 1h burns through about 2% of the 30-day error budget);
  • Burn rate ≥ 6 sustained for 6 hours → alert.

2. Tiered alerting to avoid alert fatigue ​

SeverityMeaningExamplesResponse
P1Service down / criticalInference 5xx > 5%, GPU card failureImmediate on-call page
P2Degraded functionalityDrift alerts, sustained P99 breachesHandle during working hours
P3Potential riskVRAM watermark at 80%, error budget down to 10%Log and follow up

Alert fatigue is a real incident

Too many alerts → nobody reads them carefully → a real P1 gets ignored too. Rule of thumb: each on-call engineer should be paged at most 1-2 times per week. Beyond that, raise thresholds or delete metrics.

6. Post-Launch Monitoring Checklist (the First 30 Days) ​

A freshly launched model needs different monitoring than a mature one. In the first 30 days, focus on:

Time windowFocusWhy
Day 1Loading/warmup success, first-hour latency distributionScaling and cold caches expose configuration problems fastest
Week 1P99 stability, error rate, VRAM watermarkCatch resource estimation errors
Week 2Feature distribution vs training distribution, quality baselineThe train/serve distribution gap shows up now
Day 30Full A/B quality comparison against the old versionDecide whether to roll back or cement the new version

Connection to MLOps: for launch and rollback workflows see Release Strategies: Canary and Rollback; for the full pipeline see The MLOps Deployment Pipeline.

Trade-offs ​

DecisionOptionsHow to choose
Number of metricsLean vs everythingStart with the golden signals; every added metric needs a purpose
Sampling vs full captureCheaper vs completeFull capture for slow requests, sampling for normal ones
Drift detection frequencyReal-time vs hourlyReal-time for high-value online scenarios; hourly is enough otherwise
Alert thresholdsSensitive vs looseWork backwards from "≤ 1-2 pages per week"
Self-hosted vs managedBuild the stack yourself vs SaaSTeam size and budget decide; for tooling see Resources

One-line summary: an inference service needs both ledgers watched — system health and model health. Golden signals keep it "alive", drift detection keeps it "still accurate", the three pillars give you the ability to investigate, and tiered alerting guarantees someone responds.

Further Reading ​

References ​