Appearance
Monitoring and Observability
One-sentence definition: monitoring means continuously collecting and alerting on metrics, logs, and traces from production inference services, so you find out the model is failing before it's too late — observability turns "is the model actually doing its job?" from a matter of faith into a matter of data.
Industry insight: monitoring a model is far harder than monitoring an ordinary service, because an ordinary service's failure signature is "the code doesn't change, so behavior doesn't change". A model's "code" (the weights) also doesn't change, but the data does — user habits shift, recommendation hot spots move, the style of incoming text drifts, and model quality silently decays. By the time the business side reports "recommendations have been getting worse lately", it has usually been slipping for a week or two. The industry consensus: inference service monitoring must watch two ledgers at once — "system health" and "model health". If the system goes down you need to know within minutes; if the model degrades you need to find out as early as possible (hours to days, via drift detection plus outcome recovery).
1. Why Model Monitoring Is Harder
text
Regular service: code ──► behavior (fixed) failure = broken code / infrastructure;
restart or rollback restores it
Model service: code + weights (fixed) ──► behavior (drifts with the data)
input distribution ──► quality (can degrade gradually)Model monitoring must answer three classes of questions:
- System level: is the service alive? Fast enough? Does it still have resources?
- Data level: has the input distribution changed (data drift)? Has the target concept changed (concept drift)?
- Quality level: are the model's metrics (AUC / click-through rate / translation quality) declining?
2. The Four Golden Signals (USE Method)
The classic four golden signals (the RED/latency school) apply to inference services just as well:
| Category | Example metrics | Alerting rules of thumb |
|---|---|---|
| Latency | P50/P99/P999, TTFT/TPOT | Alert when P99 exceeds the SLO threshold for 5 consecutive minutes |
| Traffic | QPS, concurrency, input bytes | Investigate any jump or drop of 5x or more |
| Errors | 5xx rate, timeout rate, OOM count | Error rate > 1% is a P1 |
| Saturation | GPU utilization, queue depth, VRAM watermark | Queue depth keeps growing / VRAM > 85% |
Using the USE method on GPU services
For inference services, track saturation with two metrics: GPU utilization (SM%) and VRAM usage. But remember that inference is usually bandwidth-bound — a low SM% does not mean idle. What really matters is whether the queued-and-waiting count is growing. See Performance Optimization and Capacity Planning.
3. ML-Specific Monitoring: Drift Detection
1. Data drift (feature drift) vs concept drift
| Type | Definition | Examples | Consequence |
|---|---|---|---|
| Data drift | The input feature distribution changed | A surge of new users, sudden weather shifts, changes in text style | The mix of distributions the model has and hasn't seen becomes unbalanced |
| Concept drift | Input distribution unchanged, but the "input → label" relationship changed | User tastes shift wholesale; exchange-rate policy changes | The patterns the model learned have gone stale |
The distinction matters because the remedies differ: data drift → resample the training data, strengthen normalization; concept drift → retrain the model.
2. Detection methods
| Method | Applies to | Threshold rules of thumb |
|---|---|---|
| PSI (Population Stability Index) | Distribution shift in a single feature | Alert when PSI > 0.2 (0.1-0.2 watch, < 0.1 normal) |
| KS test (Kolmogorov-Smirnov) | Whether two distributions differ significantly | p < 0.05 counts as drift |
| Embedding distribution distance (MMD, KL) | High-dimensional semantic distributions | Set relative to a baseline |
| Prediction distribution monitoring | Changes in the output probability distribution | Often the earliest signal of concept drift |
python
# Sketch of PSI computation (bin-count based)
def psi(expected, actual, bins=10):
# Cut both distributions into 10 bins using the same quantiles
# psi += (actual_ratio - expected_ratio) * ln(actual_ratio / expected_ratio)
# Rule of thumb: > 0.2 counts as significant drift
return psi_valueTwo traps in drift detection
- What to monitor: monitor only "features the model truly depends on and that change in production" — monitoring everything just spams alerts;
- Refresh the baseline regularly: baselines only make sense relative to "last month"; if you keep the "training set" as a permanent baseline, normal product evolution will trigger alerts.
3. Outcome recovery: the complement to drift detection
Drift detection is a proxy signal; in the end it has to land on quality metrics. How to do it:
- Online A/B tests or shadow scoring that continuously log model outputs against real feedback (e.g., click-through rate);
- Sampled human labeling (e.g., translation quality, moderation accuracy);
- Periodic offline evaluation (batch scoring), running an evaluation set against production data.
4. The Three Pillars of Observability
| Pillar | Tools | What matters for inference services |
|---|---|---|
| Metrics | Prometheus + Grafana | Golden signals, drift metrics, GPU metrics |
| Logs | Structured logs + Loki/ELK | Per-request detail, error stacks, feature snapshots |
| Traces | OpenTelemetry + Jaeger/Tempo | The full path of one inference request across gateway / service / model |
Practical essentials for landing the three pillars in inference services:
- Traces must carry the model version and an input summary: when debugging, "did this request hit the v2 model or v3?" is the single most common question;
- Structured logs: JSON with
trace_id,model_version, andlatency_msfields, so they aggregate cleanly; - Request sampling: full logging is expensive. A common combo: log every slow request (P99 outliers) + sample normal requests at 1%.
For complete implementation guidance see Observability in Practice.
5. Alert Design: SLO Burn Rates and Severity Tiers
1. SLO burn rate
Google SRE's alerting framework: alert based on how fast the error budget is being burned. If the SLO is "P99 < 100ms, met 99.9% of the time over 30 days":
- Burn rate = actual error rate ÷ allowed error rate (0.1%);
- Burn rate ≥ 14.4 sustained for 1 hour → page on-call (because 14.4 × 1h burns through about 2% of the 30-day error budget);
- Burn rate ≥ 6 sustained for 6 hours → alert.
2. Tiered alerting to avoid alert fatigue
| Severity | Meaning | Examples | Response |
|---|---|---|---|
| P1 | Service down / critical | Inference 5xx > 5%, GPU card failure | Immediate on-call page |
| P2 | Degraded functionality | Drift alerts, sustained P99 breaches | Handle during working hours |
| P3 | Potential risk | VRAM watermark at 80%, error budget down to 10% | Log and follow up |
Alert fatigue is a real incident
Too many alerts → nobody reads them carefully → a real P1 gets ignored too. Rule of thumb: each on-call engineer should be paged at most 1-2 times per week. Beyond that, raise thresholds or delete metrics.
6. Post-Launch Monitoring Checklist (the First 30 Days)
A freshly launched model needs different monitoring than a mature one. In the first 30 days, focus on:
| Time window | Focus | Why |
|---|---|---|
| Day 1 | Loading/warmup success, first-hour latency distribution | Scaling and cold caches expose configuration problems fastest |
| Week 1 | P99 stability, error rate, VRAM watermark | Catch resource estimation errors |
| Week 2 | Feature distribution vs training distribution, quality baseline | The train/serve distribution gap shows up now |
| Day 30 | Full A/B quality comparison against the old version | Decide whether to roll back or cement the new version |
Connection to MLOps: for launch and rollback workflows see Release Strategies: Canary and Rollback; for the full pipeline see The MLOps Deployment Pipeline.
Trade-offs
| Decision | Options | How to choose |
|---|---|---|
| Number of metrics | Lean vs everything | Start with the golden signals; every added metric needs a purpose |
| Sampling vs full capture | Cheaper vs complete | Full capture for slow requests, sampling for normal ones |
| Drift detection frequency | Real-time vs hourly | Real-time for high-value online scenarios; hourly is enough otherwise |
| Alert thresholds | Sensitive vs loose | Work backwards from "≤ 1-2 pages per week" |
| Self-hosted vs managed | Build the stack yourself vs SaaS | Team size and budget decide; for tooling see Resources |
One-line summary: an inference service needs both ledgers watched — system health and model health. Golden signals keep it "alive", drift detection keeps it "still accurate", the three pillars give you the ability to investigate, and tiered alerting guarantees someone responds.
Further Reading
- Observability in Practice — a hands-on guide to building the three pillars from scratch
- The MLOps Deployment Pipeline — where monitoring gates the pipeline
- Performance Optimization and Capacity Planning — saturation metrics and capacity math
- Common Pitfalls and Anti-Patterns — classic incidents caused by missing monitoring
- Security, Privacy, and Compliance — the privacy boundaries of monitoring logs