Appearance
Model Gateway and Canary Releases: Routing, Rate Limiting, Canaries, and A/B Testing
In one sentence: a model gateway is the single traffic entry in front of a model-serving cluster; it routes requests to the right model/version and layers rate limiting, circuit breaking, gradual traffic shifting, caching, and auditing on top — it's the layer you grow into once "model versions start multiplying" and direct-to-service access no longer scales.
Why it's worth doing: the moment you have several models and frequent updates, the questions arrive: how do you let 5% of traffic try the new model without affecting the rest? If the new version underperforms, how do you switch back within 10 seconds? Several teams each hang their own model service — who owns the unified entry point? The answer is a gateway. This article presents two concrete implementations — Nginx (weighted routing + gradual rollout) and Envoy (advanced routing) — then covers the complete canary release and A/B experiment workflow.
1. What a Model Gateway Does
| Responsibility | Description |
|---|---|
| Routing | Send requests to different models or versions by path/header |
| Rate limiting | Protect downstream model services from being crushed by traffic spikes (token bucket / QPS caps) |
| Circuit breaking | Fail fast when the downstream keeps failing, instead of piling on |
| Canary / gradual rollout | Shift traffic to a new version by weight or by user |
| Caching | Reuse results for identical inputs (e.g. face retrieval, templated queries) |
| Auditing | Log caller, model version, latency, and result for A/B analysis and reconciliation |
| Auth | Validate API keys and control who can call which model |
Conclusion first: small teams start with Nginx (config as code, zero maintenance); mid-to-large teams use Envoy (programmable, observable, canary-capable); when you need "model-level" semantics (versions, automatic rollback, multiple frameworks), use a dedicated inference platform like KServe. For the full decision tree, see Choosing Frameworks and Platforms.
2. Comparing the Options
| Option | Strengths | Weaknesses | Best for |
|---|---|---|---|
| Simple Nginx routing | Zero dependencies, intuitive config | No built-in circuit breaking/retries; canary by weight only | 2-3 models, quick launch |
| Advanced Envoy routing | Header/weight/shadow-traffic routing, built-in circuit breaking and retries, xDS dynamic updates | Steep learning curve, complex config DSL | Microservices, multiple teams, K8s-native |
| Application-level gateway (custom/KServe) | Understands model semantics: version management, autoscaling, batch entry point | High cost to build; KServe ties you to K8s | Teams whose product is the model |
3. Walkthrough 1: Routing Between Two Model Versions (Nginx)
Scenario: the CTR model is upgrading from v1 to v2, and 10% of traffic should go to v2 first.
nginx
# /etc/nginx/conf.d/model-gateway.conf
upstream ctr_v1 {
server 10.0.1.10:8000; # old model service (FastAPI/Triton, say)
}
upstream ctr_v2 {
server 10.0.1.11:8000; # new model service
}
server {
listen 8080;
# Header-based canary: requests carrying x-canary: v2 are forced to the new version (internal testing / specific user groups)
location = /predict {
if ($http_x_canary = "v2") {
proxy_pass http://ctr_v2;
}
proxy_pass http://ctr_v1;
}
}For weighted rollouts, Nginx's split_clients hashes a request field to produce a stable percentage split — the same user always lands on the same version, which matters for experiment consistency:
nginx
# Hash by user ID: 10% of user_id hash values land on v2
split_clients "${http_x_user_id}" $ctr_backend {
10% ctr_v2;
* ctr_v1;
}
location = /predict {
proxy_pass http://$ctr_backend;
}Why hash by user instead of round-robining by request count
Both canary releases and A/B tests require that the same user sees consistent results (no drift). Round-robining by request count sends the same user's previous request to v1 and the next one to v2, and the business side ends up confused about "what the new model actually does". So split traffic either by user hash or by cookie/header.
4. Walkthrough 2: The Canary Release Workflow (5% → 20% → 50% → 100%)
The goal of a canary is to ramp up the new version gradually and stop at the first sign of degradation. The standard workflow:
- Deploy the new version with zero traffic: v2 comes up and passes a smoke test (correctness + latency);
- 5% of traffic: run for 30-60 minutes, watching core metrics v1 vs v2;
- 20% → 50%: observe each step for at least 30 minutes; continue only if metrics hold;
- 100% full rollout: after full rollout, keep observing for a while, then retire the old version.
Metrics to watch while shifting traffic (definitions must be identical for v1/v2, or comparison is meaningless):
| Metric | What to check |
|---|---|
| P99 latency | The new version must not be slower than the old beyond a set threshold (e.g. +20%) |
| Error rate | 5xx/timeout ratio; stop on any degradation |
| Business outcome | e.g. click-through rate, conversion — the canary's final judge |
| Resources | GPU utilization, GPU memory, CPU — confirm the new model can carry the load |
Changing weights is a one-line edit in Nginx (10% → 20%), or a route-config change in Envoy (see below). For the complete release methodology, see Release Strategies: Canary and Rollback.
5. Walkthrough 3: Traffic-Based Canary and A/B with Envoy
Envoy's strength is built-in routing by header, weight, and shadow traffic. The config below routes 10% of traffic to v2 and additionally supports an x-ab: b header that forces traffic to v2 (for A/B experiment cohorts):
yaml
# envoy.yaml (excerpt)
static_resources:
clusters:
- name: ctr_v1
connect_timeout: 0.25s
type: STRICT_DNS
lb_policy: ROUND_ROBIN
load_assignment:
cluster_name: ctr_v1
endpoints:
- lb_endpoints:
- endpoint: { address: { socket_address: { address: ctr-v1, port_value: 8000 } } }
- name: ctr_v2
connect_timeout: 0.25s
type: STRICT_DNS
lb_policy: ROUND_ROBIN
load_assignment:
cluster_name: ctr_v2
endpoints:
- lb_endpoints:
- endpoint: { address: { socket_address: { address: ctr-v2, port_value: 8000 } } }
http_filters:
- name: envoy.filters.http.router
typed_config:
"@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router
# Route rules (in practice these live in the VirtualHost)
route_config:
virtual_hosts:
- name: model_api
domains: ["*"]
routes:
# A/B: requests carrying x-ab: b are forced to v2 (control group)
- match:
headers:
- name: x-ab
exact_match: "b"
prefix: /predict
route: { cluster: ctr_v2 }
# Canary: all other requests split 90/10 by weight
- match: { prefix: /predict }
route:
weighted_clusters:
clusters:
- name: ctr_v1
weight: 90
- name: ctr_v2
weight: 10Envoy's built-in circuit breaking and retries are one config block:
yaml
route:
cluster: ctr_v1
retry_policy:
retry_on: "5xx"
num_retries: 2
timeout: 2sConclusion first: for production rollouts, prefer Envoy or a K8s gateway (Istio/Argo Rollouts) over hand-written Nginx if blocks — the config updates dynamically (xDS), and you get circuit breaking, retries, and metrics integration. Nginx fits the "get it done in 5 minutes" temporary canary.
Rate Limiting: The Gateway's Second Lifesaver
Downstream model capacity is fixed; gateway rate limiting keeps a burst from crushing it. Nginx token-bucket limiting:
nginx
# At most 10 requests per second per IP, bursts up to 20; over-limit returns 429
limit_req_zone $binary_remote_addr zone=model_api:10m rate=10r/s;
location = /predict {
limit_req zone=model_api burst=20 nodelay;
limit_req_status 429;
proxy_pass http://$ctr_backend;
}Envoy's local rate limiting (envoy.filters.http.local_ratelimit) works the same way and can also limit by header dimensions like x-user-id, letting you set "N QPS per model" against downstream capacity. How to pick the numbers: measure per-instance capacity with the methods in Load Testing and Capacity Planning, then keep a 20% margin.
Circuit Breaking: Protecting Yourself and the Downstream
Default circuit-breaking parameters, conclusion first: failure threshold of 5-10 consecutive failures, break window of 10-30 seconds, fail fast during the break (return 503). Envoy config:
yaml
clusters:
- name: ctr_v1
circuit_breakers:
thresholds:
- priority: DEFAULT
max_connections: 1024
max_pending_requests: 100 # queue cap
max_requests: 2000
max_retries: 3Circuit breaking vs rate limiting in one line: rate limiting blocks "too much coming in"; circuit breaking blocks "nothing getting through" — when the downstream is already failing, circuit breaking stops you waiting around and fails fast so upstream can degrade. Combined with the cascading degradation described in Online Inference for Recommender Systems, the whole chain becomes resilient.
trace_id: The Foundation for Rollout and A/B Reconciliation
After shifting traffic you must be able to answer "which model version handled this request", so the gateway must generate and propagate a trace_id and write the model_version into the response headers:
nginx
# Put the model version in Nginx response headers so the business side can reconcile afterward
add_header X-Model-Version "v1" always;
add_header X-Trace-Id $request_id always;The trace_id flows through gateway → model service → logs → analytics events; A/B analysis joins business outcomes on it. For the implementation details of this tracing chain, see Observability in Practice.
6. A/B Experiment Design: Splitting, Instrumentation, Significance, Duration
A canary is "outcome defense" (the new version must not be worse); an A/B test is "outcome proof" (is the new version significantly better). The full design has four steps:
- Split: 50/50 experiment/control (or per your traffic budget), using user-ID hashing for stability (same user, same group) — the
x-abscheme from Section 5 or the gateway's built-in splitting; - Instrument: the gateway stamps
model_version(v1/v2) andtrace_idinto responses, and the business side joins model versions against business outcomes (clicks/conversions). The version must come back with the response, or reconciliation fails afterward; - Significance testing: use a two-sample proportion test / chi-squared test (conversion-type metrics) or a t-test (numeric metrics); only conclude when
p < 0.05, and watch the effect size (how many percentage points); - Duration: cover at least one full business cycle (a week including the weekend, for example). Estimate the sample size in advance with a power analysis — the classic mistake is concluding after 2 days with an undersized sample and no significant difference to show for it.
A/B splitting, instrumentation, and result review all ultimately land on the metric system in Monitoring and Observability; otherwise the data won't line up and experiment conclusions can't be trusted.
7. Rollback: One Switch Back to Safety
When the canary shows degradation, the rollback action is a single move: change the weight from 10% back to 0%; the v2 service can stay up for investigation.
nginx
# Rollback: v2 weight to zero
split_clients "${http_x_user_id}" $ctr_backend {
0% ctr_v2;
* ctr_v1;
}Key points:
- Don't delete the service when rolling back — just cut the traffic. Keeping v2 around to pull logs and inspect errors is much faster than rebuilding the environment;
- Set an automatic rollback threshold at the gateway: if the new version's error rate > 1% or P99 exceeds the threshold for N consecutive minutes, switch back automatically (Envoy/Istio or a homegrown script can do this);
- If a regression appears after full rollout (covering traffic patterns the canary didn't), cut back the same way and run a postmortem. For the complete rollback playbook, see Release Strategies: Canary and Rollback.
Common Pitfalls and Troubleshooting
| Pitfall | Symptom | Fix |
|---|---|---|
| Inconsistent metric definitions | v1/v2 metrics can't be compared | Unify metric definitions and instrumentation fields — see Section 6 |
| Cache pollution | Gateway cache mixes old and new version results | Include the model version in the cache key; disable or segregate caching for the new version during rollout |
| Non-sticky splitting | The same user bounces between two versions | Hash by user ID (split_clients), don't round-robin by request |
| Weight change not taking effect | Nginx edited but not reloaded | nginx -s reload; Envoy needs an xDS push or a config reload |
| Undersized experiment sample | Declaring conclusions on a non-significant p-value | Compute the sample size first, run the full cycle, then conclude |
| Gateway timeout shorter than model latency | Slow model requests get 504 from the gateway | The gateway timeout must exceed the model's P99 latency |
| Canary checks latency but not outcomes | Latency is fine but business metrics drop | Instrument business metrics for both canary and A/B; the gateway gives you traffic, not conclusions |
Further Reading
- Release Strategies: Canary and Rollback — methodology and decision framework for canary/blue-green/rolling releases
- MLOps Deployment Pipeline — the full path from training to production behind gateway traffic shifting
- Serving and Inference APIs — what the model services behind the gateway should look like
- Deployment Architecture Patterns — where the gateway sits in the overall deployment architecture and how it evolves
- Monitoring and Observability — the metrics, logs, and tracing needed for A/B and canary reconciliation
- Online Serving with FastAPI + Docker — a typical model service behind the gateway
- LLM Serving with vLLM — LLM services benefit from gateway canaries too, and often need model-based routing
References
- Nginx documentation: https://nginx.org/en/docs/
- Envoy documentation: https://www.envoyproxy.io/docs
- KServe documentation: https://kserve.github.io/website/
- Argo Rollouts (K8s canary releases): https://argoproj.github.io/rollouts/
- Nginx split_clients module: https://nginx.org/en/docs/http/ngx_http_split_clients_module.html