Skip to content

Model Gateway and Canary Releases

At a glance When multiple models or model versions coexist, the gateway owns routing, rate limiting, gradual rollouts, and A/B testing. This hands-on article builds a model gateway with Nginx/Envoy plus application-layer routing, and walks through the complete canary release and A/B experiment workflow.

Model Gateway and Canary Releases: Routing, Rate Limiting, Canaries, and A/B Testing ​

In one sentence: a model gateway is the single traffic entry in front of a model-serving cluster; it routes requests to the right model/version and layers rate limiting, circuit breaking, gradual traffic shifting, caching, and auditing on top — it's the layer you grow into once "model versions start multiplying" and direct-to-service access no longer scales.

Why it's worth doing: the moment you have several models and frequent updates, the questions arrive: how do you let 5% of traffic try the new model without affecting the rest? If the new version underperforms, how do you switch back within 10 seconds? Several teams each hang their own model service — who owns the unified entry point? The answer is a gateway. This article presents two concrete implementations — Nginx (weighted routing + gradual rollout) and Envoy (advanced routing) — then covers the complete canary release and A/B experiment workflow.

1. What a Model Gateway Does ​

ResponsibilityDescription
RoutingSend requests to different models or versions by path/header
Rate limitingProtect downstream model services from being crushed by traffic spikes (token bucket / QPS caps)
Circuit breakingFail fast when the downstream keeps failing, instead of piling on
Canary / gradual rolloutShift traffic to a new version by weight or by user
CachingReuse results for identical inputs (e.g. face retrieval, templated queries)
AuditingLog caller, model version, latency, and result for A/B analysis and reconciliation
AuthValidate API keys and control who can call which model

Conclusion first: small teams start with Nginx (config as code, zero maintenance); mid-to-large teams use Envoy (programmable, observable, canary-capable); when you need "model-level" semantics (versions, automatic rollback, multiple frameworks), use a dedicated inference platform like KServe. For the full decision tree, see Choosing Frameworks and Platforms.

2. Comparing the Options ​

OptionStrengthsWeaknessesBest for
Simple Nginx routingZero dependencies, intuitive configNo built-in circuit breaking/retries; canary by weight only2-3 models, quick launch
Advanced Envoy routingHeader/weight/shadow-traffic routing, built-in circuit breaking and retries, xDS dynamic updatesSteep learning curve, complex config DSLMicroservices, multiple teams, K8s-native
Application-level gateway (custom/KServe)Understands model semantics: version management, autoscaling, batch entry pointHigh cost to build; KServe ties you to K8sTeams whose product is the model

3. Walkthrough 1: Routing Between Two Model Versions (Nginx) ​

Scenario: the CTR model is upgrading from v1 to v2, and 10% of traffic should go to v2 first.

nginx
# /etc/nginx/conf.d/model-gateway.conf
upstream ctr_v1 {
    server 10.0.1.10:8000;   # old model service (FastAPI/Triton, say)
}

upstream ctr_v2 {
    server 10.0.1.11:8000;   # new model service
}

server {
    listen 8080;

    # Header-based canary: requests carrying x-canary: v2 are forced to the new version (internal testing / specific user groups)
    location = /predict {
        if ($http_x_canary = "v2") {
            proxy_pass http://ctr_v2;
        }
        proxy_pass http://ctr_v1;
    }
}

For weighted rollouts, Nginx's split_clients hashes a request field to produce a stable percentage split — the same user always lands on the same version, which matters for experiment consistency:

nginx
# Hash by user ID: 10% of user_id hash values land on v2
split_clients "${http_x_user_id}" $ctr_backend {
    10%    ctr_v2;
    *      ctr_v1;
}

location = /predict {
    proxy_pass http://$ctr_backend;
}

Why hash by user instead of round-robining by request count

Both canary releases and A/B tests require that the same user sees consistent results (no drift). Round-robining by request count sends the same user's previous request to v1 and the next one to v2, and the business side ends up confused about "what the new model actually does". So split traffic either by user hash or by cookie/header.

4. Walkthrough 2: The Canary Release Workflow (5% → 20% → 50% → 100%) ​

The goal of a canary is to ramp up the new version gradually and stop at the first sign of degradation. The standard workflow:

  1. Deploy the new version with zero traffic: v2 comes up and passes a smoke test (correctness + latency);
  2. 5% of traffic: run for 30-60 minutes, watching core metrics v1 vs v2;
  3. 20% → 50%: observe each step for at least 30 minutes; continue only if metrics hold;
  4. 100% full rollout: after full rollout, keep observing for a while, then retire the old version.

Metrics to watch while shifting traffic (definitions must be identical for v1/v2, or comparison is meaningless):

MetricWhat to check
P99 latencyThe new version must not be slower than the old beyond a set threshold (e.g. +20%)
Error rate5xx/timeout ratio; stop on any degradation
Business outcomee.g. click-through rate, conversion — the canary's final judge
ResourcesGPU utilization, GPU memory, CPU — confirm the new model can carry the load

Changing weights is a one-line edit in Nginx (10% → 20%), or a route-config change in Envoy (see below). For the complete release methodology, see Release Strategies: Canary and Rollback.

5. Walkthrough 3: Traffic-Based Canary and A/B with Envoy ​

Envoy's strength is built-in routing by header, weight, and shadow traffic. The config below routes 10% of traffic to v2 and additionally supports an x-ab: b header that forces traffic to v2 (for A/B experiment cohorts):

yaml
# envoy.yaml (excerpt)
static_resources:
  clusters:
    - name: ctr_v1
      connect_timeout: 0.25s
      type: STRICT_DNS
      lb_policy: ROUND_ROBIN
      load_assignment:
        cluster_name: ctr_v1
        endpoints:
          - lb_endpoints:
              - endpoint: { address: { socket_address: { address: ctr-v1, port_value: 8000 } } }
    - name: ctr_v2
      connect_timeout: 0.25s
      type: STRICT_DNS
      lb_policy: ROUND_ROBIN
      load_assignment:
        cluster_name: ctr_v2
        endpoints:
          - lb_endpoints:
              - endpoint: { address: { socket_address: { address: ctr-v2, port_value: 8000 } } }

http_filters:
  - name: envoy.filters.http.router
    typed_config:
      "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router

# Route rules (in practice these live in the VirtualHost)
route_config:
  virtual_hosts:
    - name: model_api
      domains: ["*"]
      routes:
        # A/B: requests carrying x-ab: b are forced to v2 (control group)
        - match:
            headers:
              - name: x-ab
                exact_match: "b"
            prefix: /predict
          route: { cluster: ctr_v2 }
        # Canary: all other requests split 90/10 by weight
        - match: { prefix: /predict }
          route:
            weighted_clusters:
              clusters:
                - name: ctr_v1
                  weight: 90
                - name: ctr_v2
                  weight: 10

Envoy's built-in circuit breaking and retries are one config block:

yaml
route:
  cluster: ctr_v1
  retry_policy:
    retry_on: "5xx"
    num_retries: 2
  timeout: 2s

Conclusion first: for production rollouts, prefer Envoy or a K8s gateway (Istio/Argo Rollouts) over hand-written Nginx if blocks — the config updates dynamically (xDS), and you get circuit breaking, retries, and metrics integration. Nginx fits the "get it done in 5 minutes" temporary canary.

Rate Limiting: The Gateway's Second Lifesaver ​

Downstream model capacity is fixed; gateway rate limiting keeps a burst from crushing it. Nginx token-bucket limiting:

nginx
# At most 10 requests per second per IP, bursts up to 20; over-limit returns 429
limit_req_zone $binary_remote_addr zone=model_api:10m rate=10r/s;

location = /predict {
    limit_req zone=model_api burst=20 nodelay;
    limit_req_status 429;
    proxy_pass http://$ctr_backend;
}

Envoy's local rate limiting (envoy.filters.http.local_ratelimit) works the same way and can also limit by header dimensions like x-user-id, letting you set "N QPS per model" against downstream capacity. How to pick the numbers: measure per-instance capacity with the methods in Load Testing and Capacity Planning, then keep a 20% margin.

Circuit Breaking: Protecting Yourself and the Downstream ​

Default circuit-breaking parameters, conclusion first: failure threshold of 5-10 consecutive failures, break window of 10-30 seconds, fail fast during the break (return 503). Envoy config:

yaml
clusters:
  - name: ctr_v1
    circuit_breakers:
      thresholds:
        - priority: DEFAULT
          max_connections: 1024
          max_pending_requests: 100     # queue cap
          max_requests: 2000
          max_retries: 3

Circuit breaking vs rate limiting in one line: rate limiting blocks "too much coming in"; circuit breaking blocks "nothing getting through" — when the downstream is already failing, circuit breaking stops you waiting around and fails fast so upstream can degrade. Combined with the cascading degradation described in Online Inference for Recommender Systems, the whole chain becomes resilient.

trace_id: The Foundation for Rollout and A/B Reconciliation ​

After shifting traffic you must be able to answer "which model version handled this request", so the gateway must generate and propagate a trace_id and write the model_version into the response headers:

nginx
# Put the model version in Nginx response headers so the business side can reconcile afterward
add_header X-Model-Version "v1" always;
add_header X-Trace-Id $request_id always;

The trace_id flows through gateway → model service → logs → analytics events; A/B analysis joins business outcomes on it. For the implementation details of this tracing chain, see Observability in Practice.

6. A/B Experiment Design: Splitting, Instrumentation, Significance, Duration ​

A canary is "outcome defense" (the new version must not be worse); an A/B test is "outcome proof" (is the new version significantly better). The full design has four steps:

  1. Split: 50/50 experiment/control (or per your traffic budget), using user-ID hashing for stability (same user, same group) — the x-ab scheme from Section 5 or the gateway's built-in splitting;
  2. Instrument: the gateway stamps model_version (v1/v2) and trace_id into responses, and the business side joins model versions against business outcomes (clicks/conversions). The version must come back with the response, or reconciliation fails afterward;
  3. Significance testing: use a two-sample proportion test / chi-squared test (conversion-type metrics) or a t-test (numeric metrics); only conclude when p < 0.05, and watch the effect size (how many percentage points);
  4. Duration: cover at least one full business cycle (a week including the weekend, for example). Estimate the sample size in advance with a power analysis — the classic mistake is concluding after 2 days with an undersized sample and no significant difference to show for it.

A/B splitting, instrumentation, and result review all ultimately land on the metric system in Monitoring and Observability; otherwise the data won't line up and experiment conclusions can't be trusted.

7. Rollback: One Switch Back to Safety ​

When the canary shows degradation, the rollback action is a single move: change the weight from 10% back to 0%; the v2 service can stay up for investigation.

nginx
# Rollback: v2 weight to zero
split_clients "${http_x_user_id}" $ctr_backend {
    0%     ctr_v2;
    *      ctr_v1;
}

Key points:

  • Don't delete the service when rolling back — just cut the traffic. Keeping v2 around to pull logs and inspect errors is much faster than rebuilding the environment;
  • Set an automatic rollback threshold at the gateway: if the new version's error rate > 1% or P99 exceeds the threshold for N consecutive minutes, switch back automatically (Envoy/Istio or a homegrown script can do this);
  • If a regression appears after full rollout (covering traffic patterns the canary didn't), cut back the same way and run a postmortem. For the complete rollback playbook, see Release Strategies: Canary and Rollback.

Common Pitfalls and Troubleshooting ​

PitfallSymptomFix
Inconsistent metric definitionsv1/v2 metrics can't be comparedUnify metric definitions and instrumentation fields — see Section 6
Cache pollutionGateway cache mixes old and new version resultsInclude the model version in the cache key; disable or segregate caching for the new version during rollout
Non-sticky splittingThe same user bounces between two versionsHash by user ID (split_clients), don't round-robin by request
Weight change not taking effectNginx edited but not reloadednginx -s reload; Envoy needs an xDS push or a config reload
Undersized experiment sampleDeclaring conclusions on a non-significant p-valueCompute the sample size first, run the full cycle, then conclude
Gateway timeout shorter than model latencySlow model requests get 504 from the gatewayThe gateway timeout must exceed the model's P99 latency
Canary checks latency but not outcomesLatency is fine but business metrics dropInstrument business metrics for both canary and A/B; the gateway gives you traffic, not conclusions

Further Reading ​

References ​