Appearance
Deployment Architecture Patterns
The one-sentence definition: deployment architecture patterns are the five basic shapes of "when, how, and where model inference runs"—online, batch, streaming, edge, and Serverless, each with its own latency, cost, and operations profile.
Industry insight: choosing the wrong deployment pattern is the most expensive mistake in deployment engineering, and it's hard to fix after the fact. Serving a million-DAU recommendation flow as batch processing leaves user profiles a day stale and the business collapses; hanging a once-a-day report scoring job on an always-on GPU service burns the price of an A100 in a year. Industry experience: describe your traffic pattern and latency requirement in one sentence first, then match a pattern to them—once the pattern is right, every later optimization reinforces a correct skeleton.
1. The Five Patterns at a Glance
| Pattern | Trigger | Latency | Throughput | Cost shape | Typical scenarios | Representative tools |
|---|---|---|---|---|---|---|
| Online inference | Real-time requests | Milliseconds | High (concurrent) | Always-on resources | Recommendations, search, chat, risk control | FastAPI, Triton, vLLM |
| Batch processing | Scheduled/manual | Hours | Extremely high | Pay per run | Profiles, daily reports, offline evaluation | Airflow + Spark/PyTorch |
| Streaming inference | Event-driven | Seconds | Continuous | Always-on (scalable) | Real-time risk control, real-time personalization | Kafka + Flink |
| Edge deployment | On-device trigger | Local milliseconds | Low | One-time hardware | Offline availability, privacy-sensitive | TensorRT, TFLite |
| Serverless | Real-time requests | Seconds (cold start) | Elastic | Pay per invocation | Low-frequency, bursty, internal tools | KServe, AWS Lambda |
2. The Patterns, One by One
2.1 Online Inference: Real-Time APIs
Traits: synchronous request-response, millisecond latency budgets, always-on services for concurrency. Best for latency-sensitive, steady, high-volume traffic (recommendations, risk control, chat).
text
Client ──► Gateway ──► Inference service (multiple always-on replicas) ──► GPU/CPU
▲ dynamic batching / result cachingCost profile: replicas stay up even when overnight traffic goes to zero (unless you configure autoscaling). The optimization focus is Performance Optimization and Capacity Planning and Serving and Inference APIs.
2.2 Batch Processing: Scheduled Offline Scoring
Traits: scheduled jobs, full datasets, hour-level tolerance, takes nothing away from online resources. Fits profiling, daily reports, offline experiment evaluation.
text
Airflow scheduled trigger ──► Spark/DataFrame reads tens of millions of samples
──► batch forward passes (large batches, GPU saturated)
──► results written back to the warehouse/feature tablesKey point: large batches are the core of batch performance—GPU utilization easily hits 90%+, the lowest unit cost of the five patterns. When batch and online share one model, watch "feature parity"—see Batch Inference Pipelines.
2.3 Streaming Inference: Event-Driven
Traits: processes event streams record by record, second-level latency, backpressure and checkpointing keep it from blocking. Fits real-time risk control, real-time personalization, IoT alerting.
text
Kafka event stream ──► Flink operators (with a model inference UDF)
──► decisions/alerts ──► results written back to Kafka/DBVersus online inference: streaming is data coming to you (push), latency requirements are looser (seconds, not milliseconds), but it demands state consistency and exactly-once semantics. Model inference becomes a UDF inside Flink, usually with a small model (to keep per-record latency down).
2.4 Edge Deployment: On-Device Inference
Traits: the model runs on the user's device or on-site hardware—works with no network, data never leaves the device. Fits offline tools, privacy-sensitive scenarios, vehicles/industrial sites. See TensorRT and Edge Deployment for a representative walkthrough.
text
Phone / vehicle / industrial PC
├─ TFLite / Core ML / TensorRT / edge NPU
└─ local inference → local results; optionally upload anonymized samples for federated learningThe price: hardware fragmentation (operator compatibility must be verified for every device), no hot model updates (only via app/firmware releases), and limited compute (small models + quantization are mandatory). What you get in return: zero network latency and strong privacy.
2.5 Serverless: Pay per Invocation
Traits: the platform scales automatically, billing by invocation count/duration, cold-start cost. Fits low-frequency, bursty, internal tools, and experimental services. See Serverless Inference for a hands-on walkthrough.
text
Request ──► KServe/Gateway ──► scale from 0 to 1 (cold start: pull image + load model, seconds to tens of seconds)
──► scale back to 0 when idleCold Start Is Serverless's Worst Enemy
Put a 7B model that takes 30 seconds to load on Serverless and the first request may wait 40 seconds. Mitigations: pre-warmed instances (minimum replicas = 1), model on a shared cache volume—or just don't use Serverless. "Serverless only fits small models that load in seconds" is the rule of thumb.
3. The Selection Decision Tree
text
First ask: what's the latency requirement?
├─ Milliseconds, user-perceivable (recommendations/chat/risk control) → online inference
├─ Seconds, event-driven → streaming inference
├─ Minutes–hours, full datasets → batch processing
└─ No network / strong privacy / on-device → edge deployment
Next ask: what's the traffic pattern?
├─ Steady and high → always-on online + autoscaling
├─ Low-frequency / bursty / erratic → Serverless (small model) or Serverless + pre-warming
└─ Online by day + full batch at night → hybrid deployment (below)
Then ask: where does the data live?
├─ Data in the datacenter → cloud inference
├─ Data on user devices (privacy) → edge inference
└─ Data in a streaming pipeline → streaming inferenceDecision Weight Ordering
text
Latency requirement > traffic pattern > cost constraints > data locationLatency comes before everything: if hours are tolerable, never consider an online service; if milliseconds are mandatory, pay for always-on GPUs without flinching.
4. Hybrid Deployment: Online + Batch Combined
Real systems are almost always hybrid. The classic combination:
text
Daytime: online inference (small batches, low latency) serves real-time business
Night: the same GPUs switch to batch jobs (profiling, model re-evaluation, offline experiments)
or: the online service scales down while batch scales upThe payoff: one set of GPU resources sits idle 0 hours out of 24—online by day, batch by night, overall utilization doubled. Engineering prerequisites:
- Online and batch share the same deployment unit (K8s with scaling policies switching between them);
- Model files and feature definitions are unified (see MLOps Deployment Pipeline);
- Batch jobs are lower priority than online and can be preempted during traffic spikes.
5. Cost Model Comparison
| Pattern | Fixed cost | Variable cost | Idle-waste risk |
|---|---|---|---|
| Online inference | High (always-on GPU/CPU) | Low | High (idle overnight) |
| Batch processing | Low | Per run duration | Low (released when done) |
| Streaming inference | Medium (always-on streaming cluster) | Medium | Medium |
| Edge deployment | One-time hardware | Maintenance | Low |
| Serverless | Low | Per invocation | Extremely low (zero when idle) |
A quick worked example: 10,000 calls a day, 50ms per inference, model fits on a T4 (~1 yuan/hour on demand):
- Always-on: 24×1 ≈ 24 yuan/day, actual utilization < 2%;
- Serverless (per invocation): 10,000 × 50ms ≈ 500 seconds ≈ 0.14 yuan/day, a 170× difference.
But if calls grow to 10 million a day, Serverless's per-invocation fees overtake an always-on monthly plan—that's the reality of "patterns migrate with scale."
Trade-offs
| Decision point | Options | How to choose |
|---|---|---|
| Real-time-ness | Online/streaming vs batch | The latency requirement decides |
| Cost | Always-on vs Serverless | Serverless for low-frequency bursts; always-on for steady high volume |
| Data location | Cloud vs device | Privacy and offline needs come first |
| Operational complexity | Always-on vs Serverless vs batch | The more managed, the less to operate—but the less customizable |
| Resource utilization | Single pattern vs hybrid | Hybrid if utilization is the priority |
One-line summary: the pattern is the first decision in deployment—match your latency, traffic, cost, and data location to a pattern in four steps, then combine patterns to keep resources fully utilized.
Further Reading
- Serving and Inference APIs — the service engineering foundation for the online pattern
- Batch Inference Pipelines — a complete batch processing case study
- Serverless Inference — cold starts and elastic scaling in practice
- TensorRT and Edge Deployment — compilation and deployment for the edge pattern
- Performance Optimization and Capacity Planning — load testing and capacity math for the online pattern