Skip to content

Deployment Architecture Patterns

At a glance Online, batch, streaming, edge, Serverless—five deployment architectures, each with its own sweet spot. This article provides a full comparison table, a selection decision tree, and a discussion of hybrid deployments.

Deployment Architecture Patterns ​

The one-sentence definition: deployment architecture patterns are the five basic shapes of "when, how, and where model inference runs"—online, batch, streaming, edge, and Serverless, each with its own latency, cost, and operations profile.

Industry insight: choosing the wrong deployment pattern is the most expensive mistake in deployment engineering, and it's hard to fix after the fact. Serving a million-DAU recommendation flow as batch processing leaves user profiles a day stale and the business collapses; hanging a once-a-day report scoring job on an always-on GPU service burns the price of an A100 in a year. Industry experience: describe your traffic pattern and latency requirement in one sentence first, then match a pattern to them—once the pattern is right, every later optimization reinforces a correct skeleton.

1. The Five Patterns at a Glance ​

PatternTriggerLatencyThroughputCost shapeTypical scenariosRepresentative tools
Online inferenceReal-time requestsMillisecondsHigh (concurrent)Always-on resourcesRecommendations, search, chat, risk controlFastAPI, Triton, vLLM
Batch processingScheduled/manualHoursExtremely highPay per runProfiles, daily reports, offline evaluationAirflow + Spark/PyTorch
Streaming inferenceEvent-drivenSecondsContinuousAlways-on (scalable)Real-time risk control, real-time personalizationKafka + Flink
Edge deploymentOn-device triggerLocal millisecondsLowOne-time hardwareOffline availability, privacy-sensitiveTensorRT, TFLite
ServerlessReal-time requestsSeconds (cold start)ElasticPay per invocationLow-frequency, bursty, internal toolsKServe, AWS Lambda

2. The Patterns, One by One ​

2.1 Online Inference: Real-Time APIs ​

Traits: synchronous request-response, millisecond latency budgets, always-on services for concurrency. Best for latency-sensitive, steady, high-volume traffic (recommendations, risk control, chat).

text
Client ──► Gateway ──► Inference service (multiple always-on replicas) ──► GPU/CPU
                        ▲ dynamic batching / result caching

Cost profile: replicas stay up even when overnight traffic goes to zero (unless you configure autoscaling). The optimization focus is Performance Optimization and Capacity Planning and Serving and Inference APIs.

2.2 Batch Processing: Scheduled Offline Scoring ​

Traits: scheduled jobs, full datasets, hour-level tolerance, takes nothing away from online resources. Fits profiling, daily reports, offline experiment evaluation.

text
Airflow scheduled trigger ──► Spark/DataFrame reads tens of millions of samples
                       ──► batch forward passes (large batches, GPU saturated)
                       ──► results written back to the warehouse/feature tables

Key point: large batches are the core of batch performance—GPU utilization easily hits 90%+, the lowest unit cost of the five patterns. When batch and online share one model, watch "feature parity"—see Batch Inference Pipelines.

2.3 Streaming Inference: Event-Driven ​

Traits: processes event streams record by record, second-level latency, backpressure and checkpointing keep it from blocking. Fits real-time risk control, real-time personalization, IoT alerting.

text
Kafka event stream ──► Flink operators (with a model inference UDF)
                  ──► decisions/alerts ──► results written back to Kafka/DB

Versus online inference: streaming is data coming to you (push), latency requirements are looser (seconds, not milliseconds), but it demands state consistency and exactly-once semantics. Model inference becomes a UDF inside Flink, usually with a small model (to keep per-record latency down).

2.4 Edge Deployment: On-Device Inference ​

Traits: the model runs on the user's device or on-site hardware—works with no network, data never leaves the device. Fits offline tools, privacy-sensitive scenarios, vehicles/industrial sites. See TensorRT and Edge Deployment for a representative walkthrough.

text
Phone / vehicle / industrial PC
 ├─ TFLite / Core ML / TensorRT / edge NPU
 └─ local inference → local results; optionally upload anonymized samples for federated learning

The price: hardware fragmentation (operator compatibility must be verified for every device), no hot model updates (only via app/firmware releases), and limited compute (small models + quantization are mandatory). What you get in return: zero network latency and strong privacy.

2.5 Serverless: Pay per Invocation ​

Traits: the platform scales automatically, billing by invocation count/duration, cold-start cost. Fits low-frequency, bursty, internal tools, and experimental services. See Serverless Inference for a hands-on walkthrough.

text
Request ──► KServe/Gateway ──► scale from 0 to 1 (cold start: pull image + load model, seconds to tens of seconds)
                        ──► scale back to 0 when idle

Cold Start Is Serverless's Worst Enemy

Put a 7B model that takes 30 seconds to load on Serverless and the first request may wait 40 seconds. Mitigations: pre-warmed instances (minimum replicas = 1), model on a shared cache volume—or just don't use Serverless. "Serverless only fits small models that load in seconds" is the rule of thumb.

3. The Selection Decision Tree ​

text
First ask: what's the latency requirement?
 ├─ Milliseconds, user-perceivable (recommendations/chat/risk control) → online inference
 ├─ Seconds, event-driven                                              → streaming inference
 ├─ Minutes–hours, full datasets                                       → batch processing
 └─ No network / strong privacy / on-device                            → edge deployment

Next ask: what's the traffic pattern?
 ├─ Steady and high → always-on online + autoscaling
 ├─ Low-frequency / bursty / erratic → Serverless (small model) or Serverless + pre-warming
 └─ Online by day + full batch at night → hybrid deployment (below)

Then ask: where does the data live?
 ├─ Data in the datacenter → cloud inference
 ├─ Data on user devices (privacy) → edge inference
 └─ Data in a streaming pipeline → streaming inference

Decision Weight Ordering ​

text
Latency requirement > traffic pattern > cost constraints > data location

Latency comes before everything: if hours are tolerable, never consider an online service; if milliseconds are mandatory, pay for always-on GPUs without flinching.

4. Hybrid Deployment: Online + Batch Combined ​

Real systems are almost always hybrid. The classic combination:

text
Daytime: online inference (small batches, low latency) serves real-time business
Night:   the same GPUs switch to batch jobs (profiling, model re-evaluation, offline experiments)
             or: the online service scales down while batch scales up

The payoff: one set of GPU resources sits idle 0 hours out of 24—online by day, batch by night, overall utilization doubled. Engineering prerequisites:

  1. Online and batch share the same deployment unit (K8s with scaling policies switching between them);
  2. Model files and feature definitions are unified (see MLOps Deployment Pipeline);
  3. Batch jobs are lower priority than online and can be preempted during traffic spikes.

5. Cost Model Comparison ​

PatternFixed costVariable costIdle-waste risk
Online inferenceHigh (always-on GPU/CPU)LowHigh (idle overnight)
Batch processingLowPer run durationLow (released when done)
Streaming inferenceMedium (always-on streaming cluster)MediumMedium
Edge deploymentOne-time hardwareMaintenanceLow
ServerlessLowPer invocationExtremely low (zero when idle)

A quick worked example: 10,000 calls a day, 50ms per inference, model fits on a T4 (~1 yuan/hour on demand):

  • Always-on: 24×1 ≈ 24 yuan/day, actual utilization < 2%;
  • Serverless (per invocation): 10,000 × 50ms ≈ 500 seconds ≈ 0.14 yuan/day, a 170× difference.

But if calls grow to 10 million a day, Serverless's per-invocation fees overtake an always-on monthly plan—that's the reality of "patterns migrate with scale."

Trade-offs ​

Decision pointOptionsHow to choose
Real-time-nessOnline/streaming vs batchThe latency requirement decides
CostAlways-on vs ServerlessServerless for low-frequency bursts; always-on for steady high volume
Data locationCloud vs devicePrivacy and offline needs come first
Operational complexityAlways-on vs Serverless vs batchThe more managed, the less to operate—but the less customizable
Resource utilizationSingle pattern vs hybridHybrid if utilization is the priority

One-line summary: the pattern is the first decision in deployment—match your latency, traffic, cost, and data location to a pattern in four steps, then combine patterns to keep resources fully utilized.

Further Reading ​

References ​