Appearance
Choosing a deployment stack means finding the most cost-effective option across three tiers — a self-built serving layer, dedicated inference frameworks, and cloud-hosted platforms — based on model type, latency requirements, team size, and ops capability. Why does this choice deserve serious effort? Because picking the wrong tier usually costs you a rewrite a few months later: a small-model team grinding through KServe's Kubernetes Operator, or a high-traffic service running FastAPI bare with none of Triton's dynamic batching, are both spending money in the wrong place. This page lays out a decision framework: answer five questions first, narrow the field with four comparison tables, then land on a scenario → solution recommendation matrix. For the conceptual background on deployment modes, see Deployment Patterns; for how frameworks work under the hood, see Model Serving.
The First Rule of Selection
Define the problem first, then pick the tool. A stack isn't better for being heavier or lighter — it's better when it exactly solves your problem at a cost your team can sustain.
1. Five Questions to Answer Before Choosing
Your answers determine which tier you should be shopping in. In priority order:
| # | Question | Why it matters | Where typical answers point |
|---|---|---|---|
| 1 | Model type? CV / traditional ML / LLM | LLMs require a continuous batching engine (vLLM); traditional models are fine with ONNX Runtime | LLM → engine tier; otherwise → serving tier |
| 2 | Latency requirement? What's your P99 budget | Single-digit ms needs an engine plus dedicated hardware; ~100ms is fine with FastAPI | Extreme → Triton/TensorRT; relaxed → lightweight framework |
| 3 | Team size and skills? Dedicated ops or not | A 2-person team can't sustain Kubernetes; a 10-person team actually needs an orchestration platform | Small team → platform / lightweight framework; large team → self-built orchestration |
| 4 | Cloud environment? On-prem / single cloud / multi-cloud | Managed platforms lock you in; on-prem has hardware constraints | Multi-cloud → self-built / KServe; single cloud → platform |
| 5 | Ops capability? Can you staff 24/7 on-call | Unattended → managed platform (the vendor owns the SLO) | No ops → managed SaaS |
The Most Common Selection Mistake
Letting team size substitute for all the other judgment calls. A big team ≠ build it yourself — with only two models and steady traffic, KServe's K8s complexity is a liability. A small team ≠ go managed — if the model is an LLM and latency-sensitive, managed platforms can actually be pricier and harder to tune. Read all five questions together.
2. The Three Tiers: Self-Built, Framework, Platform
text
┌──────────────────────────────────────────────────────────────┐
│ Cloud-hosted platforms (SageMaker / Vertex / Ark / PAI) │ ← buy a service; the vendor owns the SLO
├──────────────────────────────────────────────────────────────┤
│ Inference frameworks (Triton / vLLM / BentoML / Ray Serve) │ ← buy components; you do the orchestration
├──────────────────────────────────────────────────────────────┤
│ Self-built serving (FastAPI + Docker + your own ops) │ ← everything is on you
└──────────────────────────────────────────────────────────────┘
Complexity / cost ▲ Control / flexibility ▲| Dimension | Self-built (FastAPI + Docker) | Inference framework (Triton/vLLM etc.) | Cloud platform (SageMaker etc.) |
|---|---|---|---|
| Time to first service | Fast (a service in a day) | Medium (1-3 days) | Fastest (guided setup) |
| Performance ceiling | Low (no batching or engine optimizations) | High (dynamic batching, operator optimizations) | Depends on the underlying framework |
| Control | Total | High, but bounded by the framework | Low — the platform calls the shots |
| Ops cost | Highest (you own everything) | Medium (you handle deployment and scaling) | Lowest (managed) |
| Flexibility | Highest (write whatever you want) | Medium (multi-model / multi-framework) | Low (boxed in by the platform's API) |
| Best for | Prototypes, internal tools, low traffic | Production, high throughput, many models | Teams with no ops, fast time-to-market |
3. Framework Head-to-Head: Seven Mainstream Options
| Framework | Positioning | Key features | Learning curve | Best for |
|---|---|---|---|---|
| FastAPI | Python web framework | Async, type validation, broad ecosystem; not an inference engine itself | Low | Single model, lots of custom logic, leverages the team's web skills |
| TorchServe | Official PyTorch serving | Model management, batching, torch hub integration | Medium | Inside the PyTorch ecosystem, wanting official support |
| BentoML | Model packaging + serving framework | Standardized Bento packaging, multiple runtimes, cloud-native deployment | Low-medium | Fast delivery, one-click cloud deploys, teams wanting uniform packaging |
| Ray Serve | Distributed serving layer | Elastic scaling, multi-model, integrates with the Ray ecosystem | Medium-high | Existing Ray clusters, complex pipelines |
| Triton | NVIDIA inference server | Multi-framework backends, dynamic batching, concurrent model instances, ensembles | Medium | High GPU throughput, mixed models, production-grade |
| KServe | Inference platform on K8s | Serverless, autoscaling, canary releases, multi-framework | High (needs K8s) | Teams already on K8s that want to platformize |
| Seldon Core | K8s inference orchestration | Canary releases, monitoring, explainability integration | High | K8s teams that want advanced release capabilities |
What the table is telling you:
- FastAPI is the foundation, not the rival: most other frameworks are HTTP services internally anyway. FastAPI fits the "self-built + small model" path; once you chase high throughput, move the model into Triton and FastAPI demotes itself to a front gateway;
- Triton's killer feature is dynamic batching: it merges scattered requests into batches, which can multiply GPU throughput 2-5x — something a bare FastAPI service simply cannot do;
- BentoML's value is standardized packaging: model + dependencies + service packaged once into a Bento, then exported to various deployment targets — ideal when a team needs to ship many models quickly;
- KServe vs Seldon: both ride on K8s. KServe leans serverless + autoscaling; Seldon leans advanced release management and explainability. Choose based on how mature your K8s operations are.
4. Cloud Platform Comparison: Managed Serving Across the Clouds
| Platform | Features | Billing | Ease of adoption | Best for |
|---|---|---|---|---|
| AWS SageMaker | End-to-end (training/deployment/monitoring), multi-framework, autoscaling endpoints | Billed per endpoint instance-hour; idle capacity gets expensive | Medium | Already on AWS, want fully managed |
| Azure ML | Deep Azure ecosystem integration, MLOps toolchain | Billed per compute/endpoint | Medium | Already on Azure, using Azure DevOps |
| GCP Vertex AI | Unified MLOps, Vertex AI Endpoints, BigQuery integration | Pay-as-you-go, decent free tier | Medium | Already on GCP, data in BigQuery |
| Volcengine Ark | LLM inference and fine-tuning platform, China compliance | Billed by tokens/resources | Low | China-based business, LLM workloads, compliance requirements |
| Alibaba Cloud PAI | Integrated training + inference, EAS online serving, domestic hardware support | Resource packs / pay-as-you-go | Medium | China-based business, already on Alibaba Cloud |
Vendor Lock-In Is a Real Cost
Every convenience a managed platform offers is backed by "you will rewrite against our API." Assess lock-in before you launch: model artifact formats (SageMaker uses its own model packaging), scaling APIs, monitoring integrations. A simple litmus test: if you had to migrate to another cloud a year from now, how many person-weeks would it take? If the answer is more than two, seriously consider a platform-neutral option like KServe.
5. Recommendation Matrix: Scenario → Solution
Combine your five answers into a typical scenario and read straight off the table:
| Scenario | Recommended | Why | Alternative |
|---|---|---|---|
| Ship a small model fast, team of 2-3 | BentoML or FastAPI + Docker | Quick to start, good enough, no ops burden | Managed endpoint on a cloud platform |
| Extreme performance, multi-model GPU sharing | Triton | Dynamic batching + multi-framework backends, highest throughput ceiling | Self-built + Triton backend |
| LLM serving | vLLM (or a managed platform) | Continuous batching / PagedAttention is the key to LLM throughput — see the vLLM case study | Volcengine Ark / cloud LLM APIs |
| No ops team, business on a single cloud | Cloud managed platform (SageMaker/Vertex etc.) | Scaling, SLO, and monitoring all managed | In China: Volcengine Ark / Alibaba Cloud PAI |
| Already on K8s, want a platform and canaries | KServe | Serverless scaling + native canary releases | Seldon Core |
| Batch / offline inference | Cloud batch processing (e.g. SageMaker Batch / self-built pipeline) | On-demand compute, no idle cost — see the batch pipeline case study | Self-built Celery/Argo |
| Stateless, spiky traffic | Serverless inference | Zero idle footprint, trade cold starts for elasticity — see Serverless Inference | Cloud platform serverless endpoints |
| Multi-model routing + A/B canary | Gateway + framework (e.g. gateway canary releases) | Routing, traffic splitting, and observability unified at the gateway layer | Built-in platform traffic splitting |
6. Migration Cost and Lock-In
Selection isn't just about "works well today" — price out three years of total cost. Lock-in comes in three layers:
- Framework lock-in: your code depends on Triton's Python backend or BentoML's packaging format → migration means rewriting the inference code;
- Platform lock-in: you use SageMaker's pipeline API or Vertex's endpoint orchestration → migration is a full rebuild;
- Data/monitoring lock-in: metrics flow into the platform's own monitoring and logs into cloud-native storage → historical observability is lost when you exit.
General strategies to soften lock-in:
- Make your inference wrapper the single dependency point: serving code depends only on your own
inference.py(see Deploy a Model from Scratch); switching engines or platforms touches one file; - Prefer standard formats: export models to ONNX (Model Formats); ONNX is the most portable intermediate format across engines and platforms;
- Standardize on Prometheus for monitoring: the Prometheus format is the de facto standard, almost every platform supports exporting it, and metrics survive a platform move.
7. The Decision Flow as a Tree
text
Start
│
├─ Is the model an LLM? ──yes──▶ Latency-sensitive & need control? ──yes──▶ vLLM + (KServe / self-built)
│ │ │ └─no──▶ Managed LLM platform (Ark / cloud API)
│ └─no
│
├─ P99 target < 20ms? ──yes──▶ GPU + engine (Triton / TensorRT), see the TensorRT case study
│
├─ Team can operate K8s? ──yes──▶ Multi-model / platform ambitions? ──yes──▶ KServe / Seldon
│ │ └─no──▶ Single model → BentoML / TorchServe
│ └─no
│
├─ Already on a cloud? ──yes──▶ Use that cloud's managed platform (SageMaker / Vertex / PAI / Ark)
│ └─no
│
└─ Self-build → FastAPI + Docker (start from the build-your-own skeleton)Every "choose this" on the decision tree can be traced back to the five questions from Section 1, which keeps the decision auditable.
Trade-Offs
Selection never comes with a free lunch. Four trade-offs are unavoidable: control ↔ ops cost (self-built is the most flexible and the most tiring); performance ↔ complexity (Triton's batching gains come at the price of its API); fast start ↔ low ceiling (FastAPI gets you going quickly, but high throughput means learning an engine); managed convenience ↔ lock-in risk. And don't forget: selection is not a one-shot decision. The right path for many teams is "start on FastAPI → switch to Triton as traffic grows → adopt KServe at scale," and each upgrade has a clear trigger (P99 degradation, a throughput plateau, instance counts spiraling). Writing those upgrade triggers into your monitoring alerts is far more practical than trying to pick the "ultimate solution" in one shot.
Checklist
- [ ] Write down answers to the five questions (model / latency / team / cloud / ops) one by one;
- [ ] Mark at most 2 candidates in the comparison tables that fit your constraints;
- [ ] Confirm the plan against the recommendation matrix, and write down "why not the alternative";
- [ ] Have assessed lock-in cost (person-time to migrate);
- [ ] Have defined upgrade triggers (which metric degrading moves you to the next tier);
- [ ] A standard model format (ONNX) and a wrapped inference class are in place to keep you portable.
Further Reading
- Tools and Resources — official entry points and community resources for each framework
- FastAPI + Docker Online Serving — the complete self-built walkthrough
- NVIDIA Triton Multi-Model Serving — production practice on the engine path
- vLLM LLM Inference Service — deployment and tuning for LLM workloads
- Deployment Patterns — selection background for serverless/edge/batch
- Model Serving — how serving frameworks and engine tiers differ under the hood
References
- FastAPI docs: https://fastapi.tiangolo.com/
- TorchServe: https://pytorch.org/serve/
- BentoML docs: https://docs.bentoml.com/
- NVIDIA Triton Inference Server: https://github.com/triton-inference-server/server
- KServe docs: https://kserve.github.io/website/
- Seldon Core: https://www.seldon.io/
- AWS SageMaker: https://aws.amazon.com/sagemaker/
- Volcengine Ark: https://www.volcengine.com/product/ark
- Alibaba Cloud PAI: https://www.aliyun.com/product/bigdata/learn