Skip to content

Choosing Frameworks and Platforms

At a glance FastAPI, BentoML, Triton, KServe, SageMaker, Volcengine Ark... how do you pick a deployment stack? This guide gives a decision framework based on team size, latency requirements, and model type, plus a scenario-to-solution recommendation matrix.

Choosing a deployment stack means finding the most cost-effective option across three tiers — a self-built serving layer, dedicated inference frameworks, and cloud-hosted platforms — based on model type, latency requirements, team size, and ops capability. Why does this choice deserve serious effort? Because picking the wrong tier usually costs you a rewrite a few months later: a small-model team grinding through KServe's Kubernetes Operator, or a high-traffic service running FastAPI bare with none of Triton's dynamic batching, are both spending money in the wrong place. This page lays out a decision framework: answer five questions first, narrow the field with four comparison tables, then land on a scenario → solution recommendation matrix. For the conceptual background on deployment modes, see Deployment Patterns; for how frameworks work under the hood, see Model Serving.

The First Rule of Selection

Define the problem first, then pick the tool. A stack isn't better for being heavier or lighter — it's better when it exactly solves your problem at a cost your team can sustain.

1. Five Questions to Answer Before Choosing ​

Your answers determine which tier you should be shopping in. In priority order:

#QuestionWhy it mattersWhere typical answers point
1Model type? CV / traditional ML / LLMLLMs require a continuous batching engine (vLLM); traditional models are fine with ONNX RuntimeLLM → engine tier; otherwise → serving tier
2Latency requirement? What's your P99 budgetSingle-digit ms needs an engine plus dedicated hardware; ~100ms is fine with FastAPIExtreme → Triton/TensorRT; relaxed → lightweight framework
3Team size and skills? Dedicated ops or notA 2-person team can't sustain Kubernetes; a 10-person team actually needs an orchestration platformSmall team → platform / lightweight framework; large team → self-built orchestration
4Cloud environment? On-prem / single cloud / multi-cloudManaged platforms lock you in; on-prem has hardware constraintsMulti-cloud → self-built / KServe; single cloud → platform
5Ops capability? Can you staff 24/7 on-callUnattended → managed platform (the vendor owns the SLO)No ops → managed SaaS

The Most Common Selection Mistake

Letting team size substitute for all the other judgment calls. A big team ≠ build it yourself — with only two models and steady traffic, KServe's K8s complexity is a liability. A small team ≠ go managed — if the model is an LLM and latency-sensitive, managed platforms can actually be pricier and harder to tune. Read all five questions together.

2. The Three Tiers: Self-Built, Framework, Platform ​

text
┌──────────────────────────────────────────────────────────────┐
│ Cloud-hosted platforms (SageMaker / Vertex / Ark / PAI)      │  ← buy a service; the vendor owns the SLO
├──────────────────────────────────────────────────────────────┤
│ Inference frameworks (Triton / vLLM / BentoML / Ray Serve)   │  ← buy components; you do the orchestration
├──────────────────────────────────────────────────────────────┤
│ Self-built serving (FastAPI + Docker + your own ops)         │  ← everything is on you
└──────────────────────────────────────────────────────────────┘
   Complexity / cost ▲                                Control / flexibility ▲
DimensionSelf-built (FastAPI + Docker)Inference framework (Triton/vLLM etc.)Cloud platform (SageMaker etc.)
Time to first serviceFast (a service in a day)Medium (1-3 days)Fastest (guided setup)
Performance ceilingLow (no batching or engine optimizations)High (dynamic batching, operator optimizations)Depends on the underlying framework
ControlTotalHigh, but bounded by the frameworkLow — the platform calls the shots
Ops costHighest (you own everything)Medium (you handle deployment and scaling)Lowest (managed)
FlexibilityHighest (write whatever you want)Medium (multi-model / multi-framework)Low (boxed in by the platform's API)
Best forPrototypes, internal tools, low trafficProduction, high throughput, many modelsTeams with no ops, fast time-to-market

3. Framework Head-to-Head: Seven Mainstream Options ​

FrameworkPositioningKey featuresLearning curveBest for
FastAPIPython web frameworkAsync, type validation, broad ecosystem; not an inference engine itselfLowSingle model, lots of custom logic, leverages the team's web skills
TorchServeOfficial PyTorch servingModel management, batching, torch hub integrationMediumInside the PyTorch ecosystem, wanting official support
BentoMLModel packaging + serving frameworkStandardized Bento packaging, multiple runtimes, cloud-native deploymentLow-mediumFast delivery, one-click cloud deploys, teams wanting uniform packaging
Ray ServeDistributed serving layerElastic scaling, multi-model, integrates with the Ray ecosystemMedium-highExisting Ray clusters, complex pipelines
TritonNVIDIA inference serverMulti-framework backends, dynamic batching, concurrent model instances, ensemblesMediumHigh GPU throughput, mixed models, production-grade
KServeInference platform on K8sServerless, autoscaling, canary releases, multi-frameworkHigh (needs K8s)Teams already on K8s that want to platformize
Seldon CoreK8s inference orchestrationCanary releases, monitoring, explainability integrationHighK8s teams that want advanced release capabilities

What the table is telling you:

  • FastAPI is the foundation, not the rival: most other frameworks are HTTP services internally anyway. FastAPI fits the "self-built + small model" path; once you chase high throughput, move the model into Triton and FastAPI demotes itself to a front gateway;
  • Triton's killer feature is dynamic batching: it merges scattered requests into batches, which can multiply GPU throughput 2-5x — something a bare FastAPI service simply cannot do;
  • BentoML's value is standardized packaging: model + dependencies + service packaged once into a Bento, then exported to various deployment targets — ideal when a team needs to ship many models quickly;
  • KServe vs Seldon: both ride on K8s. KServe leans serverless + autoscaling; Seldon leans advanced release management and explainability. Choose based on how mature your K8s operations are.

4. Cloud Platform Comparison: Managed Serving Across the Clouds ​

PlatformFeaturesBillingEase of adoptionBest for
AWS SageMakerEnd-to-end (training/deployment/monitoring), multi-framework, autoscaling endpointsBilled per endpoint instance-hour; idle capacity gets expensiveMediumAlready on AWS, want fully managed
Azure MLDeep Azure ecosystem integration, MLOps toolchainBilled per compute/endpointMediumAlready on Azure, using Azure DevOps
GCP Vertex AIUnified MLOps, Vertex AI Endpoints, BigQuery integrationPay-as-you-go, decent free tierMediumAlready on GCP, data in BigQuery
Volcengine ArkLLM inference and fine-tuning platform, China complianceBilled by tokens/resourcesLowChina-based business, LLM workloads, compliance requirements
Alibaba Cloud PAIIntegrated training + inference, EAS online serving, domestic hardware supportResource packs / pay-as-you-goMediumChina-based business, already on Alibaba Cloud

Vendor Lock-In Is a Real Cost

Every convenience a managed platform offers is backed by "you will rewrite against our API." Assess lock-in before you launch: model artifact formats (SageMaker uses its own model packaging), scaling APIs, monitoring integrations. A simple litmus test: if you had to migrate to another cloud a year from now, how many person-weeks would it take? If the answer is more than two, seriously consider a platform-neutral option like KServe.

5. Recommendation Matrix: Scenario → Solution ​

Combine your five answers into a typical scenario and read straight off the table:

ScenarioRecommendedWhyAlternative
Ship a small model fast, team of 2-3BentoML or FastAPI + DockerQuick to start, good enough, no ops burdenManaged endpoint on a cloud platform
Extreme performance, multi-model GPU sharingTritonDynamic batching + multi-framework backends, highest throughput ceilingSelf-built + Triton backend
LLM servingvLLM (or a managed platform)Continuous batching / PagedAttention is the key to LLM throughput — see the vLLM case studyVolcengine Ark / cloud LLM APIs
No ops team, business on a single cloudCloud managed platform (SageMaker/Vertex etc.)Scaling, SLO, and monitoring all managedIn China: Volcengine Ark / Alibaba Cloud PAI
Already on K8s, want a platform and canariesKServeServerless scaling + native canary releasesSeldon Core
Batch / offline inferenceCloud batch processing (e.g. SageMaker Batch / self-built pipeline)On-demand compute, no idle cost — see the batch pipeline case studySelf-built Celery/Argo
Stateless, spiky trafficServerless inferenceZero idle footprint, trade cold starts for elasticity — see Serverless InferenceCloud platform serverless endpoints
Multi-model routing + A/B canaryGateway + framework (e.g. gateway canary releases)Routing, traffic splitting, and observability unified at the gateway layerBuilt-in platform traffic splitting

6. Migration Cost and Lock-In ​

Selection isn't just about "works well today" — price out three years of total cost. Lock-in comes in three layers:

  1. Framework lock-in: your code depends on Triton's Python backend or BentoML's packaging format → migration means rewriting the inference code;
  2. Platform lock-in: you use SageMaker's pipeline API or Vertex's endpoint orchestration → migration is a full rebuild;
  3. Data/monitoring lock-in: metrics flow into the platform's own monitoring and logs into cloud-native storage → historical observability is lost when you exit.

General strategies to soften lock-in:

  • Make your inference wrapper the single dependency point: serving code depends only on your own inference.py (see Deploy a Model from Scratch); switching engines or platforms touches one file;
  • Prefer standard formats: export models to ONNX (Model Formats); ONNX is the most portable intermediate format across engines and platforms;
  • Standardize on Prometheus for monitoring: the Prometheus format is the de facto standard, almost every platform supports exporting it, and metrics survive a platform move.

7. The Decision Flow as a Tree ​

text
Start
│
├─ Is the model an LLM? ──yes──▶ Latency-sensitive & need control? ──yes──▶ vLLM + (KServe / self-built)
│   │                           │                                  └─no──▶ Managed LLM platform (Ark / cloud API)
│   └─no
│
├─ P99 target < 20ms? ──yes──▶ GPU + engine (Triton / TensorRT), see the TensorRT case study
│
├─ Team can operate K8s? ──yes──▶ Multi-model / platform ambitions? ──yes──▶ KServe / Seldon
│   │                              └─no──▶ Single model → BentoML / TorchServe
│   └─no
│
├─ Already on a cloud? ──yes──▶ Use that cloud's managed platform (SageMaker / Vertex / PAI / Ark)
│   └─no
│
└─ Self-build → FastAPI + Docker (start from the build-your-own skeleton)

Every "choose this" on the decision tree can be traced back to the five questions from Section 1, which keeps the decision auditable.

Trade-Offs ​

Selection never comes with a free lunch. Four trade-offs are unavoidable: control ↔ ops cost (self-built is the most flexible and the most tiring); performance ↔ complexity (Triton's batching gains come at the price of its API); fast start ↔ low ceiling (FastAPI gets you going quickly, but high throughput means learning an engine); managed convenience ↔ lock-in risk. And don't forget: selection is not a one-shot decision. The right path for many teams is "start on FastAPI → switch to Triton as traffic grows → adopt KServe at scale," and each upgrade has a clear trigger (P99 degradation, a throughput plateau, instance counts spiraling). Writing those upgrade triggers into your monitoring alerts is far more practical than trying to pick the "ultimate solution" in one shot.

Checklist ​

  • [ ] Write down answers to the five questions (model / latency / team / cloud / ops) one by one;
  • [ ] Mark at most 2 candidates in the comparison tables that fit your constraints;
  • [ ] Confirm the plan against the recommendation matrix, and write down "why not the alternative";
  • [ ] Have assessed lock-in cost (person-time to migrate);
  • [ ] Have defined upgrade triggers (which metric degrading moves you to the next tier);
  • [ ] A standard model format (ONNX) and a wrapped inference class are in place to keep you portable.

Further Reading ​

References ​