Skip to content

Inference Engine Comparison

At a glance How do you choose among vLLM, SGLang, TensorRT-LLM, TGI, Triton, ONNX Runtime, OpenVINO, llama.cpp, and TFLite? Starting from five dimensions — model type, hardware, deployment form, performance target, and team capability — this article gives a nine-engine comparison table and selection decisions for three major scenarios: LLM online serving, classic model serving, and on-device.

Inference Engine Comparison ​

Choosing an inference engine is not about picking "the most popular" — it is about picking "the one that fits the current problem, the hardware, and the team." Choose wrong and you are reworking the integration three weeks after launch; choose right and both performance and operational cost halve.

In the LLM era, the number of inference engines exploded from 2 (TensorRT, ONNX Runtime) to more than 10, and every one claims to be "the fastest" — the vLLM paper reports 24× over HuggingFace, TensorRT-LLM tops NVIDIA's leaderboards in throughput, SGLang crushes the rest at complex scheduling, llama.cpp is unmatched on-device... These numbers are all true, and none of them should drive your selection. The reason is simple: every number was measured in the scenario that engine is best at, and its transferability to your scenario is close to zero.

This article provides a systematic framework for inference engine selection: ask 5 questions first, then compare 9 engines, and finally decide by 3 scenarios. After reading it, you should be able to produce a "which engine, why, and what the risks are" judgment for any inference requirement within 30 minutes. This is the engine decision basis for Deploy an Inference Service from Scratch and Portfolio Projects.

Scope of this article

This article focuses on "which engine to pick," not "how to use a given engine" — for usage of each engine, see the corresponding case study. If you want to build the overall framework of inference optimization first, read What Is Inference Acceleration? and Anatomy of the Overall Architecture.

1. General Principles: Ask Five Questions First ​

1. The Five Real Decision Dimensions ​

GitHub stars, first place on NVIDIA's leaderboard, 24× speedups in papers — these are outcome indicators, not decision criteria. Only five dimensions truly determine the selection:

            ┌─────────────────────────────────────┐
            │         The model you deploy        │
            └─────────────────────────────────────┘
                      │            │            │            │            │
        ┌─────────────▼─┐  ┌──────▼──────┐  ┌──▼─────┐  ┌──▼──────┐  ┌──▼─────────┐
        │ 1 Model type  │  │ 2 Hardware  │  │ 3      │  │ 4       │  │ 5 Team       │
        │ CNN/Transformer│  │ NVIDIA/AMD/ │  │ Deploy │  │ Perf    │  │ capability   │
        │ /LLM/Diffusion │  │ Intel CPU/  │  │ form:  │  │ target: │  │ Python/C++/  │
        │                │  │ mobile      │  │ online │  │ latency │  │ ops skills/  │
        │                │  │             │  │ /batch/│  │ vs      │  │ time budget  │
        │                │  │             │  │ stream/│  │ throughput│ │              │
        │                │  │             │  │ device │  │         │  │              │
        └────────────────┘  └─────────────┘  └─────────┘  └─────────┘  └─────────────┘
                      │            │            │            │            │
                      └─────►  Engine family  ◄──────┘
  • Model type: determines the "engine family." LLM → vLLM/TensorRT-LLM family; classic CV/NLP models → ONNX Runtime/TensorRT; on-device → llama.cpp/TFLite.
  • Hardware: determines the available set of engines. NVIDIA → everything works; AMD → vLLM/TensorRT-LLM ROCm branches; Intel CPU → OpenVINO; mobile → TFLite/CoreML/llama.cpp.
  • Deployment form: determines the "last mile" of the engine. Online serving → vLLM serve / TGI / Triton; offline batch → any engine's Python API; streaming → vLLM stream; on-device → llama.cpp static library.
  • Performance target: latency-first or throughput-first? Latency-first → speculative decoding + small batch; throughput-first → continuous batching + large batch.
  • Team capability: sets the ceiling on tool complexity. A Python team picks up vLLM fast; a C++ team won't suffer with TensorRT/TRT-LLM; weak ops → stay away from engine-build solutions.

2. The Five-Question Checklist ​

Q1: Is your model an LLM (≥ 1B params, Transformer decoder)?
   Yes → LLM engine selection (§2.1)
   No  → classic engine selection (§2.2)

Q2: What is your hardware?
   NVIDIA GPU → vLLM / TensorRT-LLM / TGI / SGLang
   AMD GPU → vLLM (ROCm) / TensorRT-LLM (ROCm, not full-featured)
   Intel CPU → OpenVINO / ONNX Runtime
   Mobile → TFLite / CoreML / llama.cpp
   Edge devices (Jetson etc.) → TensorRT / llama.cpp

Q3: What is your deployment form?
   Online serving → needs HTTP API + concurrency management
   Offline batch → direct Python API
   Streaming generation → needs a stream interface
   On-device app → static library + quantized model

Q4: Is the performance target latency or throughput?
   Latency-first (chat, completion) → speculative decoding + small batch
   Throughput-first (batch jobs, document summarization) → continuous batching + large batch
   Both matter → autoscale; decouple at cluster level

Q5: Team capability?
   Python only → vLLM / TGI
   Python + C++ → TensorRT-LLM / custom kernels
   With NVIDIA background pick the TensorRT family; without it, pick vLLM

The minimal action for engine selection

Don't write an "engine research" slide deck. Use the five questions to narrow down to 2 candidate engines, spend 4 hours each getting a minimal prototype running, measure them with the Inference Benchmarking in Practice methodology, and pick the one that "prototypes fastest and hurts least to operate." One 8-hour hands-on experiment beats three weeks of trend reports.

2. Nine-Engine Head-to-Head Comparison ​

1. Comparison Table ​

Below is a head-to-head comparison of nine mainstream inference engines. The dimensions chosen are the seven that most affect selection: LLM support, quantization, batching strategy, serving, cross-hardware support, learning curve, and community.

EngineLLM supportQuantizationBatching strategyServingCross-hardwareLearning curveCommunity
vLLM★★★ First-classAWQ/GPTQ/FP8/INT8PagedAttention + continuous batchingBuilt-in OpenAI-compatible APINVIDIA / AMD (ROCm)★★ Easy★★★ Extremely active
SGLang★★★ First-classAWQ/GPTQ/FP8RadixAttention + complex schedulingBuilt-in OpenAI-compatible APINVIDIA mainly, partial AMD★★☆ Medium★★☆ Rising
TensorRT-LLM★★★ First-classINT8/INT4/FP8/SmoothQuantIn-flight batchingVia TritonNVIDIA only★★★ Hard (engine build required)★★ Medium, NVIDIA-led
TGI (HuggingFace)★★★ First-classAWQ/GPTQ/INT8/FP8Continuous batchingBuilt-in Rust HTTP serverNVIDIA / AMD★★ Easy★★ Medium
Triton (NVIDIA)No direct LLM support (backend framework)Via backendVia backend★★★ StandardizedNVIDIA / CPU / anything★★★ Hard★★★ Industrial standard
ONNX Runtime★★ Weak (not its focus)INT8/INT4Static / dynamic batchVia Triton or PythonNVIDIA / AMD / Intel / mobile★★ Easy★★★ Cross-platform ecosystem
OpenVINO★★ WeakINT8/INT4Static batchVia OVMSIntel CPU/GPU/VPU★★☆ Medium★★ Intel-led
llama.cpp★★★ First-classGGUF Q4_K_M and moreSingle-request oriented, no continuous batchingVia llama-serverNVIDIA / AMD / CPU / mobile / Mac★☆ Extremely easy★★★ Dominant on-device
TFLite★ WeakINT8 / float16Static batchVia ML Service / Task LibraryAndroid / iOS / edge★★ Easy★★★ Mobile standard

2. Engine Deep Dives ​

vLLM ​

vLLM is the open-source LLM inference engine from UC Berkeley (2023) that became famous overnight on PagedAttention and continuous batching. It is the de facto standard for cloud LLM inference today.

Strengths:

  • PagedAttention eliminates KV cache memory fragmentation, admitting far more concurrency
  • Continuous batching makes multiple requests truly parallel instead of queued
  • OpenAI-compatible API works out of the box — migration cost is minimal
  • Extremely active community; new models and quantization methods land quickly

Weaknesses:

  • Memory gets tight at large batch (default gpu_memory_utilization=0.9 fills the card)
  • Complex scheduling (multi-turn dialogue, structured output) is not as good as SGLang
  • Slightly more deployment complexity than TGI (many parameters)

Best for: the first choice for cloud online LLM inference — the most reliable path from 0 to 1.

SGLang ​

SGLang is an engine that rose in 2024; its core innovation is RadixAttention — pushing KV cache reuse for multi-turn dialogue and shared prefixes to the extreme. See "complex scheduling scenarios" in Batching and Request Scheduling.

Strengths:

  • RadixAttention delivers significantly higher throughput than vLLM in multi-turn dialogue and shared-prefix scenarios
  • Native structured output support (JSON mode, regex-constrained generation)
  • Clear throughput advantage on complex prompt orchestration (multi-turn + few-shot + RAG)

Weaknesses:

  • Smaller community than vLLM, fewer ecosystem plugins
  • Slightly steeper learning curve; the complex API is less intuitive than vLLM's
  • Single-request latency is on par with vLLM — no crushing advantage

Best for: multi-turn dialogue, structured output, and complex prompt orchestration scenarios.

TensorRT-LLM ​

TensorRT-LLM is NVIDIA's officially maintained LLM inference engine, built on TensorRT, with the goal of "extreme performance on NVIDIA hardware."

Strengths:

  • Typically 20–50% higher throughput than vLLM on NVIDIA hardware
  • FP8 quantization delivers 2× speedup on H100/H200 (other engines are not yet complete here)
  • The most complete Triton integration — the steadiest enterprise-grade deployment

Weaknesses:

  • Steepest learning curve: engine build required, with hundreds of tuning parameters
  • NVIDIA-only — deep hardware lock-in
  • Small community; new models land slowly (waiting on NVIDIA's adaptation)
  • Hard to debug; error messages are cryptic

Best for: NVIDIA hardware, extreme latency/throughput sensitivity, teams with NVIDIA engineering background.

TGI ​

TGI (Text Generation Inference) is Hugging Face's LLM serving offering, similar in positioning to vLLM but earlier.

Strengths:

  • Native Hugging Face ecosystem support — new models land fastest
  • Rust HTTP server with stable performance
  • Simplest deployment (one Docker command)

Weaknesses:

  • Usually 10–20% slower than vLLM
  • Weaker complex scheduling than SGLang
  • Commercial use requires a Hugging Face commercial license

Best for: deep Hugging Face ecosystem users, scenarios demanding extreme deployment simplicity.

Triton ​

Triton is NVIDIA's general-purpose inference serving framework. It is not an LLM engine itself, but it can act as a backend framework integrating vLLM, TensorRT-LLM, ONNX Runtime, and more.

Strengths:

  • Unified orchestration of multiple models and versions (LLM + embedder + reranker coexisting)
  • Standard metrics port, seamless Prometheus integration
  • Multiple backends (Python / TensorRT / ONNX / vLLM)
  • Enterprise features: model ensembling, gray releases, A/B testing

Weaknesses:

  • Steep learning curve (complex configuration, many concepts)
  • No direct LLM optimizations (relies on backends)
  • Long debugging chain

Best for: multi-model coexistence, enterprise-grade operations, mid-to-large teams needing unified metrics.

ONNX Runtime ​

ONNX Runtime is Microsoft's cross-platform inference engine, positioned around "model format standardization + cross-hardware."

Strengths:

  • Strongest cross-platform story (NVIDIA / AMD / Intel / mobile all covered)
  • Models exported to ONNX can migrate across engines
  • The EP (Execution Provider) mechanism makes hardware acceleration pluggable

Weaknesses:

  • LLM optimization lags vLLM/TRT-LLM (not its focus)
  • Models occasionally drift numerically after ONNX export (see Common Pitfalls and Anti-Patterns)
  • Complex models fail to export at high rates

Best for: cross-platform deployment of classic CV/NLP models, embedded inference, scenarios requiring hardware portability.

OpenVINO ​

OpenVINO is Intel's inference engine, focused on Intel CPU/GPU/VPU.

Strengths:

  • Extremely strong performance on Intel CPUs (deep AVX-512 optimization)
  • INT8 quantization works well on Intel hardware
  • Seamless integration with the Intel hardware ecosystem (CPU/iGPU/Movidius VPU)

Weaknesses:

  • Completely useless on non-Intel hardware
  • Weak LLM support (not the focus)
  • Small community, slow updates

Best for: pure Intel CPU deployment, edge devices (Movidius), enterprise intranet deployment (no GPUs allowed).

llama.cpp ​

llama.cpp is ggerganov's C++ LLM inference engine — the de facto standard for on-device LLMs.

Strengths:

  • Broadest hardware coverage (CPU / NVIDIA / AMD / Mac M-series / iOS / Android)
  • The GGUF format compresses massively while preserving quality (Q4_K_M is typically 1/4 the original size)
  • Single binary, no dependencies — extremely simple deployment
  • On Mac M-series, the Metal backend beats every other engine by a wide margin

Weaknesses:

  • No continuous batching (single-request oriented)
  • Large-batch performance far below vLLM/TRT-LLM
  • No native HTTP server (llama-server is a basic version)

Best for: on-device deployment, personal devices, Mac users, low-concurrency local serving.

TFLite ​

TFLite is Google's mobile inference framework — the de facto standard for on-device inference on Android.

Strengths:

  • System-level integration on Android, the most stable performance
  • Seamless migration from TensorFlow models
  • NNAPI / GPU Delegate for cross-SoC adaptation

Weaknesses:

  • No LLM support (not in this race)
  • Inferior to CoreML on iOS
  • Occasional model conversion issues

Best for: on-device CV/NLP model deployment on Android, mobile app integration. See Mobile Deployment.

3. LLM Online Inference: A Three-Way Race ​

LLM online inference is the hottest track today, led by three engines: vLLM, TensorRT-LLM, and SGLang.

1. Positioning Differences ​

DimensionvLLMTensorRT-LLMSGLang
One-liner"Open, easy, broad ecosystem""Extreme performance on NVIDIA""Complex-scheduling specialist"
PerformanceMedium (10–30% below TRT-LLM)Highest (on NVIDIA)On par with or slightly above vLLM
Ease of useEasiest (one-line CLI)Hardest (engine build)Medium (complex API)
Deployment time1 hour1–3 days2 hours
Best-fit teamAny Python teamWith NVIDIA engineering backgroundWith complex prompt orchestration needs
Typical usersSmall-to-mid companies, startupsLarge companies, extreme latency sensitivityMulti-turn dialogue, structured output

2. Selection for Three Scenarios ​

Scenario A: Ship an LLM Service from 0 to 1 ​

Pick vLLM first.

Reasons:

  • Simplest deployment (one-line CLI)
  • Most active community — answers are easy to find
  • New models land fast (vLLM usually supports a model on release day)
  • Performance is not the highest, but the SLA rarely demands that last bit

When to switch to TensorRT-LLM: when vLLM's performance misses the SLA, the team has NVIDIA engineering background, and the deployment budget allows ≥ 1 week.

Scenario B: Extreme Performance / Cost Sensitive ​

Pick TensorRT-LLM.

Reasons:

  • Usually the highest throughput on NVIDIA hardware (see the TensorRT-LLM case study)
  • FP8 brings an extra 2× on H100
  • At large scale (dozens of GPUs), the per-instance performance gap multiplies

When to switch back to vLLM: when the team lacks NVIDIA engineering background, when engine-build failure rates are high, or when a new model must ship urgently.

Scenario C: Complex Scheduling Scenarios ​

Pick SGLang.

Reasons:

  • Multi-turn dialogue + shared system prompts: RadixAttention lifts KV cache reuse dramatically
  • Structured output (JSON mode, regex-constrained) natively supported
  • Clear throughput advantage on complex prompt orchestration (few-shot + RAG + multi-turn)

When to switch back to vLLM: when the scenario is single-turn Q&A, structured output doesn't matter, or community support outweighs the performance delta.

3. Real-World Comparison ​

Same Llama-2-7B-Chat, A100 40GB, 32 concurrency, prompt 1024 / output 256:

EngineTTFT (ms)TPOT (ms)Throughput (tok/s)Memory (GB)
vLLM 0.6.320020110036
TensorRT-LLM 0.1318016145032
SGLang 0.319519115035

These numbers only illustrate the differences among the three; absolute values vary widely with configuration. Reproduction: Inference Benchmarking in Practice.

Don't be fooled by a single benchmark

The numbers above reflect only the "single-turn Q&A" scenario. Switch to "multi-turn dialogue + shared system prompt" and SGLang may come back 50% ahead of vLLM. Selection must be benchmarked on your real business scenario — never copy a public leaderboard. See the "anti-cheat checklist" in Inference Benchmarking in Practice.

4. Classic Model Serving: Triton + ONNX Runtime / TensorRT ​

For non-LLM models (CV, classic NLP, tabular models), the mainstream serving path is Triton + ONNX Runtime or Triton + TensorRT.

1. Why Triton Instead of Raw Engines ​

Classic model serving has different requirements from LLMs:

  • Many models (one product may have dozens of CV models)
  • Fast single-model inference (millisecond level); overhead lives in the serving layer
  • Multi-version coexistence (A/B testing, gray releases)

What Triton solves is "unified orchestration of many models and many versions" — performance is secondary. See Triton Inference Server.

2. ONNX Runtime vs TensorRT ​

DimensionONNX RuntimeTensorRT
Cross-hardware★★★ NVIDIA/AMD/Intel/mobileNVIDIA only
PerformanceMedium (10–30% below TRT on NVIDIA)Highest (on NVIDIA)
Model supportExports from almost any frameworkComplex models occasionally hit unsupported ops
DeploymentSimple (no engine build)Engine build required
Best forCross-hardware, cross-platformPure NVIDIA, extreme latency sensitivity

3. Practical Selection ​

Single NVIDIA hardware + extreme latency sensitivity → TensorRT + Triton
Cross-hardware (incl. CPU / mobile)                  → ONNX Runtime + Triton
Many models + multi-version coexistence               → Triton (any backend)
Team without NVIDIA background                        → ONNX Runtime

5. On-Device: llama.cpp / TFLite / CoreML ​

On-device inference faces constraints completely unlike the cloud: tight memory, battery sensitivity, limited compute, multi-SoC compatibility. See llama.cpp and GGUF and Mobile Deployment.

1. Positioning of the Three ​

EnginePlatformMain model typesNotes
llama.cppiOS / Android / Mac / Linux / WindowsLLMs (1B–7B after quantization)Dominant for LLMs on-device
TFLiteAndroid / iOS (second)Classic CV / NLP modelsSystem-level integration on Android
CoreMLiOS onlyCV / NLP / small LLMsBest integration with Apple's stack

2. Practical Selection ​

LLM on iOS            → llama.cpp (Metal backend) + GGUF Q4_K_M
LLM on Android        → llama.cpp (OpenCL backend) + GGUF Q4_K_M
CV model on Android   → TFLite + NNAPI / GPU Delegate
CV model on iOS       → CoreML (best optimized for Apple hardware)
Cross-platform        → llama.cpp (most unified)

3. On-Device-Specific Considerations ​

  • Memory budget: the iPhone 15 Pro has 8 GB; the model can take at most 2–3 GB (system processes need the rest) — a 7B model at Q4 is still 4 GB, so pick a 1.5B
  • Battery: 30 minutes of continuous inference drains the battery — design a low-battery speed-reduction strategy
  • SoC adaptation: Apple A17 / Snapdragon 8 Gen 3 / Dimensity 9300 differ by 2–3× — detect the backend at runtime
  • Privacy compliance: on-device inference is a compliance advantage (data never leaves the device), but protecting model weights is another challenge

See the "on-device LLM app" project in Portfolio Projects, and "calculate the memory budget first" in Deployment Design Principles.

6. Selection Decision Tree ​

Everything above condenses into one decision tree:

What model are you deploying?
│
├─ LLM (≥ 1B params)
│  │
│  ├─ In the cloud
│  │  │
│  │  ├─ NVIDIA + team with NVIDIA background + extreme performance needs
│  │  │  → TensorRT-LLM
│  │  │
│  │  ├─ Multi-turn dialogue / structured output / complex prompts
│  │  │  → SGLang
│  │  │
│  │  └─ Everything else (the default)
│  │     → vLLM
│  │
│  ├─ On-device
│  │  │
│  │  └─ → llama.cpp + GGUF
│  │
│  └─ On Intel CPU
│     → vLLM CPU backend / OpenVINO
│
├─ Classic CV/NLP models
│  │
│  ├─ On NVIDIA GPU
│  │  ├─ Extreme latency sensitivity → TensorRT + Triton
│  │  └─ Performance adequate       → ONNX Runtime + Triton
│  │
│  ├─ On Intel CPU → OpenVINO + OVMS
│  │
│  └─ On mobile
│     ├─ Android → TFLite
│     ├─ iOS → CoreML
│     └─ Cross-platform → ONNX Runtime Mobile / llama.cpp
│
└─ Diffusion models (image generation)
   │
   └─ In the cloud → TensorRT for Diffusion / vLLM-like serving (Stable Diffusion-specific)
      On-device    → CoreML (iOS) / NCNN (Android)

How to use the decision tree

Walking this tree avoids 80% of selection mistakes. The remaining 20% are edge cases (multimodal, exotic hardware) that require checking specific case studies. The decision tree should not replace hands-on benchmarking — after narrowing to 2 candidates, still measure them with the Inference Benchmarking in Practice methodology.

7. Common "Selection Anti-Patterns" ​

Below are selection failure patterns seen in the wild:

1. "NVIDIA's leaderboard has it first, pick it" ​

TensorRT-LLM tops NVIDIA's public leaderboards in throughput, but the leaderboard measures offline batch processing — your scenario is online serving, where p99 latency may matter more than throughput. A leaderboard measures one dimension; your business measures another. See "what to measure" in Inference Benchmarking in Practice.

2. "The team only knows Python, pick vLLM" ​

vLLM is indeed easy to pick up, but when your business needs FP8 quantization or extreme H100 performance, a team without NVIDIA engineering background forcing TensorRT-LLM will actually go slower — high engine-build failure rates, long tuning time, launch slipping away. Selection must follow team capability, not community sentiment — this is "respect your team's stack" from Deployment Design Principles.

3. "New model must ship now, pick vLLM" ​

vLLM does land new models fast, but when your team has already accumulated a stack of TensorRT deployment scripts and monitoring, switching to vLLM for one new model is a stack change — monitoring, ops, and gray releases all get rebuilt. New-model adaptation cost < the cost of rebuilding the whole ops stack; in that case, wait for NVIDIA's TensorRT-LLM adaptation instead of switching engines.

4. "Pick vLLM for on-device" ​

vLLM is not designed for on-device — no Metal/OpenCL backend, no concept of a phone memory budget, no battery optimization. Forcing vLLM onto an iPhone runs 10–20× slower than llama.cpp. On-device belongs to llama.cpp and GGUF.

5. "Run LLMs with ONNX Runtime" ​

ONNX Runtime is not an LLM optimization target — its strength is cross-platform classic models. Forcing it onto LLMs, you will find: no continuous batching, no PagedAttention, primitive KV cache management. Performance lands 5–10× below vLLM. See ONNX Runtime: Cross-Platform.

8. Further Reading ​

References ​