Appearance
Inference Engine Comparison
Choosing an inference engine is not about picking "the most popular" — it is about picking "the one that fits the current problem, the hardware, and the team." Choose wrong and you are reworking the integration three weeks after launch; choose right and both performance and operational cost halve.
In the LLM era, the number of inference engines exploded from 2 (TensorRT, ONNX Runtime) to more than 10, and every one claims to be "the fastest" — the vLLM paper reports 24× over HuggingFace, TensorRT-LLM tops NVIDIA's leaderboards in throughput, SGLang crushes the rest at complex scheduling, llama.cpp is unmatched on-device... These numbers are all true, and none of them should drive your selection. The reason is simple: every number was measured in the scenario that engine is best at, and its transferability to your scenario is close to zero.
This article provides a systematic framework for inference engine selection: ask 5 questions first, then compare 9 engines, and finally decide by 3 scenarios. After reading it, you should be able to produce a "which engine, why, and what the risks are" judgment for any inference requirement within 30 minutes. This is the engine decision basis for Deploy an Inference Service from Scratch and Portfolio Projects.
Scope of this article
This article focuses on "which engine to pick," not "how to use a given engine" — for usage of each engine, see the corresponding case study. If you want to build the overall framework of inference optimization first, read What Is Inference Acceleration? and Anatomy of the Overall Architecture.
1. General Principles: Ask Five Questions First
1. The Five Real Decision Dimensions
GitHub stars, first place on NVIDIA's leaderboard, 24× speedups in papers — these are outcome indicators, not decision criteria. Only five dimensions truly determine the selection:
┌─────────────────────────────────────┐
│ The model you deploy │
└─────────────────────────────────────┘
│ │ │ │ │
┌─────────────▼─┐ ┌──────▼──────┐ ┌──▼─────┐ ┌──▼──────┐ ┌──▼─────────┐
│ 1 Model type │ │ 2 Hardware │ │ 3 │ │ 4 │ │ 5 Team │
│ CNN/Transformer│ │ NVIDIA/AMD/ │ │ Deploy │ │ Perf │ │ capability │
│ /LLM/Diffusion │ │ Intel CPU/ │ │ form: │ │ target: │ │ Python/C++/ │
│ │ │ mobile │ │ online │ │ latency │ │ ops skills/ │
│ │ │ │ │ /batch/│ │ vs │ │ time budget │
│ │ │ │ │ stream/│ │ throughput│ │ │
│ │ │ │ │ device │ │ │ │ │
└────────────────┘ └─────────────┘ └─────────┘ └─────────┘ └─────────────┘
│ │ │ │ │
└─────► Engine family ◄──────┘- Model type: determines the "engine family." LLM → vLLM/TensorRT-LLM family; classic CV/NLP models → ONNX Runtime/TensorRT; on-device → llama.cpp/TFLite.
- Hardware: determines the available set of engines. NVIDIA → everything works; AMD → vLLM/TensorRT-LLM ROCm branches; Intel CPU → OpenVINO; mobile → TFLite/CoreML/llama.cpp.
- Deployment form: determines the "last mile" of the engine. Online serving → vLLM serve / TGI / Triton; offline batch → any engine's Python API; streaming → vLLM stream; on-device → llama.cpp static library.
- Performance target: latency-first or throughput-first? Latency-first → speculative decoding + small batch; throughput-first → continuous batching + large batch.
- Team capability: sets the ceiling on tool complexity. A Python team picks up vLLM fast; a C++ team won't suffer with TensorRT/TRT-LLM; weak ops → stay away from engine-build solutions.
2. The Five-Question Checklist
Q1: Is your model an LLM (≥ 1B params, Transformer decoder)?
Yes → LLM engine selection (§2.1)
No → classic engine selection (§2.2)
Q2: What is your hardware?
NVIDIA GPU → vLLM / TensorRT-LLM / TGI / SGLang
AMD GPU → vLLM (ROCm) / TensorRT-LLM (ROCm, not full-featured)
Intel CPU → OpenVINO / ONNX Runtime
Mobile → TFLite / CoreML / llama.cpp
Edge devices (Jetson etc.) → TensorRT / llama.cpp
Q3: What is your deployment form?
Online serving → needs HTTP API + concurrency management
Offline batch → direct Python API
Streaming generation → needs a stream interface
On-device app → static library + quantized model
Q4: Is the performance target latency or throughput?
Latency-first (chat, completion) → speculative decoding + small batch
Throughput-first (batch jobs, document summarization) → continuous batching + large batch
Both matter → autoscale; decouple at cluster level
Q5: Team capability?
Python only → vLLM / TGI
Python + C++ → TensorRT-LLM / custom kernels
With NVIDIA background pick the TensorRT family; without it, pick vLLMThe minimal action for engine selection
Don't write an "engine research" slide deck. Use the five questions to narrow down to 2 candidate engines, spend 4 hours each getting a minimal prototype running, measure them with the Inference Benchmarking in Practice methodology, and pick the one that "prototypes fastest and hurts least to operate." One 8-hour hands-on experiment beats three weeks of trend reports.
2. Nine-Engine Head-to-Head Comparison
1. Comparison Table
Below is a head-to-head comparison of nine mainstream inference engines. The dimensions chosen are the seven that most affect selection: LLM support, quantization, batching strategy, serving, cross-hardware support, learning curve, and community.
| Engine | LLM support | Quantization | Batching strategy | Serving | Cross-hardware | Learning curve | Community |
|---|---|---|---|---|---|---|---|
| vLLM | ★★★ First-class | AWQ/GPTQ/FP8/INT8 | PagedAttention + continuous batching | Built-in OpenAI-compatible API | NVIDIA / AMD (ROCm) | ★★ Easy | ★★★ Extremely active |
| SGLang | ★★★ First-class | AWQ/GPTQ/FP8 | RadixAttention + complex scheduling | Built-in OpenAI-compatible API | NVIDIA mainly, partial AMD | ★★☆ Medium | ★★☆ Rising |
| TensorRT-LLM | ★★★ First-class | INT8/INT4/FP8/SmoothQuant | In-flight batching | Via Triton | NVIDIA only | ★★★ Hard (engine build required) | ★★ Medium, NVIDIA-led |
| TGI (HuggingFace) | ★★★ First-class | AWQ/GPTQ/INT8/FP8 | Continuous batching | Built-in Rust HTTP server | NVIDIA / AMD | ★★ Easy | ★★ Medium |
| Triton (NVIDIA) | No direct LLM support (backend framework) | Via backend | Via backend | ★★★ Standardized | NVIDIA / CPU / anything | ★★★ Hard | ★★★ Industrial standard |
| ONNX Runtime | ★★ Weak (not its focus) | INT8/INT4 | Static / dynamic batch | Via Triton or Python | NVIDIA / AMD / Intel / mobile | ★★ Easy | ★★★ Cross-platform ecosystem |
| OpenVINO | ★★ Weak | INT8/INT4 | Static batch | Via OVMS | Intel CPU/GPU/VPU | ★★☆ Medium | ★★ Intel-led |
| llama.cpp | ★★★ First-class | GGUF Q4_K_M and more | Single-request oriented, no continuous batching | Via llama-server | NVIDIA / AMD / CPU / mobile / Mac | ★☆ Extremely easy | ★★★ Dominant on-device |
| TFLite | ★ Weak | INT8 / float16 | Static batch | Via ML Service / Task Library | Android / iOS / edge | ★★ Easy | ★★★ Mobile standard |
2. Engine Deep Dives
vLLM
vLLM is the open-source LLM inference engine from UC Berkeley (2023) that became famous overnight on PagedAttention and continuous batching. It is the de facto standard for cloud LLM inference today.
Strengths:
- PagedAttention eliminates KV cache memory fragmentation, admitting far more concurrency
- Continuous batching makes multiple requests truly parallel instead of queued
- OpenAI-compatible API works out of the box — migration cost is minimal
- Extremely active community; new models and quantization methods land quickly
Weaknesses:
- Memory gets tight at large batch (default gpu_memory_utilization=0.9 fills the card)
- Complex scheduling (multi-turn dialogue, structured output) is not as good as SGLang
- Slightly more deployment complexity than TGI (many parameters)
Best for: the first choice for cloud online LLM inference — the most reliable path from 0 to 1.
SGLang
SGLang is an engine that rose in 2024; its core innovation is RadixAttention — pushing KV cache reuse for multi-turn dialogue and shared prefixes to the extreme. See "complex scheduling scenarios" in Batching and Request Scheduling.
Strengths:
- RadixAttention delivers significantly higher throughput than vLLM in multi-turn dialogue and shared-prefix scenarios
- Native structured output support (JSON mode, regex-constrained generation)
- Clear throughput advantage on complex prompt orchestration (multi-turn + few-shot + RAG)
Weaknesses:
- Smaller community than vLLM, fewer ecosystem plugins
- Slightly steeper learning curve; the complex API is less intuitive than vLLM's
- Single-request latency is on par with vLLM — no crushing advantage
Best for: multi-turn dialogue, structured output, and complex prompt orchestration scenarios.
TensorRT-LLM
TensorRT-LLM is NVIDIA's officially maintained LLM inference engine, built on TensorRT, with the goal of "extreme performance on NVIDIA hardware."
Strengths:
- Typically 20–50% higher throughput than vLLM on NVIDIA hardware
- FP8 quantization delivers 2× speedup on H100/H200 (other engines are not yet complete here)
- The most complete Triton integration — the steadiest enterprise-grade deployment
Weaknesses:
- Steepest learning curve: engine build required, with hundreds of tuning parameters
- NVIDIA-only — deep hardware lock-in
- Small community; new models land slowly (waiting on NVIDIA's adaptation)
- Hard to debug; error messages are cryptic
Best for: NVIDIA hardware, extreme latency/throughput sensitivity, teams with NVIDIA engineering background.
TGI
TGI (Text Generation Inference) is Hugging Face's LLM serving offering, similar in positioning to vLLM but earlier.
Strengths:
- Native Hugging Face ecosystem support — new models land fastest
- Rust HTTP server with stable performance
- Simplest deployment (one Docker command)
Weaknesses:
- Usually 10–20% slower than vLLM
- Weaker complex scheduling than SGLang
- Commercial use requires a Hugging Face commercial license
Best for: deep Hugging Face ecosystem users, scenarios demanding extreme deployment simplicity.
Triton
Triton is NVIDIA's general-purpose inference serving framework. It is not an LLM engine itself, but it can act as a backend framework integrating vLLM, TensorRT-LLM, ONNX Runtime, and more.
Strengths:
- Unified orchestration of multiple models and versions (LLM + embedder + reranker coexisting)
- Standard metrics port, seamless Prometheus integration
- Multiple backends (Python / TensorRT / ONNX / vLLM)
- Enterprise features: model ensembling, gray releases, A/B testing
Weaknesses:
- Steep learning curve (complex configuration, many concepts)
- No direct LLM optimizations (relies on backends)
- Long debugging chain
Best for: multi-model coexistence, enterprise-grade operations, mid-to-large teams needing unified metrics.
ONNX Runtime
ONNX Runtime is Microsoft's cross-platform inference engine, positioned around "model format standardization + cross-hardware."
Strengths:
- Strongest cross-platform story (NVIDIA / AMD / Intel / mobile all covered)
- Models exported to ONNX can migrate across engines
- The EP (Execution Provider) mechanism makes hardware acceleration pluggable
Weaknesses:
- LLM optimization lags vLLM/TRT-LLM (not its focus)
- Models occasionally drift numerically after ONNX export (see Common Pitfalls and Anti-Patterns)
- Complex models fail to export at high rates
Best for: cross-platform deployment of classic CV/NLP models, embedded inference, scenarios requiring hardware portability.
OpenVINO
OpenVINO is Intel's inference engine, focused on Intel CPU/GPU/VPU.
Strengths:
- Extremely strong performance on Intel CPUs (deep AVX-512 optimization)
- INT8 quantization works well on Intel hardware
- Seamless integration with the Intel hardware ecosystem (CPU/iGPU/Movidius VPU)
Weaknesses:
- Completely useless on non-Intel hardware
- Weak LLM support (not the focus)
- Small community, slow updates
Best for: pure Intel CPU deployment, edge devices (Movidius), enterprise intranet deployment (no GPUs allowed).
llama.cpp
llama.cpp is ggerganov's C++ LLM inference engine — the de facto standard for on-device LLMs.
Strengths:
- Broadest hardware coverage (CPU / NVIDIA / AMD / Mac M-series / iOS / Android)
- The GGUF format compresses massively while preserving quality (Q4_K_M is typically 1/4 the original size)
- Single binary, no dependencies — extremely simple deployment
- On Mac M-series, the Metal backend beats every other engine by a wide margin
Weaknesses:
- No continuous batching (single-request oriented)
- Large-batch performance far below vLLM/TRT-LLM
- No native HTTP server (llama-server is a basic version)
Best for: on-device deployment, personal devices, Mac users, low-concurrency local serving.
TFLite
TFLite is Google's mobile inference framework — the de facto standard for on-device inference on Android.
Strengths:
- System-level integration on Android, the most stable performance
- Seamless migration from TensorFlow models
- NNAPI / GPU Delegate for cross-SoC adaptation
Weaknesses:
- No LLM support (not in this race)
- Inferior to CoreML on iOS
- Occasional model conversion issues
Best for: on-device CV/NLP model deployment on Android, mobile app integration. See Mobile Deployment.
3. LLM Online Inference: A Three-Way Race
LLM online inference is the hottest track today, led by three engines: vLLM, TensorRT-LLM, and SGLang.
1. Positioning Differences
| Dimension | vLLM | TensorRT-LLM | SGLang |
|---|---|---|---|
| One-liner | "Open, easy, broad ecosystem" | "Extreme performance on NVIDIA" | "Complex-scheduling specialist" |
| Performance | Medium (10–30% below TRT-LLM) | Highest (on NVIDIA) | On par with or slightly above vLLM |
| Ease of use | Easiest (one-line CLI) | Hardest (engine build) | Medium (complex API) |
| Deployment time | 1 hour | 1–3 days | 2 hours |
| Best-fit team | Any Python team | With NVIDIA engineering background | With complex prompt orchestration needs |
| Typical users | Small-to-mid companies, startups | Large companies, extreme latency sensitivity | Multi-turn dialogue, structured output |
2. Selection for Three Scenarios
Scenario A: Ship an LLM Service from 0 to 1
Pick vLLM first.
Reasons:
- Simplest deployment (one-line CLI)
- Most active community — answers are easy to find
- New models land fast (vLLM usually supports a model on release day)
- Performance is not the highest, but the SLA rarely demands that last bit
When to switch to TensorRT-LLM: when vLLM's performance misses the SLA, the team has NVIDIA engineering background, and the deployment budget allows ≥ 1 week.
Scenario B: Extreme Performance / Cost Sensitive
Pick TensorRT-LLM.
Reasons:
- Usually the highest throughput on NVIDIA hardware (see the TensorRT-LLM case study)
- FP8 brings an extra 2× on H100
- At large scale (dozens of GPUs), the per-instance performance gap multiplies
When to switch back to vLLM: when the team lacks NVIDIA engineering background, when engine-build failure rates are high, or when a new model must ship urgently.
Scenario C: Complex Scheduling Scenarios
Pick SGLang.
Reasons:
- Multi-turn dialogue + shared system prompts: RadixAttention lifts KV cache reuse dramatically
- Structured output (JSON mode, regex-constrained) natively supported
- Clear throughput advantage on complex prompt orchestration (few-shot + RAG + multi-turn)
When to switch back to vLLM: when the scenario is single-turn Q&A, structured output doesn't matter, or community support outweighs the performance delta.
3. Real-World Comparison
Same Llama-2-7B-Chat, A100 40GB, 32 concurrency, prompt 1024 / output 256:
| Engine | TTFT (ms) | TPOT (ms) | Throughput (tok/s) | Memory (GB) |
|---|---|---|---|---|
| vLLM 0.6.3 | 200 | 20 | 1100 | 36 |
| TensorRT-LLM 0.13 | 180 | 16 | 1450 | 32 |
| SGLang 0.3 | 195 | 19 | 1150 | 35 |
These numbers only illustrate the differences among the three; absolute values vary widely with configuration. Reproduction: Inference Benchmarking in Practice.
Don't be fooled by a single benchmark
The numbers above reflect only the "single-turn Q&A" scenario. Switch to "multi-turn dialogue + shared system prompt" and SGLang may come back 50% ahead of vLLM. Selection must be benchmarked on your real business scenario — never copy a public leaderboard. See the "anti-cheat checklist" in Inference Benchmarking in Practice.
4. Classic Model Serving: Triton + ONNX Runtime / TensorRT
For non-LLM models (CV, classic NLP, tabular models), the mainstream serving path is Triton + ONNX Runtime or Triton + TensorRT.
1. Why Triton Instead of Raw Engines
Classic model serving has different requirements from LLMs:
- Many models (one product may have dozens of CV models)
- Fast single-model inference (millisecond level); overhead lives in the serving layer
- Multi-version coexistence (A/B testing, gray releases)
What Triton solves is "unified orchestration of many models and many versions" — performance is secondary. See Triton Inference Server.
2. ONNX Runtime vs TensorRT
| Dimension | ONNX Runtime | TensorRT |
|---|---|---|
| Cross-hardware | ★★★ NVIDIA/AMD/Intel/mobile | NVIDIA only |
| Performance | Medium (10–30% below TRT on NVIDIA) | Highest (on NVIDIA) |
| Model support | Exports from almost any framework | Complex models occasionally hit unsupported ops |
| Deployment | Simple (no engine build) | Engine build required |
| Best for | Cross-hardware, cross-platform | Pure NVIDIA, extreme latency sensitivity |
3. Practical Selection
Single NVIDIA hardware + extreme latency sensitivity → TensorRT + Triton
Cross-hardware (incl. CPU / mobile) → ONNX Runtime + Triton
Many models + multi-version coexistence → Triton (any backend)
Team without NVIDIA background → ONNX Runtime5. On-Device: llama.cpp / TFLite / CoreML
On-device inference faces constraints completely unlike the cloud: tight memory, battery sensitivity, limited compute, multi-SoC compatibility. See llama.cpp and GGUF and Mobile Deployment.
1. Positioning of the Three
| Engine | Platform | Main model types | Notes |
|---|---|---|---|
| llama.cpp | iOS / Android / Mac / Linux / Windows | LLMs (1B–7B after quantization) | Dominant for LLMs on-device |
| TFLite | Android / iOS (second) | Classic CV / NLP models | System-level integration on Android |
| CoreML | iOS only | CV / NLP / small LLMs | Best integration with Apple's stack |
2. Practical Selection
LLM on iOS → llama.cpp (Metal backend) + GGUF Q4_K_M
LLM on Android → llama.cpp (OpenCL backend) + GGUF Q4_K_M
CV model on Android → TFLite + NNAPI / GPU Delegate
CV model on iOS → CoreML (best optimized for Apple hardware)
Cross-platform → llama.cpp (most unified)3. On-Device-Specific Considerations
- Memory budget: the iPhone 15 Pro has 8 GB; the model can take at most 2–3 GB (system processes need the rest) — a 7B model at Q4 is still 4 GB, so pick a 1.5B
- Battery: 30 minutes of continuous inference drains the battery — design a low-battery speed-reduction strategy
- SoC adaptation: Apple A17 / Snapdragon 8 Gen 3 / Dimensity 9300 differ by 2–3× — detect the backend at runtime
- Privacy compliance: on-device inference is a compliance advantage (data never leaves the device), but protecting model weights is another challenge
See the "on-device LLM app" project in Portfolio Projects, and "calculate the memory budget first" in Deployment Design Principles.
6. Selection Decision Tree
Everything above condenses into one decision tree:
What model are you deploying?
│
├─ LLM (≥ 1B params)
│ │
│ ├─ In the cloud
│ │ │
│ │ ├─ NVIDIA + team with NVIDIA background + extreme performance needs
│ │ │ → TensorRT-LLM
│ │ │
│ │ ├─ Multi-turn dialogue / structured output / complex prompts
│ │ │ → SGLang
│ │ │
│ │ └─ Everything else (the default)
│ │ → vLLM
│ │
│ ├─ On-device
│ │ │
│ │ └─ → llama.cpp + GGUF
│ │
│ └─ On Intel CPU
│ → vLLM CPU backend / OpenVINO
│
├─ Classic CV/NLP models
│ │
│ ├─ On NVIDIA GPU
│ │ ├─ Extreme latency sensitivity → TensorRT + Triton
│ │ └─ Performance adequate → ONNX Runtime + Triton
│ │
│ ├─ On Intel CPU → OpenVINO + OVMS
│ │
│ └─ On mobile
│ ├─ Android → TFLite
│ ├─ iOS → CoreML
│ └─ Cross-platform → ONNX Runtime Mobile / llama.cpp
│
└─ Diffusion models (image generation)
│
└─ In the cloud → TensorRT for Diffusion / vLLM-like serving (Stable Diffusion-specific)
On-device → CoreML (iOS) / NCNN (Android)How to use the decision tree
Walking this tree avoids 80% of selection mistakes. The remaining 20% are edge cases (multimodal, exotic hardware) that require checking specific case studies. The decision tree should not replace hands-on benchmarking — after narrowing to 2 candidates, still measure them with the Inference Benchmarking in Practice methodology.
7. Common "Selection Anti-Patterns"
Below are selection failure patterns seen in the wild:
1. "NVIDIA's leaderboard has it first, pick it"
TensorRT-LLM tops NVIDIA's public leaderboards in throughput, but the leaderboard measures offline batch processing — your scenario is online serving, where p99 latency may matter more than throughput. A leaderboard measures one dimension; your business measures another. See "what to measure" in Inference Benchmarking in Practice.
2. "The team only knows Python, pick vLLM"
vLLM is indeed easy to pick up, but when your business needs FP8 quantization or extreme H100 performance, a team without NVIDIA engineering background forcing TensorRT-LLM will actually go slower — high engine-build failure rates, long tuning time, launch slipping away. Selection must follow team capability, not community sentiment — this is "respect your team's stack" from Deployment Design Principles.
3. "New model must ship now, pick vLLM"
vLLM does land new models fast, but when your team has already accumulated a stack of TensorRT deployment scripts and monitoring, switching to vLLM for one new model is a stack change — monitoring, ops, and gray releases all get rebuilt. New-model adaptation cost < the cost of rebuilding the whole ops stack; in that case, wait for NVIDIA's TensorRT-LLM adaptation instead of switching engines.
4. "Pick vLLM for on-device"
vLLM is not designed for on-device — no Metal/OpenCL backend, no concept of a phone memory budget, no battery optimization. Forcing vLLM onto an iPhone runs 10–20× slower than llama.cpp. On-device belongs to llama.cpp and GGUF.
5. "Run LLMs with ONNX Runtime"
ONNX Runtime is not an LLM optimization target — its strength is cross-platform classic models. Forcing it onto LLMs, you will find: no continuous batching, no PagedAttention, primitive KV cache management. Performance lands 5–10× below vLLM. See ONNX Runtime: Cross-Platform.
8. Further Reading
- What Is Inference Acceleration? — where engines sit in the inference stack
- Anatomy of the Overall Architecture — engines and their upstream/downstream neighbors
- vLLM and PagedAttention — deep dive into the vLLM engine
- TensorRT-LLM — deep dive into the TRT-LLM engine
- ONNX Runtime: Cross-Platform — deep dive into ONNX Runtime
- OpenVINO and CPU Inference — deep dive into OpenVINO
- llama.cpp and GGUF — deep dive into llama.cpp
- Triton Inference Server — deep dive into the orchestration layer
- Deploy an Inference Service from Scratch — end-to-end landing after selection
- Inference Benchmarking in Practice — methodology for selection comparisons
- Tuning and Performance Optimization — tuning after the engine is chosen
- Deployment Design Principles — the discipline of engine selection
- Common Pitfalls and Anti-Patterns — the full version of selection anti-patterns
- Batching and Request Scheduling — batching differences across engines
- Model Serving and Orchestration — the basis of the orchestration layer
References
- vLLM Documentation — vLLM official documentation
- SGLang: Documentation — SGLang official repository
- TensorRT-LLM: Documentation — TRT-LLM official documentation
- TGI: Text Generation Inference — Hugging Face TGI repository
- Triton Inference Server: User Guide — Triton official documentation
- ONNX Runtime: Documentation — ONNX Runtime official documentation
- OpenVINO: Documentation — OpenVINO official documentation
- llama.cpp: GitHub — llama.cpp repository
- TensorFlow Lite: Documentation — TFLite official documentation
- MLPerf Inference: Language — cross-engine standardized benchmark