Appearance
Learning Paths: Three Routes, from Beginner to Mastery
Model deployment isn't a field you can finish by reading one book. It's an engineering system assembled from dozens of interlocking pieces — inference engines, model formats, quantization, serving architecture, observability, MLOps — and if any one piece is missing, you can't ship a system to production. This article organizes the whole body of knowledge into three routes: job-hunt sprint (about 2 weeks), systematic deep-dive (about 8 weeks), and desk reference (look things up as needed), with a page list and checkpoints for each.
Start here
The three guide pages are the skeleton of this handbook; read them first no matter which route you take: What Is Model Deployment, Deployment vs. MLOps vs. Inference vs. Model Serving, and Anatomy of an Inference System.
1. Why You Need a Roadmap: Knowledge Dependencies
Deployment knowledge isn't a flat landscape; it has clear prerequisites. The classic symptom of reading in the wrong order: staring blankly at the INT8 quantization formulas in /concepts/quantization because you haven't yet worked out how inference actually executes a single neural-network layer in /concepts/inference.
You can draw the dependencies as a directed graph:
text
Concepts before practice; inference fundamentals before performance optimization
┌────────────────────────────┐
│ Inference fundamentals │ ← where everything starts
└──────────────┬─────────────┘
┌──────────────────────┼──────────────────────┐
▼ ▼ ▼
┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐
│ Model formats & │ │ Quantization / │ │ Serving & │
│ hardware │ │ compression │ │ architecture │
└─────────────────┘ └─────────────────┘ └─────────────────┘
│ │ │
▼ ▼ ▼
┌───────────────────────────────────────────────────────────┐
│ Performance & observability (performance + monitoring)│
└───────────────────────────────────────────────────────────┘
│ │
▼ ▼
┌─────────────────┐ ┌─────────────────┐
│ Case studies & │ │ LLM inference │
│ pitfalls │ │ (llm) │
└─────────────────┘ └─────────────────┘Three iron rules:
- Concepts before practice:
/concepts/*is the theoretical foundation;/practice/*and/case-studies/*are the structure built on top of it. Jump straight into case studies and you'll keep hitting gaps of the form "but why do it this way?". - Inference fundamentals before optimization: Without the mathematical execution path of inference, you can't understand why quantization speeds things up, why accuracy is lost, or why GPU memory is the bottleneck.
- General before specialized: LLM inference (continuous batching, KV Cache) builds on general inference concepts — learn it after the general material.
2. Comparing the Three Routes
| Route | Who it's for | Time required | Goal | Entry page |
|---|---|---|---|---|
| Job-hunt sprint | New grads and career changers preparing for interviews | About 2 weeks | Cover the high-frequency interview topics; able to answer them and to write about them | What Is Model Deployment |
| Systematic deep-dive | Engineers who want to build a complete, structured skill set | About 8 weeks | Build a full knowledge tree and complete an integrative practice project | Anatomy of an Inference System |
| Desk reference | Already doing deployment; needs to look things up when problems hit | As needed | Problem → answer, located within 15 minutes | A Brief History of Deployment |
INFO
The three routes aren't mutually exclusive: most people can run the 2-week job-hunt sprint first to get the big picture, then use the 8-week route to fill in depth, and finally keep the desk reference as a permanent companion at work.
3. Route One: Job-Hunt Sprint (About 2 Weeks)
Work through four stages: match the JD → core concepts → practice → interview questions. Start from the job description (JD) of the role you're targeting, mark its most frequent keywords (TensorRT, vLLM, quantization, K8s, serving), then knock out the checklist below item by item.
Stage 1: Match the JD and Get the Big Picture (Days 1-2)
- What Is Model Deployment — be able to state the four elements of deployment and its core challenges
- Deployment vs. MLOps vs. Inference vs. Model Serving — be able to draw the concept hierarchy
- Anatomy of an Inference System — be able to reproduce the seven-layer architecture diagram from memory
Checkpoint: Without looking at your notes, explain to an interviewer within 5 minutes what problem a deployment engineer actually solves.
Stage 2: Core Concepts (Days 3-8)
| Topic | Page | High-frequency interview questions |
|---|---|---|
| Inference | Inference Fundamentals | How inference differs from the training backward pass |
| Model formats | Model Formats | What ONNX is for; the PTH→ONNX→TensorRT pipeline |
| Hardware | Hardware Fundamentals | GPU memory bandwidth; how CPU and GPU split the work |
| Quantization | Quantization | How INT8/FP16 work; the accuracy trade-offs |
| Compression | Model Compression | Distillation / pruning / sparsification |
| Serving | Model Serving | gRPC vs REST; dynamic batching |
| Deployment patterns | Deployment Patterns | Online / offline / edge / Serverless |
| Performance | Performance Optimization | Latency/throughput metrics; locating bottlenecks |
| Monitoring | Monitoring | Prometheus metrics; drift detection |
| LLM | LLM Inference | KV Cache; continuous batching |
Checkpoint: For every topic, be able to name a real tool (ONNX Runtime, TensorRT, Triton, vLLM, KServe, MLflow, Prometheus) and explain where it fits.
Stage 3: Practice and Case Studies (Days 9-11)
Pick 2-3 case studies for a close read, focusing on the arc from problem → engine choice → solution → results:
- FastAPI Serving and Triton Case Study (required; they cover the two most mainstream routes)
- vLLM Case Study (required for LLM-oriented roles)
- TensorRT Edge Deployment (required for embedded roles)
Checkpoint: For at least one case study, explain why that engine was chosen over the alternatives (comparison criteria in Framework Comparison).
Stage 4: Interview Sprint (Days 12-14)
- Work through the interview questions answering each one yourself; map anything you can't answer back to its concept page
- Use Common Pitfalls to check whether your answers are purely theoretical
Checkpoint: You can talk for a full 2 minutes — no notes — on 80% of the questions in the list.
4. Route Two: Systematic Deep-Dive (About 8 Weeks)
One topic per week; self-test against the weekend checkpoint. The week 8 capstone is the acceptance test.
| Week | Topic | Pages | Weekend checkpoint |
|---|---|---|---|
| Week 1 | Guide & inference fundamentals | Guide, Inference Fundamentals, Hardware Fundamentals | Explain the memory and compute flow of one forward pass |
| Week 2 | Model formats & quantization | Model Formats, Quantization, Compression | Plan a PTH→ONNX→INT8 pipeline from scratch, unaided |
| Week 3 | Serving & architecture | Model Serving, Deployment Patterns, MLOps Pipeline | Draw the complete component diagram of an online serving stack |
| Week 4 | Performance & monitoring | Performance Optimization, Monitoring, Observability in Practice | Design a metrics system covering latency, throughput, and drift |
| Week 5 | LLM inference | LLM Inference, vLLM Paper | Explain clearly why continuous batching and KV Cache exist |
| Week 6 | Case studies | Triton Case Study, FastAPI Case Study, vLLM Case Study, Gateway Canary Release | Retell the key engine-selection rationale for each case |
| Week 7 | Reading papers in depth | Quantization Papers, vLLM Paper, Paper Index | State the core problem each paper tackles |
| Week 8 | Capstone & job hunting | Build Your Own, Load Testing, Rollout, Interview Questions | Independently build and ship a small service with load testing, monitoring, and canary release |
Don't skip week 8
The job-hunt sprint only asks you to know; the systematic route must get you to do. In week 8, actually run the project in Build Your Own. Those load-test numbers will show up on your resume — and in the interviewer's follow-up questions.
5. Route Three: Desk Reference (Look It Up as Needed)
Organized as problem → page. When something breaks, first figure out which layer the symptom belongs to, then jump to the matching page:
| The problem you're facing | Go straight to |
|---|---|
| Model inference got slower | Performance Optimization, Load Testing Methods |
| How to slim down a model | Quantization, Model Compression |
| No idea how to turn a model into a service | Model Serving, Build Your Own |
| Which inference engine to choose | Framework Comparison, Triton Case Study |
| Quality dropped after launch | Monitoring & Drift, Rollout Process |
| Hit a baffling failure | Common Pitfalls |
| Need a term's definition | Glossary |
| Want a ready-made tool list | Tool Inventory, Awesome List |
| Can't follow the underlying principles | System Primer, A Brief History of Deployment |
| Preparing for interviews | Interview Questions |
How to use it: Bookmark this page. When a problem comes up, check the table first, then read the "one-sentence takeaway" section of the target page. The desk-reference route is all about speed: within 15 minutes, either find the answer or figure out exactly who to ask.
6. FAQ
Should I read every concept page before building anything?
No. Read concept pages on demand: read the three guide pages first to build the skeleton, then start building immediately — get a minimal service running, and come back to the concepts when you get stuck. Build first, read second, and the concepts get "used" into memory. Read everything first and you'll have forgotten the first third by the time you write any code.
Short on time — what gets cut?
- Only 1 week: cut the week 7 papers and the week 5 LLM details from the systematic route, and keep just the first three stages of the job-hunt sprint.
- Only 3 days: read just the three guide pages plus the interview questions — the questions will sweep you back through the high-frequency topics.
- Never cut:
/concepts/inference(inference fundamentals) and/practice/pitfalls(common pitfalls). The former is the foundation every other concept sits on; the latter is the fastest way to avoid stepping on the same rakes.
Do the papers need a close reading?
Depends on the goal. Engineering roles: read the vLLM paper closely and skim the rest. Research or senior roles: work through 3-5 papers from the paper index using a problem → method → validation frame — it will meaningfully set you apart from other candidates.
Can the three routes run in parallel?
Yes. The recommended combination: keep the desk reference permanently at hand (look up problems any time) + the job-hunt sprint (the 2 weeks before you join or switch roles) + the systematic deep-dive (1-2 months after joining, to fill in depth). The deep-dive doesn't need to be done in one pass — split it by week and weave it into your day job.
Further Reading
- What Is Model Deployment — the starting point of every route: what deployment is and its four elements
- Anatomy of an Inference System — where each concept lives in the system
- A Brief History of Deployment — how today's tools came to be, so you don't miss the forest for the trees
- Build Your First Inference Service — the hands-on guide for week 8 of route two
- Model Deployment Interview Questions — the finish line and final checkpoint of the job-hunt sprint
- Glossary — a vocabulary index you can consult at any time