Appearance
Curated Resource List
How to Use This List
Browse by category. Each entry comes with a one-line review to help you judge whether it's worth clicking. Links change over time—if one is dead, look for the latest entry point on the project's official GitHub or homepage. The end includes a "start with these 5" starter list.
Official Docs and Tool Sites
| Name | Link | One-Line Review |
|---|---|---|
| ONNX website | https://onnx.ai/ | The spec and ecosystem hub for the open model exchange format—the starting point for model conversion. |
| ONNX Runtime | https://onnxruntime.ai/ | Microsoft's cross-platform inference engine supporting CPU/GPU/NPU—the go-to runtime for getting started with deployment. |
| NVIDIA TensorRT | https://developer.nvidia.com/tensorrt | Official home of NVIDIA's high-performance inference compiler, with docs, downloads, and GTC materials. |
| Triton Inference Server | https://github.com/triton-inference-server/server | A production-grade, multi-framework, multi-model, multi-GPU inference server; the benchmark for dynamic batching and concurrent execution. |
| vLLM docs | https://docs.vllm.ai/en/latest/ | The de facto standard LLM inference engine—the pioneering work behind PagedAttention + continuous batching. |
| KServe | https://kserve.github.io/website/ | The standard inference platform on Kubernetes, turning "deploy a model" declaratively into YAML. |
| BentoML | https://www.bentoml.com/ | A Python-native serving framework for inference—packaging, deployment, and observability in one place. |
| MLflow | https://mlflow.org/ | The industry standard for experiment tracking + model registry—the bridge from training to registration to deployment. |
| Hugging Face Inference Endpoints | https://huggingface.co/docs/inference-endpoints/index | Managed service that hosts models on the HF Hub as APIs—go live with a few lines of config. |
| llama.cpp | https://github.com/ggml-org/llama.cpp | A C/C++ implementation with GGUF quantization support—the universal key to running LLMs locally, on CPU, and on Apple Silicon. |
| OpenVINO | https://docs.openvino.ai/ | Intel's inference toolchain across its whole hardware line (CPU/GPU/NPU)—a strong pick for edge and AI PC scenarios. |
| TensorFlow Lite | https://www.tensorflow.org/lite | The official solution for mobile and embedded inference, with companion model conversion and quantization tools. |
| FastAPI | https://fastapi.tiangolo.com/ | The most popular API framework in Python; nearly every model-serving HTTP layer is built on it. |
| PyTorch | https://pytorch.org/ | The de facto standard training framework; deployment tools like TorchServe/ExecuTorch start from here. |
| Text Generation Inference (TGI) | https://github.com/huggingface/text-generation-inference | HF's high-performance LLM server, now in maintenance mode (the official recommendation for new projects is vLLM/SGLang). |
Courses
| Name | Link | One-Line Review |
|---|---|---|
| MIT 6.5940 TinyML and Efficient Deep Learning Computing | https://hanlab.mit.edu/courses/2024-fall-65940 | The bible-level course on efficient deep learning: pruning, quantization, distillation, and LLM deployment all covered, with labs and assignments. |
| Stanford CS 329S Machine Learning Systems Design | https://stanford-cs329s.github.io/ | An ML systems design course with a complete methodology from requirements analysis to production monitoring; companion book Designing Machine Learning Systems. |
| MLOps Zoomcamp (DataTalks.Club) | https://github.com/DataTalksClub/mlops-zoomcamp | A free 9-week hands-on MLOps course: experiment tracking, deployment, monitoring, and CI/CD end to end. |
| Full Stack Deep Learning | https://fullstackdeeplearning.com/ | A full-stack deep learning course from training to deployment to monitoring, including an LLM Bootcamp; free and public. |
Classic Papers
| Name | Link | One-Line Review |
|---|---|---|
| This handbook's Papers section | /papers/vllm-paper and /papers/quantization-papers | In-house deep dives: PagedAttention and quantization classics like GPTQ/AWQ, with commentary. |
| arXiv | https://arxiv.org/ | The first-hand source for papers; follow the latest research with keywords (LLM inference / quantization / serving). |
Tips for Reading Papers
Start with this handbook's paper deep dives to build a framework, then go back to arXiv for the originals. Note: arXiv preprints are not peer-reviewed—cross-check implementations (such as the vLLM and llama.cpp code) before productionizing.
High-Quality Blogs
| Name | Link | One-Line Review |
|---|---|---|
| NVIDIA Technical Blog | https://developer.nvidia.com/blog | Official deep dives into TensorRT/Triton/GPU performance; the quantization series and inference optimization posts deserve a careful read. |
| vLLM Blog | https://blog.vllm.ai | First-hand material on LLM inference systems engineering: the principles behind PagedAttention, continuous batching, and speculative decoding. |
| Hugging Face Blog | https://huggingface.co/blog | The home base for model, library, and ecosystem updates, often with high-quality systems and evaluation articles. |
| BentoML Blog | https://www.bentoml.com/blog | Hands-on serving and deployment, including LLM inference performance comparisons and best practices. |
| Databricks Blog | https://www.databricks.com/blog | Exceptionally good long-form posts from the MLOps/data engineering perspective—required reading for understanding "production-grade machine learning". |
| Anyscale Blog | https://www.anyscale.com/blog | The official Ray ecosystem blog, rich in practical experience with distributed inference and Ray Serve. |
Communities
| Name | Link | One-Line Review |
|---|---|---|
| Reddit r/MachineLearning | https://www.reddit.com/r/MachineLearning/ | The world's largest ML community—the first place papers drop and industry discussions happen. |
| Hacker News | https://news.ycombinator.com/ | Tech news with high signal-to-noise comments; new framework launches often spread from here. |
| awesome-production-machine-learning | https://github.com/EthicalML/awesome-production-machine-learning | A 20k+ star catalog of production-grade ML tools (⚠️ the list may be outdated; verify entries yourself). |
| Hugging Face Hub | https://huggingface.co/ | Models, datasets, Spaces, and docs in one place—the default starting point of the LLM era. |
Chinese-Language Resources
Personal Picks
The following are the author's personal preferences, for reference only; they do not represent the site's official stance. WeChat official accounts have no stable links—search for the names inside WeChat.
| Name | Link | One-Line Review |
|---|---|---|
| Zhihu topics "Model Deployment / LLM Inference Optimization" | https://www.zhihu.com/ | Search within the site for "model deployment", "LLM inference optimization", and "MLOps"—plenty of long-form war stories from front-line engineers. |
| Synced (jiqizhixin.com) | https://www.jiqizhixin.com/ | The leading Chinese AI industry and engineering news outlet, with broad coverage of LLM deployment cases. |
| QbitAI | https://www.qbitai.com/ | Fast-updating AI news, well-informed on industry trends and open-source model progress. |
| WeChat official accounts: NVIDIA China, Datawhale, OSChina, and others | Search inside WeChat | Communities like Datawhale continuously produce Chinese-language tutorial series (e.g., LLM inference, MLOps)—great for structured onboarding. |
Start With These 5
- This handbook first: the Glossary + What Is Model Deployment build a 30-minute concept map.
- The vLLM blog (https://blog.vllm.ai): understand why LLM inference is fast from the "PagedAttention and Continuous Batching" post.
- The ONNX Runtime official quickstart (https://onnxruntime.ai/): run the minimal export → convert → infer loop with the most popular runtime.
- The MIT 6.5940 course (https://hanlab.mit.edu/courses/2024-fall-65940): systematically fill in the foundations of quantization, pruning, and distillation.
- The NVIDIA Technical Blog's inference performance series (https://developer.nvidia.com/blog): turn "performance" from black magic into a measurable, optimizable engineering problem.
Further Reading
- Glossary: look up unfamiliar terms here first
- Systems Fundamentals Cheat Sheet: the Linux/containers/K8s/GPU command handbook for deployment troubleshooting
- Tool Profiles: a selection comparison table by use case
- Practice Guide: deploy a model from scratch
- Paper Deep Dive: the PagedAttention systems paper