Skip to content

Curated Resource List

At a glance Curated resources for model deployment: official documentation, classic courses, high-quality blogs, and communities—organized by purpose with a one-line review for each entry, so you never have to dredge the internet blind.

Curated Resource List ​

How to Use This List

Browse by category. Each entry comes with a one-line review to help you judge whether it's worth clicking. Links change over time—if one is dead, look for the latest entry point on the project's official GitHub or homepage. The end includes a "start with these 5" starter list.

Official Docs and Tool Sites ​

NameLinkOne-Line Review
ONNX websitehttps://onnx.ai/The spec and ecosystem hub for the open model exchange format—the starting point for model conversion.
ONNX Runtimehttps://onnxruntime.ai/Microsoft's cross-platform inference engine supporting CPU/GPU/NPU—the go-to runtime for getting started with deployment.
NVIDIA TensorRThttps://developer.nvidia.com/tensorrtOfficial home of NVIDIA's high-performance inference compiler, with docs, downloads, and GTC materials.
Triton Inference Serverhttps://github.com/triton-inference-server/serverA production-grade, multi-framework, multi-model, multi-GPU inference server; the benchmark for dynamic batching and concurrent execution.
vLLM docshttps://docs.vllm.ai/en/latest/The de facto standard LLM inference engine—the pioneering work behind PagedAttention + continuous batching.
KServehttps://kserve.github.io/website/The standard inference platform on Kubernetes, turning "deploy a model" declaratively into YAML.
BentoMLhttps://www.bentoml.com/A Python-native serving framework for inference—packaging, deployment, and observability in one place.
MLflowhttps://mlflow.org/The industry standard for experiment tracking + model registry—the bridge from training to registration to deployment.
Hugging Face Inference Endpointshttps://huggingface.co/docs/inference-endpoints/indexManaged service that hosts models on the HF Hub as APIs—go live with a few lines of config.
llama.cpphttps://github.com/ggml-org/llama.cppA C/C++ implementation with GGUF quantization support—the universal key to running LLMs locally, on CPU, and on Apple Silicon.
OpenVINOhttps://docs.openvino.ai/Intel's inference toolchain across its whole hardware line (CPU/GPU/NPU)—a strong pick for edge and AI PC scenarios.
TensorFlow Litehttps://www.tensorflow.org/liteThe official solution for mobile and embedded inference, with companion model conversion and quantization tools.
FastAPIhttps://fastapi.tiangolo.com/The most popular API framework in Python; nearly every model-serving HTTP layer is built on it.
PyTorchhttps://pytorch.org/The de facto standard training framework; deployment tools like TorchServe/ExecuTorch start from here.
Text Generation Inference (TGI)https://github.com/huggingface/text-generation-inferenceHF's high-performance LLM server, now in maintenance mode (the official recommendation for new projects is vLLM/SGLang).

Courses ​

NameLinkOne-Line Review
MIT 6.5940 TinyML and Efficient Deep Learning Computinghttps://hanlab.mit.edu/courses/2024-fall-65940The bible-level course on efficient deep learning: pruning, quantization, distillation, and LLM deployment all covered, with labs and assignments.
Stanford CS 329S Machine Learning Systems Designhttps://stanford-cs329s.github.io/An ML systems design course with a complete methodology from requirements analysis to production monitoring; companion book Designing Machine Learning Systems.
MLOps Zoomcamp (DataTalks.Club)https://github.com/DataTalksClub/mlops-zoomcampA free 9-week hands-on MLOps course: experiment tracking, deployment, monitoring, and CI/CD end to end.
Full Stack Deep Learninghttps://fullstackdeeplearning.com/A full-stack deep learning course from training to deployment to monitoring, including an LLM Bootcamp; free and public.

Classic Papers ​

NameLinkOne-Line Review
This handbook's Papers section/papers/vllm-paper and /papers/quantization-papersIn-house deep dives: PagedAttention and quantization classics like GPTQ/AWQ, with commentary.
arXivhttps://arxiv.org/The first-hand source for papers; follow the latest research with keywords (LLM inference / quantization / serving).

Tips for Reading Papers

Start with this handbook's paper deep dives to build a framework, then go back to arXiv for the originals. Note: arXiv preprints are not peer-reviewed—cross-check implementations (such as the vLLM and llama.cpp code) before productionizing.

High-Quality Blogs ​

NameLinkOne-Line Review
NVIDIA Technical Bloghttps://developer.nvidia.com/blogOfficial deep dives into TensorRT/Triton/GPU performance; the quantization series and inference optimization posts deserve a careful read.
vLLM Bloghttps://blog.vllm.aiFirst-hand material on LLM inference systems engineering: the principles behind PagedAttention, continuous batching, and speculative decoding.
Hugging Face Bloghttps://huggingface.co/blogThe home base for model, library, and ecosystem updates, often with high-quality systems and evaluation articles.
BentoML Bloghttps://www.bentoml.com/blogHands-on serving and deployment, including LLM inference performance comparisons and best practices.
Databricks Bloghttps://www.databricks.com/blogExceptionally good long-form posts from the MLOps/data engineering perspective—required reading for understanding "production-grade machine learning".
Anyscale Bloghttps://www.anyscale.com/blogThe official Ray ecosystem blog, rich in practical experience with distributed inference and Ray Serve.

Communities ​

NameLinkOne-Line Review
Reddit r/MachineLearninghttps://www.reddit.com/r/MachineLearning/The world's largest ML community—the first place papers drop and industry discussions happen.
Hacker Newshttps://news.ycombinator.com/Tech news with high signal-to-noise comments; new framework launches often spread from here.
awesome-production-machine-learninghttps://github.com/EthicalML/awesome-production-machine-learningA 20k+ star catalog of production-grade ML tools (⚠️ the list may be outdated; verify entries yourself).
Hugging Face Hubhttps://huggingface.co/Models, datasets, Spaces, and docs in one place—the default starting point of the LLM era.

Chinese-Language Resources ​

Personal Picks

The following are the author's personal preferences, for reference only; they do not represent the site's official stance. WeChat official accounts have no stable links—search for the names inside WeChat.

NameLinkOne-Line Review
Zhihu topics "Model Deployment / LLM Inference Optimization"https://www.zhihu.com/Search within the site for "model deployment", "LLM inference optimization", and "MLOps"—plenty of long-form war stories from front-line engineers.
Synced (jiqizhixin.com)https://www.jiqizhixin.com/The leading Chinese AI industry and engineering news outlet, with broad coverage of LLM deployment cases.
QbitAIhttps://www.qbitai.com/Fast-updating AI news, well-informed on industry trends and open-source model progress.
WeChat official accounts: NVIDIA China, Datawhale, OSChina, and othersSearch inside WeChatCommunities like Datawhale continuously produce Chinese-language tutorial series (e.g., LLM inference, MLOps)—great for structured onboarding.

Start With These 5 ​

  1. This handbook first: the Glossary + What Is Model Deployment build a 30-minute concept map.
  2. The vLLM blog (https://blog.vllm.ai): understand why LLM inference is fast from the "PagedAttention and Continuous Batching" post.
  3. The ONNX Runtime official quickstart (https://onnxruntime.ai/): run the minimal export → convert → infer loop with the most popular runtime.
  4. The MIT 6.5940 course (https://hanlab.mit.edu/courses/2024-fall-65940): systematically fill in the foundations of quantization, pruning, and distillation.
  5. The NVIDIA Technical Blog's inference performance series (https://developer.nvidia.com/blog): turn "performance" from black magic into a measurable, optimizable engineering problem.

Further Reading ​