Appearance
Portfolio Projects
On the resume of an inference deployment candidate, the line that shines brightest is not "used vLLM" but "built a RAG service from scratch — 50 concurrent users, p99 800 ms": the kind of project you can talk about for thirty minutes from a single sentence.
Interviews for inference deployment roles have a curious pattern: candidates who can explain their own project are rarer than candidates who can recite vLLM parameters. The reason is simple: anyone can memorize vLLM parameters, but building an end-to-end inference service from scratch, keeping it stable, finding its bottlenecks, and measuring its numbers accurately — that skill can only be honed on real projects. This article presents 5 most representative portfolio projects for inference deployment. Every project satisfies three conditions: it has motivation (solves a real problem), it has deliverables (code + report + demo), and it has selling points you can talk about (30 minutes in an interview without running dry).
Where this article sits
This article is the "expansion route" of Deploy an Inference Service from Scratch — that article is a single-model, three-version progressive end-to-end project, while this one extends it into five engineering works in different directions, each able to stand alone as a portfolio piece. Every project assumes you can already run v2 of build-your-own (vLLM + single-GPU serving), and builds on top of it. For the complete job-oriented interpretation, see Skills Benchmarking: What to Highlight on Your Resume.
1. The Five Projects at a Glance
| Project | Main technologies | Difficulty | Effort | Resume boost |
|---|---|---|---|---|
| 1. RAG inference service | embedder + vector DB + reranker + LLM | ★★ | 1–2 weeks | "I got the full RAG stack running" |
| 2. Multi-model routing | small model + large model + routing policy | ★★★ | 2 weeks | "I cut costs by 70%" |
| 3. Streaming chat API | SSE + vLLM streaming + client | ★★ | 1 week | "I understand token-level streaming" |
| 4. Quantization comparison benchmark platform | multi-engine, multi-quantization comparison + report generation | ★★★ | 2–3 weeks | "I have a quantization decision framework" |
| 5. On-device LLM app | llama.cpp + iOS/Android + GGUF | ★★★★ | 3–4 weeks | "I can handle on-device deployment" |
Match projects to yourself in three steps: job target → current ability → effort budget. If your target is cloud inference engineer, projects 1/2/4 are the top picks; if your target is on-device / Mobile AI, project 5 is the top pick; if you want the broadest coverage, the 1+2+4 trio tells the best story.
2. Project 1: RAG Inference Service
1. Motivation
RAG (Retrieval-Augmented Generation) is the most common shape of LLM adoption — any "answer questions over an enterprise knowledge base" requirement is RAG at its core. Its engineering challenge is not the LLM itself but stringing four inference services — embedder, vector retrieval, reranker, and LLM — into one pipeline while keeping end-to-end latency in check. This project extends the "single LLM" of Deploy an Inference Service from Scratch into a "multi-model pipeline," and it is the project every cloud inference engineer should have.
2. Tech Stack
User request ──→ embedder (BGE-M3) ──→ vector DB (Qdrant/Milvus)
│
▼
top-K retrieval (K=10)
│
▼
reranker (bge-reranker)
│
▼
top-N reranking (N=3)
│
▼
LLM (Llama-2-7B-Chat via vLLM)
│
▼
Answer + citations| Component | Recommended choice | Alternatives | Notes |
|---|---|---|---|
| Embedder | BGE-M3 / bge-large-en-v1.5 | sentence-transformers, OpenAI embedder | Deploy with ONNX Runtime — see ONNX Runtime: Cross-Platform |
| Vector DB | Qdrant | Milvus, Weaviate, pgvector | Start single-node with Qdrant — simple and stable |
| Reranker | bge-reranker-large | Cohere rerank API | Run directly with transformers, or export to ONNX |
| LLM | Llama-2-7B-Chat via vLLM | Qwen2.5-7B-Instruct | See Deploy an Inference Service from Scratch |
| Orchestration | Python + FastAPI | LangChain, LlamaIndex | Don't use LangChain to orchestrate production — plain FastAPI is more controllable |
3. Deliverables
rag-inference-service/
├── README.md # Architecture diagram, benchmarks, reproduction guide
├── docker-compose.yml # One-command start: Qdrant + vLLM + API
├── configs/
│ ├── embedder.yaml
│ ├── retriever.yaml
│ └── llm.yaml
├── src/
│ ├── embedder.py # ONNX inference
│ ├── retriever.py # Qdrant client
│ ├── reranker.py # transformers inference
│ ├── llm.py # vLLM client
│ ├── pipeline.py # Wire the four into FastAPI
│ └── eval.py # ragas evaluation
├── data/
│ └── docs/ # Test knowledge base (public wiki excerpts)
├── bench/
│ ├── bench_e2e.py # End-to-end latency
│ └── bench_quality.py # Answer quality evaluation
└── reports/
└── bench_2024-08-22.md4. Selling Points You Can Talk About
- "I got the full RAG stack running": four inference services — embedder, vector DB, reranker, LLM — each with its own accuracy and latency optimization
- "I have an end-to-end SLA": measured "single-query TTFT < 500 ms, p99 < 2 s, 30 QPS throughput"
- "I have an evaluation set": ran ragas on faithfulness, answer_relevancy, and context_precision — "recall 0.85, answer accuracy 0.78"
- "I understand selection trade-offs": why Qdrant over Milvus (deployment complexity), why BGE-M3 over the OpenAI embedder (cost and privacy)
- "I know bottleneck analysis": profiling showed the reranker took 60% of end-to-end latency, so I converted it to ONNX INT8 and cut overall latency in half — this is the "check bottlenecks at all three layers: model, operator, system" principle from Deployment Design Principles
3. Project 2: Multi-Model Routing
1. Motivation
LLM cost is a core KPI for SaaS companies — a service that sends every request to GPT-4 costs millions of dollars per month. But in the real request distribution, only 20% of requests truly need GPT-4-level capability; 80% are fine with a 7B model. Multi-model routing is the answer: route easy requests to the small model based on request difficulty, and leave hard ones for the large model. This project showcases the advanced application of Model Serving and Orchestration and Batching and Request Scheduling.
2. Tech Stack
User request
│
▼
┌─────────────────┐
│ Router │ ← a lightweight classifier (e.g., BERT-tiny)
│ (complexity │ or a rule-based heuristic
│ scoring) │
└────────┬────────┘
│
┌────────┴────────┐
▼ ▼
Easy requests 80% Hard requests 20%
│ │
▼ ▼
Llama-2-7B Llama-2-70B / GPT-4
(local vLLM) (API or local)| Component | Choice |
|---|---|
| Router classifier | BERT-tiny + rules (length, keywords, context complexity) |
| Small model | Llama-2-7B-Chat via vLLM (local) |
| Large model | Llama-2-70B (local, 4-GPU TP) or GPT-4 API |
| Routing framework | Triton Inference Server + ensemble models |
| Monitoring | Prometheus + Grafana, tracking traffic on both paths, error rate, degradation rate |
3. Deliverables
multi-model-router/
├── README.md
├── configs/
│ ├── router.yaml
│ ├── small_model.yaml
│ └── large_model.yaml
├── src/
│ ├── router.py # Router classifier
│ ├── clients.py # Clients for both LLMs
│ ├── orchestrator.py # FastAPI main service
│ ├── fallback.py # Degrade to the small model on large-model timeout
│ └── metrics.py # Prometheus metrics
├── triton/
│ └── model_repository/
│ ├── router_bert/
│ ├── llama2-7b/
│ └── ensemble/
├── bench/
│ ├── bench_routing.py # Routing accuracy
│ └── bench_cost.py # Cost comparison
└── reports/
└── cost_saving_2024-08-22.md4. Selling Points You Can Talk About
- "I cut costs by 70%": measured 80% of traffic on the small model — average cost is 30% of the large-model-only setup
- "I have a routing accuracy evaluation": 1,000 human-labeled "difficulty" samples, routing accuracy 0.82
- "I have a degradation mechanism": automatic fallback to the small model when the large model times out or gets rate-limited, keeping the service available
- "I have gray releases": using Triton's multi-version mechanism, a new routing policy runs on 5% of traffic first, expanding to 100% after 24 hours
- "I know why not to use LangChain Router": LangChain Router is prompt-based routing (letting an LLM decide), which costs one LLM call per routing decision and ends up more expensive; this design uses a lightweight classifier — routing itself takes only 5 ms
This is the textbook application of "gray-release new configs at 5% traffic" from Deployment Design Principles, and the counterexample to "large batch for online inference" in Common Pitfalls and Anti-Patterns (the router classifier must stay at small batch and low latency).
4. Project 3: Streaming Chat API
1. Motivation
User patience in chat scenarios is measured in seconds — if nothing appears within 3 seconds, the user concludes "it's frozen." Streaming output is the core technique for cutting time to first token, yet many teams still use a non-streaming interface that "waits for the model to finish, then returns everything at once," maxing out TTFT. This project is the engineering implementation of "streaming beats waiting for the result" from Latency, Throughput, and Concurrency.
2. Tech Stack
Browser / App
│
│ SSE (Server-Sent Events)
▼
FastAPI (async stream)
│
│ vLLM stream API
▼
vLLM (Llama-2-7B-Chat)
│
│ token-level streaming generation
▼
Return token by token to the client| Component | Choice |
|---|---|
| Inference engine | vLLM 0.6+ (native stream support) |
| Protocol | SSE (Server-Sent Events), not WebSocket |
| Web framework | FastAPI + sse-starlette |
| Client | Browser EventSource / Python httpx |
| Monitoring | TTFT distribution, per-request TPOT, concurrency |
3. Deliverables
streaming-chat-api/
├── README.md
├── src/
│ ├── server.py # FastAPI + SSE
│ ├── vllm_client.py # Async stream client
│ ├── backpressure.py # Backpressure control
│ └── metrics.py
├── client/
│ ├── web.html # Browser EventSource demo
│ └── python_cli.py # CLI client
├── bench/
│ ├── bench_ttft.py # TTFT distribution
│ └── bench_concurrent.py # Concurrent streaming
└── docker/
└── Dockerfile4. Selling Points You Can Talk About
- "I understand the SSE vs WebSocket trade-off": streaming output uses SSE, not WebSocket — SSE is one-directional, auto-reconnecting, and HTTP-compatible; WebSocket is bidirectional but more complex
- "I measured the TTFT vs TPOT gap": the streaming interface drops the perceived "first-token latency" from 2 s to 200 ms while actual generation speed stays the same — that is the essence of experience engineering
- "I understand backpressure": when the client consumes slowly, the server cannot pile up tokens without bound — a backpressure mechanism is required
- "I understand fairness across users": when 32 streaming requests arrive at once, the early ones must not hog the GPU — continuous batching + fair scheduling
- "I know how to test streaming": streaming latency is not request-response; it must be measured as chunk-level latency — see Inference Benchmarking in Practice
This is the counterexample to "streaming responses done with polling instead of SSE" in Common Pitfalls and Anti-Patterns, and the canonical implementation of "streaming beats waiting for the result" in Deployment Design Principles.
5. Project 4: Quantization Comparison Benchmark Platform
1. Motivation
Quantization is the core lever for LLM inference cost optimization, but there are too many quantization schemes (AWQ, GPTQ, SmoothQuant, INT8, INT4, FP8...), too many models (Llama, Qwen, Mistral...), and too many engines (vLLM, TRT-LLM, SGLang...) — every team must make a "which quantization scheme" decision, and without a unified comparison platform, the decision is made by gut feel. This project turns the Inference Benchmarking in Practice methodology into an automated platform — a seriously hardcore piece of engineering.
2. Tech Stack
Config (YAML)
│
▼
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ Quantization │ → │ Inference │ → │ Performance │
│ module │ │ engine │ │ collection │
│ (autoawq, │ │ (vLLM, │ │ (nvidia-smi │
│ auto-gptq, │ │ TRT-LLM, │ │ + custom) │
│ smoothquant │ │ SGLang) │ │ │
│ ) │ │ │ │ │
└─────────────┘ └─────────────┘ └─────────────┘
│
▼
┌─────────────┐
│ Quality eval │
│ (lm-eval- │
│ harness) │
└─────────────┘
│
▼
┌─────────────┐
│ Report │
│ generation │
│ (Markdown + │
│ HTML viz) │
└─────────────┘| Component | Choice |
|---|---|
| Quantization libraries | AutoAWQ, AutoGPTQ, SmoothQuant, llm-compressor |
| Engines | vLLM, TensorRT-LLM, SGLang |
| Performance collection | vLLM bench, trtllm-bench, nvidia-smi, Nsight Systems |
| Quality evaluation | lm-evaluation-harness (MMLU, HellaSwag, HumanEval, etc.) |
| Orchestration | Python + Click CLI |
| Reports | Jinja2 templates → Markdown + HTML |
3. Deliverables
quant-bench-platform/
├── README.md
├── configs/
│ ├── awq_llama2_vllm.yaml
│ ├── gptq_llama2_trtllm.yaml
│ └── ... (dozens of combinations)
├── src/
│ ├── quantize/
│ │ ├── awq_runner.py
│ │ ├── gptq_runner.py
│ │ └── smoothquant_runner.py
│ ├── engines/
│ │ ├── vllm_runner.py
│ │ ├── trtllm_runner.py
│ │ └── sglang_runner.py
│ ├── metrics/
│ │ ├── perf.py
│ │ └── quality.py
│ └── report.py
├── reports/
│ ├── 2024-08-22_llama2_quant.md
│ └── 2024-08-22_llama2_quant.html
└── cli.py # Entry point: python cli.py run --config xxx.yaml4. Selling Points You Can Talk About
- "I have a complete quantization decision framework": measured AWQ vs GPTQ vs SmoothQuant on Llama-2-7B across throughput, memory, and MMLU, producing a "when to pick which" decision tree
- "I have cross-engine comparisons": performance of the same AWQ model on vLLM / TRT-LLM / SGLang
- "I have a reproducible platform": anyone can run your cli.py + config files on their machine and reproduce the same numbers — this is the Inference Benchmarking in Practice report template turned into engineering
- "I know cheat-proofing": every benchmark separates cold/warm start, handles CPU-GPU overlap, and explicitly states padding
- "I have visual reports": the HTML report embeds interactive charts (Plotly) that an interviewer can click through on the spot
This is the engineering implementation that turns the theory of Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision into operable decisions, and the "data backend" for Inference Engine Comparison.
6. Project 5: On-Device LLM App
1. Motivation
On-device LLMs are the hottest direction of 2024–2025 — iOS 18, Android 15, and Windows Copilot+ are all moving LLMs onto the device. On-device deployment challenges are completely different from the cloud: tight memory budget (the iPhone 15 Pro has only 8 GB), battery sensitivity, limited compute, and the need to support multiple SoCs (Apple Silicon / Snapdragon / MediaTek). This project is the most on-target engineering work for on-device / Mobile AI roles.
2. Tech Stack
iOS App / Android App
│
▼
llama.cpp (C++)
│
▼
GGUF quantized model (Q4_K_M)
│
▼
Metal (iOS) / OpenCL (Android) backend
│
▼
Streaming generation + context management| Component | Choice |
|---|---|
| Inference engine | llama.cpp (see llama.cpp and GGUF) |
| Quantization format | GGUF Q4_K_M (4-bit, balanced quality and size) |
| iOS backend | Metal Performance Shaders |
| Android backend | OpenCL / Vulkan |
| Model | Qwen2.5-1.5B-Instruct (small enough to run on a phone) |
| App framework | SwiftUI (iOS) / Jetpack Compose (Android) |
| Context management | Sliding window + context compression |
3. Deliverables
on-device-llm-app/
├── README.md
├── ios/
│ ├── LlamaApp/ # SwiftUI app
│ ├── llama.cpp/ # git submodule
│ └── build_scripts/
├── android/
│ ├── app/ # Jetpack Compose
│ ├── llama.cpp/
│ └── build.gradle
├── models/
│ └── qwen2.5-1.5b-instruct-q4_k_m.gguf # Not committed; README provides a download script
├── bench/
│ ├── bench_ios.sh # Benchmark on an iPhone
│ └── bench_android.sh # Benchmark on Android
└── docs/
├── ios_memory_budget.md # Memory budget analysis per iPhone model
├── android_soc_matrix.md # Performance matrix per SoC
└── battery_test.md # Battery impact test4. Selling Points You Can Talk About
- "I run a 1.5B model on an iPhone 15 Pro at TTFT 1.2 s and TPOT 80 ms": real, demonstrable on-device performance
- "I understand battery optimization": measured X% battery drain over 30 minutes of continuous chat, and designed a "reduce generation speed in low-power mode" strategy
- "I understand model selection": why 1.5B rather than 7B (phone memory budget), why Qwen rather than Llama (higher quality in Chinese scenarios)
- "I understand cross-SoC adaptation": performance differences across Apple A17, Snapdragon 8 Gen 3, and Dimensity 9300, with a "detect the SoC at runtime and pick the backend" solution
- "I understand on-device-specific challenges": memory budget, thermal throttling, battery management, privacy compliance — problems the cloud never has
See llama.cpp and GGUF and Mobile Deployment, and the extreme-case application of "calculate the memory budget first" from Deployment Design Principles.
7. Project Combination Advice
You don't need to build every project — going deep on 2–3 beats skimming 5. Three combination suggestions:
Combination A: Cloud Inference Engineer (Most Common)
- Project 1 (RAG service): covers the multi-model pipeline
- Project 2 (multi-model routing): covers cost optimization and orchestration
- Project 4 (quantization benchmark platform): covers quantization and decision-making
This trio covers the three core cloud inference skills: end-to-end pipeline, cost optimization, and quantization decisions. You can talk about all three in an interview; see Skills Benchmarking: What to Highlight on Your Resume.
Combination B: On-Device / Mobile AI Engineer
- Project 5 (on-device LLM app): the core
- Project 3 (streaming chat API): adds cloud experience
- Project 1 (RAG service): adds pipeline experience
An on-device resume must have project 5, or it won't pass the first screen. Projects 3 and 1 fill the gap because on-device engineers often collaborate with the cloud (hybrid device-cloud architectures) — not understanding cloud inference is a real weakness.
Combination C: Performance Optimization / Accelerator Roles
- Project 4 (quantization benchmark platform): the core
- Project 2 (multi-model routing): covers system-level scheduling
- Custom project: CUDA kernel optimization (not covered here; see Kernel Fusion and Custom Kernels)
The core skill for performance roles is "measure and decide," and project 4 is its most direct expression. Add one system-level piece (multi-model routing) and one low-level piece (kernels) to complete a full "measure → decide → low-level" capability chain.
8. The "Standard Kit" Every Project Needs
Whichever project you pick, these 5 standard items are non-negotiable — this is Deployment Design Principles applied to a portfolio:
- A README written like a paper: motivation, architecture diagram, rationale for technology choices, performance numbers, reproduction commands — miss any one and it fails
- A standalone benchmark report: a complete Inference Benchmarking in Practice report template under
reports/ - A reproducible docker-compose: the interviewer can
docker compose upyour project in one command - A demo video or screenshots: code alone can't make someone understand your project in 30 seconds — a demo is more persuasive than code
- An interview script: write down in advance "if the interviewer asks for a 5-minute walkthrough, what do I say" — see the STAR framework in Interview Question Bank
9. Further Reading
- Deploy an Inference Service from Scratch — the shared foundation of all five projects
- Inference Engine Comparison — the basis for engine choices in projects 1/2/4
- Inference Benchmarking in Practice — the performance reporting methodology for all five projects
- Deployment Design Principles — the design discipline for all five projects
- Common Pitfalls and Anti-Patterns — the pitfall checklist for all five projects
- Model Serving and Orchestration — the theory behind projects 1/2/3
- llama.cpp and GGUF — the core engine of project 5
- Mobile Deployment — the extension of project 5
- Skills Benchmarking: What to Highlight on Your Resume — project combinations and resume presentation
- Interview Question Bank — reference for project scripts
References
- Qdrant Documentation — vector DB for project 1
- bge-reranker — reranker for project 1
- Ragas: RAG Evaluation — evaluation for project 1
- FastAPI + SSE — service framework for project 3
- sse-starlette — SSE implementation for project 3
- AutoAWQ — quantization for project 4
- AutoGPTQ — quantization for project 4
- lm-evaluation-harness — evaluation for project 4
- llama.cpp — on-device engine for project 5
- GGUF Format — quantization format for project 5
- Triton Inference Server: Model Ensemble — multi-model orchestration for project 2