Skip to content

Portfolio Projects

At a glance Five portfolio-worthy projects for inference and deployment roles: a RAG inference service, multi-model routing, a streaming chat API, a quantization comparison benchmark platform, and an on-device LLM app. Each project comes with motivation, tech stack, deliverables, and selling points you can talk about.

Portfolio Projects ​

On the resume of an inference deployment candidate, the line that shines brightest is not "used vLLM" but "built a RAG service from scratch — 50 concurrent users, p99 800 ms": the kind of project you can talk about for thirty minutes from a single sentence.

Interviews for inference deployment roles have a curious pattern: candidates who can explain their own project are rarer than candidates who can recite vLLM parameters. The reason is simple: anyone can memorize vLLM parameters, but building an end-to-end inference service from scratch, keeping it stable, finding its bottlenecks, and measuring its numbers accurately — that skill can only be honed on real projects. This article presents 5 most representative portfolio projects for inference deployment. Every project satisfies three conditions: it has motivation (solves a real problem), it has deliverables (code + report + demo), and it has selling points you can talk about (30 minutes in an interview without running dry).

Where this article sits

This article is the "expansion route" of Deploy an Inference Service from Scratch — that article is a single-model, three-version progressive end-to-end project, while this one extends it into five engineering works in different directions, each able to stand alone as a portfolio piece. Every project assumes you can already run v2 of build-your-own (vLLM + single-GPU serving), and builds on top of it. For the complete job-oriented interpretation, see Skills Benchmarking: What to Highlight on Your Resume.

1. The Five Projects at a Glance ​

ProjectMain technologiesDifficultyEffortResume boost
1. RAG inference serviceembedder + vector DB + reranker + LLM★★1–2 weeks"I got the full RAG stack running"
2. Multi-model routingsmall model + large model + routing policy★★★2 weeks"I cut costs by 70%"
3. Streaming chat APISSE + vLLM streaming + client★★1 week"I understand token-level streaming"
4. Quantization comparison benchmark platformmulti-engine, multi-quantization comparison + report generation★★★2–3 weeks"I have a quantization decision framework"
5. On-device LLM appllama.cpp + iOS/Android + GGUF★★★★3–4 weeks"I can handle on-device deployment"

Match projects to yourself in three steps: job target → current ability → effort budget. If your target is cloud inference engineer, projects 1/2/4 are the top picks; if your target is on-device / Mobile AI, project 5 is the top pick; if you want the broadest coverage, the 1+2+4 trio tells the best story.

2. Project 1: RAG Inference Service ​

1. Motivation ​

RAG (Retrieval-Augmented Generation) is the most common shape of LLM adoption — any "answer questions over an enterprise knowledge base" requirement is RAG at its core. Its engineering challenge is not the LLM itself but stringing four inference services — embedder, vector retrieval, reranker, and LLM — into one pipeline while keeping end-to-end latency in check. This project extends the "single LLM" of Deploy an Inference Service from Scratch into a "multi-model pipeline," and it is the project every cloud inference engineer should have.

2. Tech Stack ​

User request ──→ embedder (BGE-M3) ──→ vector DB (Qdrant/Milvus)
                                          │
                                          ▼
                                    top-K retrieval (K=10)
                                          │
                                          ▼
                                    reranker (bge-reranker)
                                          │
                                          ▼
                                    top-N reranking (N=3)
                                          │
                                          ▼
                            LLM (Llama-2-7B-Chat via vLLM)
                                          │
                                          ▼
                                       Answer + citations
ComponentRecommended choiceAlternativesNotes
EmbedderBGE-M3 / bge-large-en-v1.5sentence-transformers, OpenAI embedderDeploy with ONNX Runtime — see ONNX Runtime: Cross-Platform
Vector DBQdrantMilvus, Weaviate, pgvectorStart single-node with Qdrant — simple and stable
Rerankerbge-reranker-largeCohere rerank APIRun directly with transformers, or export to ONNX
LLMLlama-2-7B-Chat via vLLMQwen2.5-7B-InstructSee Deploy an Inference Service from Scratch
OrchestrationPython + FastAPILangChain, LlamaIndexDon't use LangChain to orchestrate production — plain FastAPI is more controllable

3. Deliverables ​

rag-inference-service/
├── README.md                # Architecture diagram, benchmarks, reproduction guide
├── docker-compose.yml       # One-command start: Qdrant + vLLM + API
├── configs/
│   ├── embedder.yaml
│   ├── retriever.yaml
│   └── llm.yaml
├── src/
│   ├── embedder.py          # ONNX inference
│   ├── retriever.py         # Qdrant client
│   ├── reranker.py          # transformers inference
│   ├── llm.py               # vLLM client
│   ├── pipeline.py          # Wire the four into FastAPI
│   └── eval.py              # ragas evaluation
├── data/
│   └── docs/                # Test knowledge base (public wiki excerpts)
├── bench/
│   ├── bench_e2e.py         # End-to-end latency
│   └── bench_quality.py     # Answer quality evaluation
└── reports/
    └── bench_2024-08-22.md

4. Selling Points You Can Talk About ​

  • "I got the full RAG stack running": four inference services — embedder, vector DB, reranker, LLM — each with its own accuracy and latency optimization
  • "I have an end-to-end SLA": measured "single-query TTFT < 500 ms, p99 < 2 s, 30 QPS throughput"
  • "I have an evaluation set": ran ragas on faithfulness, answer_relevancy, and context_precision — "recall 0.85, answer accuracy 0.78"
  • "I understand selection trade-offs": why Qdrant over Milvus (deployment complexity), why BGE-M3 over the OpenAI embedder (cost and privacy)
  • "I know bottleneck analysis": profiling showed the reranker took 60% of end-to-end latency, so I converted it to ONNX INT8 and cut overall latency in half — this is the "check bottlenecks at all three layers: model, operator, system" principle from Deployment Design Principles

3. Project 2: Multi-Model Routing ​

1. Motivation ​

LLM cost is a core KPI for SaaS companies — a service that sends every request to GPT-4 costs millions of dollars per month. But in the real request distribution, only 20% of requests truly need GPT-4-level capability; 80% are fine with a 7B model. Multi-model routing is the answer: route easy requests to the small model based on request difficulty, and leave hard ones for the large model. This project showcases the advanced application of Model Serving and Orchestration and Batching and Request Scheduling.

2. Tech Stack ​

                    User request
                        │
                        ▼
              ┌─────────────────┐
              │  Router          │  ← a lightweight classifier (e.g., BERT-tiny)
              │  (complexity     │     or a rule-based heuristic
              │   scoring)       │
              └────────┬────────┘
                       │
              ┌────────┴────────┐
              ▼                 ▼
        Easy requests 80%   Hard requests 20%
              │                 │
              ▼                 ▼
        Llama-2-7B        Llama-2-70B / GPT-4
        (local vLLM)      (API or local)
ComponentChoice
Router classifierBERT-tiny + rules (length, keywords, context complexity)
Small modelLlama-2-7B-Chat via vLLM (local)
Large modelLlama-2-70B (local, 4-GPU TP) or GPT-4 API
Routing frameworkTriton Inference Server + ensemble models
MonitoringPrometheus + Grafana, tracking traffic on both paths, error rate, degradation rate

3. Deliverables ​

multi-model-router/
├── README.md
├── configs/
│   ├── router.yaml
│   ├── small_model.yaml
│   └── large_model.yaml
├── src/
│   ├── router.py            # Router classifier
│   ├── clients.py           # Clients for both LLMs
│   ├── orchestrator.py      # FastAPI main service
│   ├── fallback.py          # Degrade to the small model on large-model timeout
│   └── metrics.py           # Prometheus metrics
├── triton/
│   └── model_repository/
│       ├── router_bert/
│       ├── llama2-7b/
│       └── ensemble/
├── bench/
│   ├── bench_routing.py     # Routing accuracy
│   └── bench_cost.py        # Cost comparison
└── reports/
    └── cost_saving_2024-08-22.md

4. Selling Points You Can Talk About ​

  • "I cut costs by 70%": measured 80% of traffic on the small model — average cost is 30% of the large-model-only setup
  • "I have a routing accuracy evaluation": 1,000 human-labeled "difficulty" samples, routing accuracy 0.82
  • "I have a degradation mechanism": automatic fallback to the small model when the large model times out or gets rate-limited, keeping the service available
  • "I have gray releases": using Triton's multi-version mechanism, a new routing policy runs on 5% of traffic first, expanding to 100% after 24 hours
  • "I know why not to use LangChain Router": LangChain Router is prompt-based routing (letting an LLM decide), which costs one LLM call per routing decision and ends up more expensive; this design uses a lightweight classifier — routing itself takes only 5 ms

This is the textbook application of "gray-release new configs at 5% traffic" from Deployment Design Principles, and the counterexample to "large batch for online inference" in Common Pitfalls and Anti-Patterns (the router classifier must stay at small batch and low latency).

4. Project 3: Streaming Chat API ​

1. Motivation ​

User patience in chat scenarios is measured in seconds — if nothing appears within 3 seconds, the user concludes "it's frozen." Streaming output is the core technique for cutting time to first token, yet many teams still use a non-streaming interface that "waits for the model to finish, then returns everything at once," maxing out TTFT. This project is the engineering implementation of "streaming beats waiting for the result" from Latency, Throughput, and Concurrency.

2. Tech Stack ​

Browser / App
     │
     │ SSE (Server-Sent Events)
     ▼
FastAPI (async stream)
     │
     │ vLLM stream API
     ▼
vLLM (Llama-2-7B-Chat)
     │
     │ token-level streaming generation
     ▼
Return token by token to the client
ComponentChoice
Inference enginevLLM 0.6+ (native stream support)
ProtocolSSE (Server-Sent Events), not WebSocket
Web frameworkFastAPI + sse-starlette
ClientBrowser EventSource / Python httpx
MonitoringTTFT distribution, per-request TPOT, concurrency

3. Deliverables ​

streaming-chat-api/
├── README.md
├── src/
│   ├── server.py            # FastAPI + SSE
│   ├── vllm_client.py       # Async stream client
│   ├── backpressure.py      # Backpressure control
│   └── metrics.py
├── client/
│   ├── web.html             # Browser EventSource demo
│   └── python_cli.py        # CLI client
├── bench/
│   ├── bench_ttft.py        # TTFT distribution
│   └── bench_concurrent.py  # Concurrent streaming
└── docker/
    └── Dockerfile

4. Selling Points You Can Talk About ​

  • "I understand the SSE vs WebSocket trade-off": streaming output uses SSE, not WebSocket — SSE is one-directional, auto-reconnecting, and HTTP-compatible; WebSocket is bidirectional but more complex
  • "I measured the TTFT vs TPOT gap": the streaming interface drops the perceived "first-token latency" from 2 s to 200 ms while actual generation speed stays the same — that is the essence of experience engineering
  • "I understand backpressure": when the client consumes slowly, the server cannot pile up tokens without bound — a backpressure mechanism is required
  • "I understand fairness across users": when 32 streaming requests arrive at once, the early ones must not hog the GPU — continuous batching + fair scheduling
  • "I know how to test streaming": streaming latency is not request-response; it must be measured as chunk-level latency — see Inference Benchmarking in Practice

This is the counterexample to "streaming responses done with polling instead of SSE" in Common Pitfalls and Anti-Patterns, and the canonical implementation of "streaming beats waiting for the result" in Deployment Design Principles.

5. Project 4: Quantization Comparison Benchmark Platform ​

1. Motivation ​

Quantization is the core lever for LLM inference cost optimization, but there are too many quantization schemes (AWQ, GPTQ, SmoothQuant, INT8, INT4, FP8...), too many models (Llama, Qwen, Mistral...), and too many engines (vLLM, TRT-LLM, SGLang...) — every team must make a "which quantization scheme" decision, and without a unified comparison platform, the decision is made by gut feel. This project turns the Inference Benchmarking in Practice methodology into an automated platform — a seriously hardcore piece of engineering.

2. Tech Stack ​

Config (YAML)
   │
   ▼
┌─────────────┐    ┌─────────────┐    ┌─────────────┐
│ Quantization │ →  │ Inference   │ →  │ Performance │
│ module       │    │ engine      │    │ collection │
│ (autoawq,    │    │ (vLLM,      │    │ (nvidia-smi │
│  auto-gptq,  │    │  TRT-LLM,   │    │  + custom) │
│  smoothquant │    │  SGLang)    │    │             │
│  )           │    │             │    │             │
└─────────────┘    └─────────────┘    └─────────────┘
                                              │
                                              ▼
                                       ┌─────────────┐
                                       │ Quality eval │
                                       │ (lm-eval-    │
                                       │  harness)    │
                                       └─────────────┘
                                              │
                                              ▼
                                       ┌─────────────┐
                                       │ Report       │
                                       │ generation   │
                                       │ (Markdown +  │
                                       │  HTML viz)   │
                                       └─────────────┘
ComponentChoice
Quantization librariesAutoAWQ, AutoGPTQ, SmoothQuant, llm-compressor
EnginesvLLM, TensorRT-LLM, SGLang
Performance collectionvLLM bench, trtllm-bench, nvidia-smi, Nsight Systems
Quality evaluationlm-evaluation-harness (MMLU, HellaSwag, HumanEval, etc.)
OrchestrationPython + Click CLI
ReportsJinja2 templates → Markdown + HTML

3. Deliverables ​

quant-bench-platform/
├── README.md
├── configs/
│   ├── awq_llama2_vllm.yaml
│   ├── gptq_llama2_trtllm.yaml
│   └── ... (dozens of combinations)
├── src/
│   ├── quantize/
│   │   ├── awq_runner.py
│   │   ├── gptq_runner.py
│   │   └── smoothquant_runner.py
│   ├── engines/
│   │   ├── vllm_runner.py
│   │   ├── trtllm_runner.py
│   │   └── sglang_runner.py
│   ├── metrics/
│   │   ├── perf.py
│   │   └── quality.py
│   └── report.py
├── reports/
│   ├── 2024-08-22_llama2_quant.md
│   └── 2024-08-22_llama2_quant.html
└── cli.py                    # Entry point: python cli.py run --config xxx.yaml

4. Selling Points You Can Talk About ​

  • "I have a complete quantization decision framework": measured AWQ vs GPTQ vs SmoothQuant on Llama-2-7B across throughput, memory, and MMLU, producing a "when to pick which" decision tree
  • "I have cross-engine comparisons": performance of the same AWQ model on vLLM / TRT-LLM / SGLang
  • "I have a reproducible platform": anyone can run your cli.py + config files on their machine and reproduce the same numbers — this is the Inference Benchmarking in Practice report template turned into engineering
  • "I know cheat-proofing": every benchmark separates cold/warm start, handles CPU-GPU overlap, and explicitly states padding
  • "I have visual reports": the HTML report embeds interactive charts (Plotly) that an interviewer can click through on the spot

This is the engineering implementation that turns the theory of Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision into operable decisions, and the "data backend" for Inference Engine Comparison.

6. Project 5: On-Device LLM App ​

1. Motivation ​

On-device LLMs are the hottest direction of 2024–2025 — iOS 18, Android 15, and Windows Copilot+ are all moving LLMs onto the device. On-device deployment challenges are completely different from the cloud: tight memory budget (the iPhone 15 Pro has only 8 GB), battery sensitivity, limited compute, and the need to support multiple SoCs (Apple Silicon / Snapdragon / MediaTek). This project is the most on-target engineering work for on-device / Mobile AI roles.

2. Tech Stack ​

iOS App / Android App
        │
        ▼
   llama.cpp (C++)
        │
        ▼
   GGUF quantized model (Q4_K_M)
        │
        ▼
   Metal (iOS) / OpenCL (Android) backend
        │
        ▼
   Streaming generation + context management
ComponentChoice
Inference enginellama.cpp (see llama.cpp and GGUF)
Quantization formatGGUF Q4_K_M (4-bit, balanced quality and size)
iOS backendMetal Performance Shaders
Android backendOpenCL / Vulkan
ModelQwen2.5-1.5B-Instruct (small enough to run on a phone)
App frameworkSwiftUI (iOS) / Jetpack Compose (Android)
Context managementSliding window + context compression

3. Deliverables ​

on-device-llm-app/
├── README.md
├── ios/
│   ├── LlamaApp/             # SwiftUI app
│   ├── llama.cpp/            # git submodule
│   └── build_scripts/
├── android/
│   ├── app/                  # Jetpack Compose
│   ├── llama.cpp/
│   └── build.gradle
├── models/
│   └── qwen2.5-1.5b-instruct-q4_k_m.gguf   # Not committed; README provides a download script
├── bench/
│   ├── bench_ios.sh          # Benchmark on an iPhone
│   └── bench_android.sh      # Benchmark on Android
└── docs/
    ├── ios_memory_budget.md  # Memory budget analysis per iPhone model
    ├── android_soc_matrix.md # Performance matrix per SoC
    └── battery_test.md       # Battery impact test

4. Selling Points You Can Talk About ​

  • "I run a 1.5B model on an iPhone 15 Pro at TTFT 1.2 s and TPOT 80 ms": real, demonstrable on-device performance
  • "I understand battery optimization": measured X% battery drain over 30 minutes of continuous chat, and designed a "reduce generation speed in low-power mode" strategy
  • "I understand model selection": why 1.5B rather than 7B (phone memory budget), why Qwen rather than Llama (higher quality in Chinese scenarios)
  • "I understand cross-SoC adaptation": performance differences across Apple A17, Snapdragon 8 Gen 3, and Dimensity 9300, with a "detect the SoC at runtime and pick the backend" solution
  • "I understand on-device-specific challenges": memory budget, thermal throttling, battery management, privacy compliance — problems the cloud never has

See llama.cpp and GGUF and Mobile Deployment, and the extreme-case application of "calculate the memory budget first" from Deployment Design Principles.

7. Project Combination Advice ​

You don't need to build every project — going deep on 2–3 beats skimming 5. Three combination suggestions:

Combination A: Cloud Inference Engineer (Most Common) ​

  • Project 1 (RAG service): covers the multi-model pipeline
  • Project 2 (multi-model routing): covers cost optimization and orchestration
  • Project 4 (quantization benchmark platform): covers quantization and decision-making

This trio covers the three core cloud inference skills: end-to-end pipeline, cost optimization, and quantization decisions. You can talk about all three in an interview; see Skills Benchmarking: What to Highlight on Your Resume.

Combination B: On-Device / Mobile AI Engineer ​

  • Project 5 (on-device LLM app): the core
  • Project 3 (streaming chat API): adds cloud experience
  • Project 1 (RAG service): adds pipeline experience

An on-device resume must have project 5, or it won't pass the first screen. Projects 3 and 1 fill the gap because on-device engineers often collaborate with the cloud (hybrid device-cloud architectures) — not understanding cloud inference is a real weakness.

Combination C: Performance Optimization / Accelerator Roles ​

  • Project 4 (quantization benchmark platform): the core
  • Project 2 (multi-model routing): covers system-level scheduling
  • Custom project: CUDA kernel optimization (not covered here; see Kernel Fusion and Custom Kernels)

The core skill for performance roles is "measure and decide," and project 4 is its most direct expression. Add one system-level piece (multi-model routing) and one low-level piece (kernels) to complete a full "measure → decide → low-level" capability chain.

8. The "Standard Kit" Every Project Needs ​

Whichever project you pick, these 5 standard items are non-negotiable — this is Deployment Design Principles applied to a portfolio:

  1. A README written like a paper: motivation, architecture diagram, rationale for technology choices, performance numbers, reproduction commands — miss any one and it fails
  2. A standalone benchmark report: a complete Inference Benchmarking in Practice report template under reports/
  3. A reproducible docker-compose: the interviewer can docker compose up your project in one command
  4. A demo video or screenshots: code alone can't make someone understand your project in 30 seconds — a demo is more persuasive than code
  5. An interview script: write down in advance "if the interviewer asks for a 5-minute walkthrough, what do I say" — see the STAR framework in Interview Question Bank

9. Further Reading ​

References ​