Theme
Resume Analysis: What to Highlight
Remember this one-liner from this page: a resume is not a laundry list of experience — it's a chain of evidence, translating every requirement from the JD List into "here's proof I can deliver." This is especially true for LLM roles: the market is flooded with resumes that say "I know how to call APIs." Resumes that clearly articulate technical challenges, tradeoffs, and quantified results are rare.
All project examples are generic templates
All project examples in this page are template frameworks designed to demonstrate structure and writing style — they do not represent any real person's experience, and you shouldn't copy them verbatim. Copying example projects will be exposed after one round of deep-dive questions.
1. The Essence of a Resume: Chain of Evidence
An interviewer spends minutes reading a resume, and what they're really doing is just one thing: looking for evidence that you match the JD. All resume-writing methods compress into a single chain:
JD Requirement → Capability Claim → Evidence
"Familiar with RAG" → RAG capability → "Built enterprise knowledge-base Q&A system; optimized chunk strategy to improve recall from 68% to 82%"- JD Requirement: from the target role's JD (see JD List);
- Capability Claim: the ability you're claiming to have;
- Evidence: specific facts from projects/experience that support that capability, preferably with numbers.
Among these three, evidence is the only real currency. A claim like "proficient in XX" without evidence is worth zero to an interviewer — and can even be negative (if you can't explain it when they dig deep).
One iron rule from interviewers
Every technical term that appears on your resume is a 10-minute follow-up waiting to happen. Before writing anything down, ask yourself: can I explain this term's principles, tradeoffs, and a real example? If not, delete it or downgrade it to "familiar with."
2. The Golden Structure for LLM Resumes
1. Four-Part Project Structure: Context → Challenge → Solution → Quantified Results
This is the most effective structure for LLM role projects. Each section has a distinct purpose:
| Part | What to Write | What Interviewers Want to Hear |
|---|---|---|
| Project Context | Business scenario, what problem it solves, why it matters | You understand the problem, not just the code |
| Technical Challenge | What was genuinely hard about this project (data, quality, cost, stability) | You've identified real difficulties |
| Solution & Tradeoffs | What you chose, why, and what alternatives you compared | You have judgment and comparative evidence |
| Quantified Results | Model metrics + business metrics + resource cost | You can prove effectiveness; numbers survive follow-up |
2. Explaining One Project Deeply Beats Listing Five Superficially
LLM interviews almost always include a "deep-dive your project" segment: continuous follow-up questions on a single project for 20–40 minutes. It's better to have two or three projects you can defend thoroughly than five projects you barely touched. The test: can you walk through a project from start to finish — why you chose that model/solution, where the data came from, what pitfalls you hit, how you diagnosed them, and how you calculated the quantified results?
3. "What You've Done" vs "How Deep You Can Explain"
In the LLM field, "I've done XX" depreciates fast (open-source tutorials let everyone claim "I've done it"). "How deep you can explain" is the real moat. Three typical tiers:
| Tier | Resume Wording | Interview Performance | Assessment |
|---|---|---|---|
| Used it | "Used GPT-4 to implement XX feature" | Can only say "called the API, tried a few prompts" | Hits every candidate; no differentiation |
| Built it | "Built a RAG Q&A system with 75% recall" | Can explain the flow and answer detail questions | Passes the bar, but homogenous |
| Mastered it | "Compared BM25/vector/hybrid retrieval, chose hybrid + reranking; identified chunk boundary issues causing recall failure, switched to chapter-based splitting, improved recall from 68%→82%, reduced errors by 31%" | Can explain tradeoffs, pitfall stories, numbers survive follow-up | Has a moat, makes it to final rounds |
To move from "built it" to "mastered it": add three things to every project — comparison experiments (you chose A, why didn't B work), failure stories (which approach failed and how you found out), and quantified results (one number each for quality, cost, and latency).
4. How to Write Five Types of Projects
The five types below are the most common project categories from the Knowledge Breakdown. Each gives "how interviewers will ask" + "good writing example."
1. Fine-tuning Projects
How interviewers will ask: Why LoRA over full-parameter fine-tuning? How did you choose rank r and α? Where did the training data come from, and how many examples? Any overfitting? Did the model get dumber after fine-tuning, and how did you detect and fix it? How do you prove the fine-tuning was effective?
Weak example: Used LoRA to fine-tune Llama-3-8B, improving customer service intent recognition accuracy to 90%.
Strong example:
Intent Classification Fine-tuning (LoRA)
Context: Rule-based + prompt approach achieved only 71% accuracy on long-tail intents, creating high user churn risk.
Challenge: Training data was only 2,000 examples, prone to overfitting; base model performed poorly on Chinese business corpus.
Solution: LoRA fine-tuning on Qwen base (r=16, α=32); built instruction-response dataset;
froze base model, trained only low-rank increments; early-stopped using 10% validation set.
Results: Long-tail intent accuracy improved from 71%→86% (+15 points); general capability
benchmarks remained flat before and after, proving no catastrophic forgetting;
cost ≈30 RMB per epoch on a single A100.Why it's good: Context → challenge → solution (with hyperparameters and rationale) → quantified results (model metrics + forgetting detection + cost). Every sentence opens a door for deep-dive questions.
2. RAG Projects
How interviewers will ask: Why RAG over fine-tuning? How did you chunk? How big? What retriever did you use, and why? How do you evaluate retrieval quality? When the answer is wrong, is it a retrieval problem or a generation problem? How do you diagnose it?
Weak example: Implemented RAG knowledge-base Q&A with a vector database, achieving 85% answer accuracy.
Strong example:
Enterprise Knowledge Base Q&A (RAG)
Context: Customer service knowledge base had 20K documents, updated frequently; fine-tuning couldn't keep up, requiring external knowledge injection.
Challenge: Documents contained many tables and long paragraphs; naive chunking caused fragmented recall; answers frequently hallucinated citations.
Solution: Semantic paragraph chunking (500 tokens, 50-token overlap) + hybrid retrieval (BM25+vector)
with reranking; enforced citation in generation; built 300-question golden set.
Results: Retrieval recall improved from 68%→82%; answer faithfulness (citation verifiability) reached 93%,
reducing hallucination rate by 41% vs. baseline; avg. first-token latency 1.2s.Why it's good: Shows judgment on "RAG vs fine-tuning," retrieval optimization details, and crucial evaluation systems (golden set, faithfulness metric).
3. Agent Projects
How interviewers will ask: How is the Agent loop designed? How do you ensure tool calling produces correct formats? What happens when the Agent freezes or loops? How is memory stored? How is cost controlled? How do you prevent prompt injection?
Weak example: Developed a data analysis Agent based on LangChain that supports SQL tool calling.
Strong example:
Data Analysis Agent (Tool Calling)
Context: Business users had to wait for data warehouse scheduling to query data; wanted natural language to data access.
Challenge: SQL generation accuracy was unstable; multi-turn conversations produced extremely long context and tool results; injection risks existed.
Solution: ReAct loop + tool schema validation (auto-retry invalid parameters once);
compressed critical context into summaries; SQL whitelist + read-only account for permission isolation.
Results: SQL generation accuracy improved from 81% (baseline 62%); avg. 3.2 tool calls per task,
auto-recovery covered 87% of failures; token cost ≈0.08 RMB per task.Why it's good: Demonstrates reliability and cost awareness — exactly what Agent roles in 2025 are looking for.
4. Evaluation Projects
How interviewers will ask: How did you build the evaluation set? Why these benchmarks? How did you handle LLM-as-a-judge bias? How do you prevent data contamination? How do you validate in production?
Weak example: Responsible for model evaluation, using MMLU and other benchmarks.
Strong example:
Business Evaluation System for LLMs
Context: Models iterated frequently; the team relied on manual spot-checking to judge quality, making regression impossible.
Challenge: No standard answers for business tasks; needed to balance objective metrics and subjective quality; prevented benchmark contamination.
Solution: Built 3 evaluation sets (500 questions from public benchmarks + 800 custom business questions + human standard set),
introduced LLM-as-a-judge (dual-model cross-evaluation + position shuffling to eliminate bias), integrated into CI regression.
Results: Each post-deployment iteration produced evaluation reports within 20 minutes, catching 6 quality regressions;
human-model agreement reached 89%, approaching the ~90% baseline of human-human agreement.Why it's good: Evaluation projects are the scarcest — clearly describing your methodology (bias prevention, contamination prevention) is itself a differentiator.
5. Deployment Projects
How interviewers will ask: How did you estimate memory? Why this quantization bit-width? How did you handle batching? What were your TTFT/throughput numbers? How did you stress test? What happened when you hit OOM?
Weak example: Deployed the model with vLLM, improving QPS by 3x.
Strong example:
Inference Serving & Optimization
Context: Business needed 100 QPS for online inference; the original service could only support 20 QPS per instance.
Challenge: KV cache memory exploded in long-context scenarios; FP16 weights + cache exceeded available memory.
Solution: INT8 weight quantization (AWQ, <0.5-point accuracy loss) + vLLM continuous
batching + KV cache upper-bound constraints with PagedAttention reuse;
stress testing showed the bottleneck shifted from GPU compute to memory bandwidth.
Results: Throughput per instance improved from 20→110 QPS (~5.5x), P95 latency <2s,
cost per token reduced by ~60%; accuracy regression was confirmed via evaluation-set regression.Why it's good: Elevates from "deployed it" to "optimization chain + quantization selection + bottleneck identification + cost numbers."
5. Project Deep-Dive Self-Test: Six Categories of Follow-up Questions
"Project deep-dive" in LLM interviews isn't casual conversation — it's six categories of questions used to repeatedly pressure-test a single project to see if you can handle it. After writing each project, use the table below to self-test:
| Question Type | Typical Questions | What "Can't Answer" Looks Like | How to Prepare |
|---|---|---|---|
| Motivation | Why RAG over fine-tuning? Why this base model? | "Everyone else does it that way" | Write a one-line rationale + one counterexample |
| Detail | How big is your chunk? How much overlap? Why? | Can't give specific numbers | Write down your key hyperparameters and rationale |
| Data | Where did the data come from? How many examples? How was it cleaned? | "Found it online" | Write data source, scale, and processing pipeline |
| Tradeoffs | Did you try other approaches? Why did you abandon them? | "Never tried anything else" | Run one comparison experiment, even small-scale |
| Evaluation | How do you prove it works? What's the baseline? Where's the eval set from? | Just says "90% accuracy" | Write baseline, metrics definition, and eval set composition |
| Cost | How much does training/inference cost per run? What's the latency? | "Never calculated it" | Estimate at least an order of magnitude (RMB/day, seconds) |
The three project questions that most commonly trip people up (with example answers):
- "What's your baseline for this approach?" — Example: "Base model with direct prompting, 71% accuracy on the same 300-question business set; improved to 86% after fine-tuning."
- "Is there data leakage? Could your eval set overlap with training data?" — Example: "We ran MinHash approximate dedup between the training set and the eval set, identified and removed XX overlapping samples; the eval set was split chronologically, later than the training data."
- "If you had to redo this project, what would you change?" — Example: "We built the eval set too late. The first two iterations were judged entirely by subjective opinion. If I redid it, I'd build the golden set on day one and integrate it into CI for regression."
Create a "Project Health Checklist" from the six question types
For your two strongest projects, write a one-page "Project Profile" each: one-line context, three challenges, a solution tradeoff table, evaluation methodology, cost numbers, and one failure retrospective. Read only these profiles in the three days before interviews — they're your deep-dive defense line.
Project Profile Template (ready to use):
text
Project name: ____________________
One-line summary: ____________________
Business context (3 lines): ____________________
Three technical challenges: ①____ ②____ ③____
Solution selection & rationale (including abandoned options): ____
Data source & scale: ____
Evaluation methodology (eval set composition / baseline / metrics): ____
Quantified results (model metrics / business metrics / cost & latency): ____
Pitfalls learned & retrospective: ____6. How to Write Quantified Results
1. Model Metrics + Business Metrics: You Need Both
| Type | Examples | Why It Matters |
|---|---|---|
| Model metrics | Accuracy, recall, pass@k, faithfulness, hallucination rate | Proves "the technique itself works" |
| Business metrics | Conversion rate, complaint reduction, headcount saved, user retention | Proves "it has value to the business" (what interviewers value most) |
| Resource metrics | Latency, throughput, token cost, GPU memory | Unique differentiator for LLM roles: proves "it runs and is affordable" |
2. The Honesty Principle: Numbers Must Survive Three Follow-up Questions
"How did you measure that? How large was the sample? What was the baseline?" LLM interviewers are naturally suspicious of claims like "improved by 15 points." Three rules for resume writing:
- Write the baseline: "71%→86%" is ten times more credible than just "86%";
- Write the methodology: note the eval set size and composition ("300-question business golden set");
- Don't fabricate: It's better to leave a number out than to write one that collapses under one follow-up question. Interviewers will spot fabricated numbers on the spot, and one exposed fabrication zeroes out the credibility of your entire resume.
3. Three Sources of Project Numbers and Their Credibility
| Source | Example | Credibility | Interview Follow-up |
|---|---|---|---|
| Comparison experiments you ran yourself | "Same eval set, baseline 71% → fine-tuned 86%" | High — most recommended | "Where did the eval set come from? How many samples?" |
| Team/production data | "Customer complaints dropped ~30% after deployment" | Medium — must be able to explain methodology | "How is this metric defined? What's the time window?" |
| Cited from official/public benchmarks | "This model's public MMLU score is 90" | Low (not your achievement) | Can only be background context, not a personal accomplishment |
The most common "number disaster" on LLM resumes is the third type: writing a model's public score as your own achievement. Public benchmark scores belong to the model vendor, not to you — your achievements must be things you personally measured. It's acceptable to write "Used this model (public MMLU 90) to achieve XX on a business set." Writing "MMLU 90" as your own achievement will collapse after one round of questioning.
7. How to List Skills: Tier Your List, Don't Stack Nouns
For LLM roles, we recommend organizing skills into four layers. Fewer is better:
| Layer | Content | Example |
|---|---|---|
| Programming languages | Python (proficient), C++ (familiar), SQL (skilled) | Write proficiency level, not just the name |
| Frameworks & tools | PyTorch, HF Transformers, vLLM, LangChain | Should align with projects; you should be able to explain principles |
| Model technologies | Transformer, LoRA, RAG, RLHF/DPO, quantization | Should align with projects; you should be able to explain tradeoffs |
| Engineering skills | Distributed training, evaluation systems, serving, CI/CD | Bonus items; only write if you have evidence |
Three pitfalls in skill lists
① Stacking 30 nouns — interviewers assume you know all of them and pick one at random for a deep-dive, causing you to crumble; ② Saying "familiar with" but having no project to back it up — disconnected skills are as good as not listed; ③ Listing outdated tools (e.g., in 2025, only mentioning "familiar with LangChain" without any deeper understanding) — exposes stale knowledge.
8. Common Resume Mistakes: 10 Anti-Patterns
- Only listing courses, no projects: LLM roles are engineering roles — without projects, you have no evidence;
- Stacking nouns without explanation: "Proficient in Transformer, RAG, LoRA, Agent" — none of them survive follow-up;
- Numbers that don't survive questioning: wrote "95% accuracy," but can't answer "where's the eval set from?";
- Listing "did a demo" as "owned a system": the gap between a demo and production exposure in 10 minutes of deep-dive;
- Only listing tools, no judgment: "Used LangChain" vs "compared LangChain with a custom pipeline and chose the latter for controllability";
- No failures or retrospectives: LLM projects always have pitfalls — only sharing successes signals lack of depth;
- Projects disconnected from target JD: applying for an Agent role but your resume's lead project is image classification;
- Writing team projects as personal projects: vague when pressed on "what exactly did you do yourself?";
- Only model metrics, no business or cost metrics: LLM roles especially value cost awareness;
- One resume for all roles: not even changing the role title — HR spots this instantly.
9. Customizing Your Resume by Role: Different Emphasis
For the same experience, when applying to different roles, highlight different evidence. Customize your resume's title and first bullet point by cross-referencing the table below (role emphasis drawn from the JD List keyword radar):
| Role | Evidence to Highlight Most | Common Pitfall |
|---|---|---|
| Algorithm Engineer (LLM) | Full production chain (selection → composition → evaluation → deployment) and tradeoff reasoning | Only writing "improved results" without "why this approach was chosen" |
| NLP Algorithm Engineer | Classical NLP foundations (tokenization/sequence labeling) + LLM composition | Only listing LLM projects, losing evidence of foundational skills |
| LLM Training Engineer | Distributed configuration, loss curve diagnostics, data pipeline scale | Only writing "ran training," without scale or problem diagnosis |
| Inference Optimization Engineer | Specific quantization/KV cache/batching numbers, bottleneck identification process | Only writing "deployed vLLM," without before/after comparison |
| AI Application Engineer | End-to-end delivery + evaluation system + cost/latency | Only has a demo, no evaluation or cost awareness |
| Agent Engineer | Reliability (retries/guardrails/injection defense), cost control | Only listing framework names, no stability design |
| Evaluation Engineer | Eval set design, bias/contamination prevention methods, regression-catch examples | Only writing "tested on MMLU," no methodology |
| Data Engineer (LLM) | Corpus scale, dedup/filtering methods, quality closed-loop | Only writing "cleaned data," no methods or numbers |
One experience, different "evidence angles" by role
The same fine-tuning project:
- For algorithm roles → emphasize "why LoRA over full-parameter, what experiments you did with rank and data mixing";
- For application roles → emphasize "how the eval set was designed, cost and latency after deployment";
- For training roles → emphasize "how memory was saved, how training curves were monitored."
This isn't fabrication — it's presenting different facets of the same facts, aimed at different role radars.
We recommend maintaining two versions of your resume: one for algorithm/training roles that highlights model selection, principles, and experiments; one for application/engineering roles that highlights engineering closure, evaluation, and cost. Don't send the same resume to every role — the first thing HR looks for is "keywords aligned with the role radar."
10. Self-Assessment Checklist
Check each box before submitting. If any ☐ is unchecked, go back and revise:
Project Layer
- [ ] Your 1–2 strongest projects can each be explained for 20+ minutes (including tradeoffs and failures)
- [ ] Every project follows the Context → Challenge → Solution → Quantified Results four-part structure
- [ ] At least one project has "comparison experiments" or "rationale for approach selection"
- [ ] Quantified results include baselines and methodology; you can answer "how was it measured, what was the sample size?"
- [ ] LLM projects cover at least: quality + cost/latency (at least one of each)
Skills Layer
- [ ] Every noun on your skills list can be explained for 10 minutes of principles
- [ ] Skills cross-validate with projects; no orphaned nouns
- [ ] Aligned with the target JD's keyword radar (see JD List)
Language Layer
- [ ] No empty claims of "proficient in" without evidence
- [ ] No typos, consistent formatting, one page
- [ ] Has GitHub/portfolio links that are real and accessible
Alignment Layer
- [ ] Per Knowledge Breakdown self-assessment, no "conceptually fuzzy" ratings for abilities claimed on your resume
- [ ] Per Interview Questions self-test, can answer 80%+ of project-related deep-dive questions
11. Further Reading
Continue within the site
- Knowledge Breakdown — every noun on your resume has deep-dive exam questions here
- Interview Questions — reference key points for project deep-dive questions
- JD List — align your resume with each keyword in the radar
- Module Overview & Career Landscape — confirm which role's story your resume is telling
References
- GitHub — home for portfolio and project repos; LLM projects should include a README and demo
- Hugging Face — model/dataset/space showcase; the best platform for fine-tuning and evaluation project portfolios
- LinkedIn — primary platform for overseas job-hunting; English JDs and industry connections
- BOSS Zhipin — primary domestic job platform; directly compare resume keywords with JD keywords