Appearance
What Your Resume Should Highlight
One-sentence definition: A resume is not a laundry list of experiences—it is a chain of evidence aimed at the target role, using verifiable, quantifiable, question-proof facts to prove "I possess every capability this role requires." In trending AI roles, this reads more literally than in any other field: in the 2025 resume pool, nearly every document says "familiar with large language models," "built RAG," or "understands agents"—nouns alone can no longer distinguish anyone.
Think of a resume as evidence being examined in court: the judge (the interviewer) only believes statements backed by evidence. You say "I'm proficient with LLMs"—where's the evidence? You say "I have RAG project experience"—where's the evidence? You say "I've solved hallucination problems"—where's the evidence? Unsupported claims are, at the AI-role screening stage, roughly equivalent to writing nothing. And AI adds one more twist: the tech stack turns over every six months, everyone is learning the new buzzwords, and interviewers have extremely little patience for being bombarded with nouns—their fastest way to verify a candidate is to pick one term and drill it all the way down.
This page is the third stop in the career module: the first two stops, the JD checklist ("what the market is hiring for") and the capability map ("what you can and cannot do"), feed into this one, whose job is to fuse the gap between the two into a resume—rewriting the experience you already have through an interviewer's eyes, not inventing from scratch. Everything here argues with evidence. Start from a brutal premise: the interviewer assumes everything on your resume may be fabricated, and your job is to make every line survive follow-up questions.
1. Common ways resumes die in trending AI roles
The tech direction keeps changing; the failure modes barely do. The six below are ordered by kill power, and each one leads directly to being filtered out or crashing in the first interview round.
| Failure mode | Typical symptom | Why it's an instant kill |
|---|---|---|
| "Proficient in LLMs" with no projects | A skills line reading "proficient in LLMs, familiar with RAG, aware of agents" while project experience is nearly blank | Unverifiable; to an interviewer, "proficient in LLMs" ≈ "can use ChatGPT," which is not a capability claim at all |
| Stacking tool names | Fifteen nouns in one line: LangChain, LlamaIndex, FAISS, vLLM, Dify, AutoGen... | One gets punctured and the credibility of the whole page drops to zero; tool names ≠ ability—tools get swapped and obsoleted, principles don't |
| No quantified results | "Built a RAG system," "improved answer quality," "worked pretty well" | No numbers, no comparison baseline, no measurement definition—equivalent to saying nothing |
| Tutorial clones passed off as projects | Reproducing an official quickstart or tutorial demo verbatim, README included by copy-paste | Three layers of follow-ups expose it: why chunk it this way? Would another vector store work? Did it ever fail? |
| Responsibilities without decisions | "Responsible for the retrieval module," "participated in prompt design," "assisted with model tuning" | Your judgment and contribution are invisible; it reads pure name-riding, and the interviewer has nothing to dig into |
| A resume that targets no role | One resume mass-mailed to algorithm, application, and product openings alike | The JD keywords don't match, so keyword filters cut you at screening; even if you pass, the interview turns into answers that miss the question |
A discipline that matters more than the six above
AI tech stacks update extremely fast, but the interviewer's verification method has not changed: pick one term and drill it to the bottom. Every noun you put on the resume should be ready to take three layers of questioning: "explain the principle → explain the project → explain the trade-offs." Better five terms you can talk about for 20 minutes than ten you're shaky on. The price of stacking nouns is "one gets punctured and the whole page resets to zero"—especially lethal in an AI job market where information is highly homogeneous.
2. The four-quadrant capability benchmark: how to write resume evidence
The capabilities behind trending AI roles (algorithm, application, product) split into four quadrants. Interviewers verify each quadrant differently, so the form of evidence you present on the resume differs too. Get the whole picture from the table first, then go quadrant by quadrant to see how the evidence is written.
| Quadrant | What is tested | How interviewers verify | Related pages on this site |
|---|---|---|---|
| Model theory | Transformers and attention, pretraining/fine-tuning principles, able to explain the "why" | Principle probing, formula derivations, architecture comparisons | Transformers and attention, Large language models (LLMs) |
| Engineering ability | Shipping end to end: RAG pipelines, agent tool calling, deployment and tuning | Project deep-dives, system design questions | Build a RAG app from scratch, Build an agent from scratch, Deployment and inference optimization in practice |
| Evaluation mindset | Can design evaluations, isn't fooled by "feels better," can quantify the gains | Follow-ups like "how did you evaluate this," metric definitions | Build an LLM eval suite, LLM evaluation and benchmarks |
| Domain understanding | Knows the business scenario and the user's problems, knows the model's limits and costs | Scenario questions, case discussions, trade-off reasoning | Concept boundaries, Perplexity and AI search, and other case pages |
1. Model-theory quadrant: replace "I studied it" with "mechanism-based decisions"
The evidence is not "I studied Transformers" but "I made mechanism-based choices inside a project and lived with the consequences." For example:
- ❌ "Familiar with attention mechanisms."—everyone writes this; it cannot be verified.
- ✅ "In a long-document retrieval-and-reading scenario, compared sliding-window attention against full attention: sliding window saves memory but loses cross-segment dependencies; measured ROUGE dropped 6%, so we kept full attention and added retrieval-first input truncation."
The subtext of the second one: "I can walk you through, right now, why multi-head attention splits into heads, why the attention formula divides by √d_k, and how the KV cache saves compute." All of these derivations are collected in Transformers and attention and Large language models (LLMs)—catch up there.
2. Engineering quadrant: the evidence is "a complete pipeline plus performance numbers"
Engineering evidence is an entire pipeline you can narrate, plus end-to-end numbers. A writing template:
30,000 documents → cleaning (tables/scans converted to text) → chunking (by heading + 500-token overlap) → Chinese-specialized embeddings → hybrid retrieval (vector store + BM25) → reranking → generation (faithfulness constraints) → serving (P95 latency 2.1s) → monitoring and alerting.
Every link in that chain can be probed, and every number has a definition. For the full hands-on path see Build a RAG app from scratch and Build an agent from scratch; for deployment and throughput/latency tuning see Deployment and inference optimization in practice.
3. Evaluation quadrant: most resumes' blank spot and the highest-ROI differentiator
The evaluation quadrant is the one with the fewest credible candidates and the strongest proof that "you really did it." The reason: someone who has never run a system—and never watched one go sideways—cannot write evaluation details at all. How to write the evidence:
Built a 120-item after-sales eval set (real user questions + labeled answers), scored on three dimensions: faithfulness / relevance / completeness; LLM-as-judge agreed with human spot checks 89% of the time, and divergent samples triggered human review and a revision of the scoring rubric.
That passage proves three things at once: you have an evaluation system, you validated the evaluation tooling itself, and you are willing to spend labeling budget on metrics. For the methodology see Build an LLM eval suite and LLM evaluation and benchmarks.
4. Domain-understanding quadrant: spell out "what the problem is, why it's worth solving, what the constraints are"
Domain evidence on a resume is not "familiar with e-commerce / finance / healthcare"—it is awareness of business constraints. For example:
Customer support: a wrong answer costs complaints and legal exposure, so the design puts faithfulness above fluency—when retrieval misses, the bot is allowed to say "not in the knowledge base, escalating to a human" instead of letting the model improvise.
That one sentence decides every technical trade-off that follows, and it is also where interview scenario questions hand out top marks. To get a feel for how different domains trade off differently, read ChatGPT and conversational AI, Perplexity and AI search, and Manus and agent applications to see "how product goals dictate technical choices."
How to allocate space across the four quadrants
Don't spread the four quadrants evenly—rank them by the target role (see Section 5): algorithm roles should thicken model theory and evaluation, application roles should thicken engineering and evaluation, product roles should thicken domain understanding and evaluation. Note that "evaluation mindset" is a plus for nearly all three role types, because it proves engineering maturity and product judgment at the same time.
3. The project description formula: context - action - result
Project experience is the load-bearing wall of a resume. The master rule in one sentence: context (the business problem + data scale + constraints) → action (what you did + why you chose it + what went wrong) → result (model metrics + business metrics + a comparison baseline). Missing any of the three, and the interviewer will skip the project or drill it into a corner.
1. Before vs. after: the same RAG project wearing two faces
Demonstrated on the same fictional project (Q&A over internal after-sales manuals):
❌ Before (a tool laundry list—every sentence is "what was done"):
An intelligent Q&A system built on LangChain. Uses FAISS as the vector database and calls the GPT-4 API. Implemented document loading, text chunking, vector retrieval, and answer generation. Worked pretty well.
The problem list for this version: no context (where did the data come from, how many documents, why does this project exist); actions with tool names but no judgment (why FAISS and not Milvus? why chunk this way?); "worked pretty well" with zero numbers; and, most damning—anyone who has read a tutorial could write this description word for word, so it proves nothing about you versus the next candidate.
✅ After (a chain of evidence, all three parts present):
Intelligent Q&A over internal after-sales manuals—turned 30,000 PDFs into a knowledge base support agents can actually use.
- Context: support handled 2,000+ repetitive inquiries per day and each manual lookup averaged 8 minutes; the PDFs were full of tables and scans, and feeding them straight into the model performed poorly.
- Action: compared three chunking strategies—by heading / recursive character / fixed length—and measured that "by heading + 500-token overlap" gave the best retrieval recall; converted table rows to markdown before ingestion, fixing table-retrieval losses; reranked the Top-20 recalled items down to Top-5 with a cross-encoder; added faithfulness constraints on the generation side so that on a retrieval miss the bot answers "not in the knowledge base, escalating to a human."
- Result: Top-5 retrieval hit rate 78%→91%; answer acceptance rate in support spot checks 61%→87%; average handling time per inquiry down from 8 minutes to 3. Pitfall: an early embedding model mismatched to the document language cratered recall; switching to a Chinese-specialized model fixed it.
In the after version, every action carries a "why," every number has a definition, and every pitfall can be expanded into a story—exactly the material interviewers want to dig into. The "pitfall" line matters most: someone who never did the work cannot invent a real pitfall, which makes it the strongest signal of "I actually did this."
2. Two iron rules for metrics
- Carry both model metrics and business metrics: "91% retrieval recall" alone is incomplete; add "support handling time cut by 62%" and the business value becomes provable. See the "evaluate at both the technical and business layers" idea in LLM evaluation and benchmarks.
- Always include the baseline and the measurement definition: "95% accuracy" carries no information; "acceptance rate up from 61% to 87% (spot-checked by 30 support agents on 200 items, same rubric as production)" does. Before writing "improved 300%," know exactly what it is relative to—interviewers will ask.
4. How to build personal projects from zero
With no internships or work experience, personal projects are the only source of evidence you fully control. The recommended path follows this site's practice pages in "one progressively deepening main line"—each step reuses the previous step's output, ending with four projects whose evidence corroborates one another:
- Build a RAG app from scratch (Build a RAG app from scratch): pick a real domain (technical docs you have actually read, an open-source project's README, policy documents) and build retrieval Q&A. This is the starting project, covering the complete engineering pipeline.
- Build an agent from scratch (Build an agent from scratch): add tool calling on top of the RAG app (query a database, call APIs, execute code) and make an application that can do things rather than just talk.
- Fine-tune your own LLM (Fine-tune your own LLM): fine-tune an open-source model with LoRA to produce a "domain style / output-format alignment" case, and absorb the theory in Fine-tuning and PEFT (LoRA) along the way.
- Build an LLM eval suite (Build an LLM eval suite): build eval sets and metrics for the projects above, quantifying "RAG on vs. off" and "before vs. after fine-tuning"—this step turns your projects from demos into evidence.
Why this order: RAG provides engineering evidence, the agent provides evidence of system complexity, fine-tuning provides theory evidence, and evaluation serves all three while being the single most probed topic in interviews. One project that is "both deep and complete" beats five things that never got past demo level.
Three criteria for picking a topic
(1) It's a problem you would genuinely hit yourself (exam-prep material Q&A, a fitness knowledge base, an assistant for your own blog) and you can explain "why build it, who benefits"; (2) the data can be legally obtained and is large enough to feel like real work; (3) you can define "good"—that is, you can build an eval set. Meet these three and the project becomes ammunition for the "project deep-dive" round of the interview question bank.
5. How resume emphasis differs by role
| Role type | Quadrants to lead with | Form of evidence | Related pages on this site |
|---|---|---|---|
| Algorithm roles (model/research) | Model theory + evaluation mindset | Fine-tuning comparison experiments, eval set construction, technical blogs / experiment logs | Fine-tuning and PEFT (LoRA), Alignment: RLHF and DPO, LLM evaluation and benchmarks |
| Application roles (engineering) | Engineering ability + evaluation mindset | End-to-end RAG/agent projects, performance metrics, pitfall retrospectives | Build a RAG app from scratch, Build an agent from scratch, Deployment and inference optimization in practice |
| Product roles (AI product) | Domain understanding + evaluation mindset | Problem definition, user-metric gains, trade-off retrospectives | Perplexity and AI search, Manus and agent applications, Prompt playbook |
1. Algorithm roles: prove depth with "why" and trade-offs
The deep-dive round for algorithm roles is mostly "principle probing plus experiment retrospective." What you can do at the resume stage is leave depth hooks inside your projects that invite probing:
- Spell out the comparison and trade-offs behind your model/fine-tuning choices: "tried LoRA ranks 8/16/32; 16 gained the most on the eval set, 32 showed diminishing returns and doubled memory usage";
- Spell out the design of your evaluation: why this metric, how the eval set was built, how data leakage was prevented;
- Spell out your understanding of the mechanisms: what long context, positional encodings, and attention variants actually did in the project.
The subtext of these hooks: "I can talk about this live for 30 minutes, and I have the experimental data." For theory foundations see Transformers and attention and Large language models (LLMs).
2. Application roles: prove breadth with "a complete pipeline and stability"
Application roles do not expect you to derive the latest attention variant, but they do expect you to turn a model into a reliable system. Highlight:
- Full-pipeline coverage (data ingestion → cleaning → retrieval → generation → deployment → monitoring);
- An understanding of "stable" (version management, metric monitoring, graceful degradation on failure, rollback plans);
- Performance numbers: P95 latency, throughput, metric comparisons before and after quantization—see Inference optimization and quantization and Deployment and inference optimization in practice;
- For a self-check on engineering pitfalls, compare against Common pitfalls and anti-patterns.
3. Product roles: prove judgment with "problem definition and a closed value loop"
Product roles highlight "who the users are, how badly it hurts, how you define 'good,' and which metrics move after launch":
- Spell out where requirements came from and how you prioritized (which problems are worth solving with AI, and which are not);
- Spell out how "good" is defined and measured (acceptance rate, human-handoff rate, hours saved—not empty phrases like "high answer quality");
- Spell out the solution trade-offs: why RAG instead of fine-tuning, why ship a minimal viable version first and iterate.
You can learn this style of narrative from the analysis of "how retrieval + generation drive product decisions" in Perplexity and AI search and "how a conversational experience gets productized" in ChatGPT and conversational AI.
One governing principle
Depth comes from drilling one point to the bottom; breadth comes from carrying one line all the way through. At least one project should reach the mechanism/trade-off layer, while every project covers the complete pipeline. One project that is "both deep and complete" beats five that all stop at demo level—this principle holds for all three role types.
Further Reading
Keep reading on this site:
- See the role clearly before writing your resume: career module guide → JD checklist → capability map
- The next gate after your resume passes screening: interview question bank
- Putting the four quadrants into practice: Build a RAG app from scratch, Build an agent from scratch, Fine-tune your own LLM, Build an LLM eval suite
- Concept catch-up: Transformers and attention, Large language models (LLMs), Retrieval-Augmented Generation (RAG), LLM evaluation and benchmarks
- Engineering pitfall self-check: Common pitfalls and anti-patterns
- Job-search main line and pacing: Learning paths; one unified terminology standard: Glossary
References
- Harvard OCS Resumes and Cover Letters — the Harvard career center's resume guide, an authoritative reference for format and wording
- Google Machine Learning Crash Course — a free course, good for shoring up fundamentals and engineering intuition before interviews
- Awesome README — a classic list of README writing examples, useful as a quality bar for portfolio repos
- The STAR Method (Indeed Career Guide) — the canonical source of the STAR methodology
- Levels.fyi — overseas leveling and salary data, for sizing up target roles
One last thing: in trending AI roles, the most valuable sentence on a resume is "the system I built moved the metric from X to Y"—everything else is a footnote to it. Time spent on a resume goes into "thinking through what you did, how well it went, and why you did it that way"—an investment that never loses.