Appearance
JD Knowledge Breakdown
You've got the job description (JD) of a company you'd love to join. It reads "deep expertise in the Transformer architecture," "familiar with RAG," "hands-on experience shipping agents," "knowledge of LoRA fine-tuning," "understanding of inference optimization"... You can read every line, yet every line is wrapped in fog: how well do you actually need to know this, and how will the interviewer test it?
What this article does is translate the "skill words" in a JD into a knowledge-point checklist you can review item by item—not vague advice like "go learn large models," but concrete test points such as "can draw the self-attention computation flow," "can explain how to chunk documents in RAG and when to bring in reranking," and "can articulate why LoRA works with so few parameters." Combined with the five-tier self-assessment table and the method for generating a study checklist at the end, you can run a complete "skills health check" in one afternoon and walk away with a learning roadmap of your own.
Get the JD ──① Extract skill words──② Map to test points──③ Five-tier self-check──④ Generate a study checklist──⑤ Weekly re-test
│ │ │ │
Proficient/Familiar/Aware Sections 2–11 of this article Section 12 Section 13This article is for candidates applying to roles such as RAG Application Engineer, Agent Engineer, LLM Algorithm Engineer, or AI Infra Engineer—as well as every learner who wants to turn "I feel like I know all of this" into "I can explain all of it."
1. What JD Skill Words Actually Test
A JD is not a course syllabus. It is the company's minimum expectation of "what you can independently deliver." The recruiter's subtext reads:
Writing "familiar with RAG" is not asking you to recite RAG concepts from memory. It is asking you to take an enterprise knowledge-base Q&A requirement and independently handle approach selection, hands-on build-out, quantitative evaluation, and failure troubleshooting.
So the first step in breaking down a JD is mapping each skill word to the capability the interviewer will actually probe. Clear away three common misconceptions first:
- Misconception 1: Treating the JD as the complete list. Of the 10 requirements on a JD, the company usually has only 3 hard gates in mind (the ones that decide whether your resume survives screening); the rest paint the picture of an "ideal candidate." Patch the gate items first.
- Misconception 2: Reading "familiar" as "aware." When an interviewer writes "familiar with X," they are testing whether you can explain the principles, implement it hands-on, and articulate the trade-offs—not whether you've "heard of X." When self-assessing, treat "can explain it thoroughly" as the passing bar for "familiar."
- Misconception 3: Reviewing single points but not combinations. Interviews almost never test one point at a time: you often get a scenario ("Build knowledge-base Q&A for an enterprise with 100,000 PDFs—how do you design it?") that demands RAG (architecture) + vector retrieval (recall) + evaluation (quality) + inference optimization (cost) all at once. So the ten sections below have to be read as one connected whole.
2. Master Table of the Ten Knowledge Points: Keyword → Concept → Test Point → Page on This Site
This is the single most important mapping table on the site. The left column holds high-frequency JD keywords; the right column holds the matching in-depth pages.
| JD Keyword | Core Concept | How the Interview Tests It (Typical Questions) | In-Depth Pages on This Site |
|---|---|---|---|
| Transformer / attention / positional encoding | Self-Attention, QKV, multi-head, positional encoding | "What are Q/K/V? Why scale by √d_k? Why is positional encoding needed?" | Transformer and Attention |
| LLM principles / pre-training / emergence | Next-token prediction, Scaling Laws, emergence | "Why can next-token prediction learn language? How do you make sense of emergence?" | Large Language Models (LLM) |
| Prompts / CoT / structured output | Prompt design, chain of thought, output constraints | "What problems do few-shot and CoT each solve? How do you guarantee structured output?" | Prompt Engineering |
| RAG / vector retrieval / reranking | Indexing, chunking, recall, rerank | "How do you troubleshoot a failure in any RAG stage? When is rerank a must?" | Retrieval-Augmented Generation (RAG) · Vector Databases and Semantic Search |
| Agent / tool calling / ReAct | Task planning, Function Calling, loops | "What do you do when an agent gets stuck? How do you design the tool-calling schema?" | AI Agents |
| LoRA / fine-tuning / data engineering | PEFT, low-rank adaptation, instruction data | "Why is LoRA effective with so few parameters? Where does fine-tuning data come from?" | Fine-Tuning and PEFT |
| RLHF / DPO / alignment | Reward model, PPO, preference optimization | "How do you train the RLHF reward model? What does DPO simplify away from RLHF?" | Alignment: RLHF and DPO |
| Quantization / inference optimization / deployment | INT8/INT4, KV Cache, serving | "Why does quantization speed things up? Where does the accuracy loss come from? What is the KV Cache for?" | Inference Optimization and Quantization · Deployment and Inference Optimization in Practice |
| Evaluation / benchmarks | Eval sets, human vs. model evaluation, metrics | "How do you evaluate generation quality? How do you quantify hallucination? How do metrics map to the business?" | LLM Evaluation and Benchmarks |
| Multimodal / diffusion | Vision-encoder alignment, diffusion process | "How do you align a vision encoder with an LLM? How do you speed up diffusion sampling?" | Multimodal Models · Diffusion Models and Generative AI |
The ten sections below expand, one by one, each knowledge point's likely follow-up questions and recommended reading pages. For the overview and the boundaries between concepts, start with What Are the Hot AI Concepts and Concept Boundaries.
3. Transformer / Attention / Positional Encoding
JD wording: "In-depth understanding of the Transformer architecture." This is the foundation question for every LLM-related role—almost guaranteed to appear, and tested in fine detail.
Likely interview follow-ups:
- What Q, K, and V each mean and where they come from: "attention from a sequence to itself—why is it called self-attention?"
- The scaling factor: why divide the dot product by √d_k? "Dot-product variance grows with dimension, which saturates the softmax."
- Positional encoding: why is it needed? What problem does rotary positional encoding (RoPE) solve?
- The point of multi-head attention: heads "look at relationships in different subspaces"—contrast this clearly with single-head attention.
- Comparison with CNN/RNN: parallelism, long-range dependencies, quadratic complexity.
Bottom line: Being able to tell the 5-minute story of "the attention computation flow + why the scaling + why positional encoding is necessary" is the passing bar for this section. For further reading see Transformer and Attention; for a hands-on way in, see ChatGPT and Conversational AI.
Recommended reading: Transformer and Attention · Large Language Models (LLM)
4. LLM Principles / Pre-training / Emergence
JD wording: "Familiar with the principles of large language models," "experience training or applying LLMs." The core question: how did large models get so strong?
Likely interview follow-ups:
- The pre-training objective: why can next-token prediction learn grammar, knowledge, and reasoning?
- Scaling Laws: as parameters, data, and compute grow together, how does capability change? Where are the limits of "bigger is stronger"?
- Emergence: are emergent abilities real capabilities or artifacts of how we evaluate? How would you explain them to a non-expert?
- In-context learning: what mechanistic explanations exist for few-shot/zero-shot behavior (the Bayesian view vs. the implicit-gradient view)?
- Hallucination: the inherent flaw of parametric memory and how to mitigate it (which is also the motivation for RAG).
Bottom line: This section doesn't test memorized conclusions—it tests whether you have built the mental model of "an LLM = compressed memory + probabilistic generation." For the full explanation see Large Language Models (LLM); for the historical arc see A Brief History.
Recommended reading: Large Language Models (LLM) · Transformer and Attention · DeepSeek-R1 and Reasoning Models
5. Prompts / CoT / Structured Output
JD wording: "Familiar with prompt engineering," "proficient with CoT." This section looks easy, but interviews love to dig deep here.
Likely interview follow-ups:
- Prompt components: what problems do the system prompt, few-shot examples, and output-format constraints each solve?
- CoT: why does "let's think step by step" work? When does CoT fail?
- Structured output: JSON mode, Function Calling schema design—how do you guarantee parsing never fails?
- Prompting vs. fine-tuning: when is it enough to change the prompt, and when must you fine-tune?
- Cost and latency: longer prompts mean pricier tokens—how do you trade quality against cost?
Bottom line: There is no "standard answer" for prompts, only "measurable iteration"—every change has to be provably better with data. That is the dividing line between "someone who tweaks prompts" and "someone who does prompt engineering." For the practical manual see The Prompt Playbook.
Recommended reading: Prompt Engineering · The Prompt Playbook · LLM Evaluation and Benchmarks
6. RAG / Vector Retrieval / Reranking
JD wording: "Proficient with RAG," "familiar with vector retrieval." The core topic for application-track roles—and the densest interview material of all.
Likely interview follow-ups:
- The full pipeline: offline indexing (parse → clean → chunk → embed → store) and online retrieval (vectorize the query → recall → rerank → generate). How do you troubleshoot a failure at each stage?
- Chunking: how do you split documents—by paragraph, heading, or semantics? How big should chunks be? Why does chunk size directly affect answer quality?
- Recall vs. rerank: vector recall is the "wide net," rerank is the "fine sort"—when is rerank a must?
- Hybrid retrieval: why does keyword (BM25) + vector search usually beat pure vector search? How do you fuse the results?
- Evaluation: RAG's two evaluation dimensions—retrieval quality (hit rate/NDCG) and generation quality (citation rate/faithfulness). How do you build the eval set?
Bottom line: The hidden test in RAG interviews is always whether you can troubleshoot—empty retrieval, irrelevant recall, generated answers that never cite the source: every failure mode has a matching fix. For the systematic breakdown see Retrieval-Augmented Generation (RAG); for the hands-on route see Build a RAG Application from Scratch; for a real-world case see Perplexity and AI Search.
Recommended reading: Retrieval-Augmented Generation (RAG) · Vector Databases and Semantic Search · Build a RAG Application from Scratch · Knowledge Graphs and Knowledge Injection
7. Agent / Tool Calling / ReAct
JD wording: "Hands-on experience shipping agents," "familiar with tool calling." A high-frequency section added after 2024, with interviews shifting from "understands the concept" toward "has run the real loop."
Likely interview follow-ups:
- The ReAct loop: reason → act → observe. Why is it more stable than calling tools directly? When does it break down?
- Tool calling: how do you design a Function Calling schema? How do you backstop parameter validation and error handling?
- Long tasks: planning strategies for multi-step tasks (plan everything up front vs. decide step by step)? What do you do when the context overflows?
- Stability: the agent loops forever, drifts off course, or hallucinates tool arguments—how do you detect and recover?
- Evaluation: how do you measure whether an agent is any good (task completion rate, step count, failure rate)? How does this differ from RAG evaluation?
Bottom line: What agent interviews punish most is armchair theorizing—if you have never run an agent that got stuck and retried in circles, you can't talk about stability solutions. For the hands-on route see Build an Agent from Scratch; for the product landscape see Manus and Agent Applications.
Recommended reading: AI Agents · Build an Agent from Scratch · Common Pitfalls and Anti-Patterns
8. LoRA / Fine-Tuning / Data Engineering
JD wording: "Familiar with LoRA/fine-tuning," "data engineering experience." The three pillars of fine-tuning interviews: methods, data, and evaluation.
Likely interview follow-ups:
- LoRA principles: why does the low-rank assumption hold? How do you choose the rank? What are the trade-offs versus full-parameter fine-tuning?
- QLoRA: 4-bit quantization + LoRA—why does this combination make fine-tuning possible on a single consumer GPU?
- Fine-tuning vs. RAG: when should you fine-tune, when should you use RAG, and how do you combine the two?
- Data engineering: where does instruction data come from (human/synthetic/distilled)? How do you clean, deduplicate, mix, and quality-check it?
- Overfitting and forgetting: the model lost its general abilities after fine-tuning—now what? How should evaluation cover this?
Bottom line: 70% of fine-tuning work is data engineering and 30% is training—when the interviewer asks "what have you fine-tuned," what they really want to know is "how did you build the data, and how did you prove the fine-tuning helped." For the full route see Fine-Tuning and PEFT and Fine-Tune Your Own LLM.
Recommended reading: Fine-Tuning and PEFT · Fine-Tune Your Own LLM · Datasets and Tools Directory
9. RLHF / DPO / Alignment
JD wording: "Familiar with RLHF," "knowledge of alignment." High-frequency words for foundation-model teams; a bonus item for application roles.
Likely interview follow-ups:
- The three-stage RLHF pipeline: SFT → reward model → PPO. How do you train the reward model (ranking preference pairs)?
- Why PPO is expensive: it has to load four models (policy/reference/reward/value)—what is the KL constraint for?
- DPO: it compresses RLHF's "train a reward model + reinforcement learning" into direct preference optimization—what does it simplify, and what does it give up?
- The traps of alignment: reward hacking, over-optimization—how do you mitigate them?
- Alignment vs. safety: alignment means making the model "do what humans intend," while AI Safety means "making sure no harm is done"—how do the two relate?
Bottom line: Every test point in this section is a "why"—once you understand that "RLHF is expensive because of online sampling plus four models," you understand why DPO exists. For the deep dive see Alignment: RLHF and DPO.
Recommended reading: Alignment: RLHF and DPO · AI Safety and Governance · Large Language Models (LLM)
10. Quantization / Inference Optimization / Deployment
JD wording: "Familiar with inference optimization," "knowledge of model deployment." The main battlefield for AI Infra roles, and a bonus item for application roles.
Likely interview follow-ups:
- Quantization: why do INT8/INT4 speed things up? Where does the accuracy loss come from (distribution compression, outliers)? How do you calibrate?
- KV Cache: why is the decode phase slower than prefill? How do you compute the KV Cache's memory footprint, and how do you manage it?
- Continuous batching: why does it beat static batching on throughput? What problem does PagedAttention solve?
- Deployment: how do you work out the latency budget, QPS estimate, and memory plan for a serving stack? How do you run canary rollouts without regressing quality?
- Distillation: how does a small model learn from a big one? When is distillation a better deal than quantization?
Bottom line: The three keywords of inference-optimization interviews are latency, throughput, and cost—for every technique you must be able to answer "which metric does it improve, and what does it cost?" For the deep dive see Inference Optimization and Quantization; for the hands-on route see Deployment and Inference Optimization in Practice.
Recommended reading: Inference Optimization and Quantization · Deployment and Inference Optimization in Practice · Transformer and Attention
11. Evaluation / Benchmarks / Multimodal and Diffusion
1. Evaluation / Benchmarks
JD wording: "Familiar with LLM evaluation," "has built an evaluation system." This is a capability that spans every role, and the line that differentiates resumes the most.
Likely interview follow-ups:
- Evaluation types: human evaluation vs. model-based evaluation (LLM-as-judge) vs. automated metrics—what are the costs and biases of each?
- Eval set design: what to test (factuality/instruction following/reasoning/safety), how to pick samples, which edge cases to cover?
- Metric mapping: how do offline metrics correspond to business goals (answer accuracy vs. user retention)?
- Quantifying hallucination: how do you compute citation rate and faithfulness? How do you attribute failures separately to retrieval vs. generation?
- The limits of benchmarks: what can't MMLU and HumanEval measure?
Bottom line: "How do you prove your change actually worked" is the most valuable sentence in any interview—evaluation skill is what turns that sentence into a methodology. For the full breakdown see LLM Evaluation and Benchmarks; for hands-on work see Build an LLM Evaluation Suite.
Recommended reading: LLM Evaluation and Benchmarks · Build an LLM Evaluation Suite
2. Multimodal and Diffusion
JD wording: "Multimodal experience," "familiar with diffusion models." Multimodal understanding and generation form a track of their own, and a new add-on for traditional CV/NLP roles.
Likely interview follow-ups (understanding track):
- How do you align a vision encoder with an LLM (the trade-offs among projection layers, Q-Former, and Adapters)?
- Multimodal hallucination: the model "sees" objects that aren't there—what causes it, and how is it evaluated?
- Cross-modal retrieval: in image-text retrieval, how does contrastive learning organize positive and negative samples?
Likely interview follow-ups (generation track):
- The diffusion process: forward noising + reverse denoising—what is the training objective?
- Sampling acceleration: why do DDIM and DPM-Solver cut the number of steps? What is the quality-speed trade-off?
- Conditional control: how do ControlNet and LoRA inject conditions into generation?
Bottom line: Multimodal roles test alignment (understanding track) or generation quality and controllability (generation track). Going deep on one track beats being shallow on both. For the understanding track see Multimodal Models; for the generation track see Diffusion Models and Generative AI, Image Generation Case Studies, and Video Generation Case Studies.
Recommended reading: Multimodal Models · Diffusion Models and Generative AI · Speech AI Case Studies
12. Not Written in the JD but Tested in Interviews: Hidden Test Points
The test points below rarely appear verbatim in a JD, yet they are the hidden dividing line of interviews—they are what move you from "can recite it" to "can do the work."
| Hidden Test Point | Why It's Not in the JD but Still Tested | Where to Start Preparing |
|---|---|---|
| The full engineering and deployment pipeline | Interviewers assume plenty of candidates "understand the concepts" but few can "get something running in production" | Deployment and Inference Optimization in Practice · Common Pitfalls and Anti-Patterns |
| Evaluation design skills | The JD says "familiar with evaluation," but the real test is "can you design an evaluation for this specific scenario" | Build an LLM Evaluation Suite · LLM Evaluation and Benchmarks |
| Data engineering fundamentals | Nearly every role spends its first month processing data, yet JDs never say so directly | Datasets and Tools Directory |
| Cost awareness | Token costs and GPU costs sit at the heart of business decisions, but JDs almost never spell them out | Inference Optimization and Quantization |
| Safety and compliance | Content safety and data compliance for generative AI are becoming hard constraints | AI Safety and Governance |
| Combining skills across scenarios | Interviews often pose composite scenario questions (knowledge-base Q&A, agent workflows) that test how multiple sections work together | Build a RAG Application from Scratch · Build an Agent from Scratch |
| Industry and product sense | "Who is this solution for, and what's the ROI" is the question that separates seniority levels | Module Guide and Role Map · AI Product Manager |
Composite scenario questions are the biggest hidden test point
Interview questions in 2025 are almost all scenario questions: "Build a customer-service knowledge base for an enterprise: 100,000 documents, answers must be traceable to their sources, 500 QPS at the evening peak—how do you design it?" This single question tests RAG architecture (Section 6), evaluation (Section 11), inference optimization (Section 10), and cost (the hidden test points) all at once. After reviewing section by section, be sure to run a "combo drill" once—stitch the knowledge from multiple sections into one end-to-end system design and present it out loud.
13. The Five-Tier Self-Assessment Table and Your Study Checklist
1. The Five-Tier Self-Assessment Scale
| Tier | Meaning | Bar to Clear (score yourself in the mirror) |
|---|---|---|
| 1 Never heard of it | Completely unfamiliar | You've never even seen the term |
| 2 Heard of it | A vague impression | You've seen or heard it, but can't define it |
| 3 Understand the concept | Knows what it is | You can give a definition and one example, but not the principles or trade-offs |
| 4 Familiar | Can use it hands-on | You can implement, tune, and debug it on your own, and know its strengths, weaknesses, and where it applies |
| 5 Can explain it thoroughly | Can teach it | You can cover principle + derivation + trade-offs + counterexamples in 5 minutes and survive any follow-up |
2. Self-Assessment Table for the Ten Sections (Printable)
| Test Point | Never heard of it | Heard of it | Understand the concept | Familiar | Can explain it thoroughly |
|---|---|---|---|---|---|
| Transformer: attention / QKV / positional encoding | □ | □ | □ | □ | □ |
| LLM principles: pre-training / Scaling Laws / emergence | □ | □ | □ | □ | □ |
| Prompts: CoT / structured output / prompt vs. fine-tuning | □ | □ | □ | □ | □ |
| RAG: chunking / recall / rerank / evaluation | □ | □ | □ | □ | □ |
| Vector databases: similarity / indexing / hybrid retrieval | □ | □ | □ | □ | □ |
| Agent: ReAct / tool calling / stability | □ | □ | □ | □ | □ |
| Fine-tuning: LoRA / QLoRA / data engineering | □ | □ | □ | □ | □ |
| Alignment: RLHF / DPO / reward hacking | □ | □ | □ | □ | □ |
| Inference optimization: quantization / KV Cache / deployment | □ | □ | □ | □ | □ |
| Evaluation: eval sets / metrics / business mapping | □ | □ | □ | □ | □ |
| Multimodal / diffusion: alignment / sampling / controllability | □ | □ | □ | □ | □ |
| Hidden test points: full engineering pipeline / cost / safety | □ | □ | □ | □ | □ |
How to print and check off
Work through the table row by row, section by section, and check the tier that matches your real current level. Honesty comes first—the goal is to surface every row below "can explain it thoroughly," not to make the table look pretty. When in doubt, check the tier below.
3. Generating Your Personal Study Checklist
Once the self-assessment is done, turn it into an executable weekly plan in four steps:
| Step | Action | Key Points |
|---|---|---|
| Step 1 | Group by tier | Tiers 1–2 = Group A (first get clear on "what it is"); Tier 3 = Group B (fill in principles and trade-offs); Tier 4 = Group C (write a 5-minute talk track) |
| Step 2 | Prioritize by section | Cover the hard-requirement sections for your target role first (see the JD list), then the bonus sections |
| Step 3 | Apply the weekly plan template | 3 Group A items + 2 Group B items + 1 talk track per week; rebalance A/B/C as you progress |
| Step 4 | Follow the three execution principles | You only know it if you can say it (record and replay to find the stumbles); deriving by hand > reading derivations; re-test weekly (re-score every Sunday) |
4. The Four Execution Principles
- Principle 1: You only know it if you can say it. After reviewing each test point, record a 5-minute self-explanation on your phone and play it back to find where you get stuck—every stumble marks something you haven't truly understood.
- Principle 2: Deriving by hand beats reading derivations. For the math-heavy points (attention scaling, the LoRA update rule, quantization error), always work the derivation out yourself—in an interview, "I've seen it" and "I can derive it" are two different things.
- Principle 3: Projects beat concepts. Finish at least one project end to end (Build a RAG Application from Scratch or Build an Agent from Scratch are good picks) to ground the concepts in real engineering.
- Principle 4: Re-test weekly. Re-score every Sunday evening, cross out the test points that moved up a tier, and flag the ones that didn't move—then ask yourself why.
One more note
The goal of the study checklist is not to "fill in the whole table"—it is to lift every gate item in the target role's JD to tier 4 or above. Sections the role doesn't need (say, a RAG application role doesn't require tier-5 pre-training) can stop at tier 4—put your time where it cuts deepest.
Further Reading
On this site
- Module Guide and Role Map · JD List · Resume Analysis · Interview Question Bank — the complete four-part career module
- Learning Paths — the full through-line for learning backward from JDs and interview questions
- In-depth pages for the ten sections: Transformer and Attention · Large Language Models (LLM) · Prompt Engineering · Retrieval-Augmented Generation (RAG) · Vector Databases and Semantic Search · AI Agents · Fine-Tuning and PEFT · Alignment: RLHF and DPO · Inference Optimization and Quantization · LLM Evaluation and Benchmarks · Multimodal Models · Diffusion Models and Generative AI
- Hands-on practice: Build a RAG Application from Scratch · Build an Agent from Scratch · Build an LLM Evaluation Suite · Fine-Tune Your Own LLM · Deployment and Inference Optimization in Practice
- Quick lookup: Glossary · Datasets and Tools Directory · Models and Leaderboards Cheat Sheet
References
- Vaswani et al. Attention Is All You Need (NeurIPS 2017) — the original Transformer paper and the authoritative source on attention
- Kaplan et al. Scaling Laws for Neural Language Models (2020) — the original scaling-laws paper
- Wei et al. Chain-of-Thought Prompting (NeurIPS 2022) — the original chain-of-thought paper
- Lewis et al. Retrieval-Augmented Generation (NeurIPS 2020) — the original RAG paper
- Yao et al. ReAct: Synergizing Reasoning and Acting (ICLR 2023) — the original ReAct paper
- Hu et al. LoRA: Low-Rank Adaptation (ICLR 2022) — the original LoRA paper
- Ouyang et al. Training language models to follow instructions with human feedback (2022) — the key InstructGPT / RLHF paper
- Rafailov et al. Direct Preference Optimization (NeurIPS 2023) — the original DPO paper
- Ho et al. Denoising Diffusion Probabilistic Models (NeurIPS 2020) — the foundational diffusion-model paper
- Radford et al. Learning Transferable Visual Models From Natural Language Supervision (CLIP) (2021) — the landmark image–text alignment paper
- Hugging Face LLM Course — a free hands-on course covering LLM fine-tuning, RAG, and evaluation