Skip to content

Knowledge Breakdown: Mapping JDs to Testable Concepts

At a glance Maps every high-frequency JD requirement (Transformer/attention, pretraining, RLHF/DPO, LoRA, evaluation, RAG, Agent, inference optimization, engineering skills) to knowledge points and site pages. Each point includes "how interviewers will ask about it," plus a five-level self-assessment table and study-plan generator.

This page contains time-sensitive content, current as of 2025-08; job descriptions, rankings, product features, and other information may have changed. Please verify with original sources before citing.

Knowledge Breakdown: Mapping JDs to Testable Concepts ​

Remember this one-liner from this page: every skill word behind a JD hides a "real exam question" — what interviewers truly want to know is not whether you can recite a definition, but whether you can explain "why" something works and connect it to engineering consequences. This page breaks the high-frequency requirements from the JD List into nine knowledge blocks, each with "must-master → how interviewers will ask → related page on this site." It ends with a self-assessment table and a study-plan template.

How to use this page

This page is not a reading manual — it's a gap detector. Start with the self-assessment table at the end, mark everything you don't know or are shaky on, then go back to the relevant pages for deep reading. Reading from top to bottom will give you a false sense of "I know a little bit of everything."

1. JD High-Frequency Requirements Overview ​

Knowledge BlockCommon JD PhrasingHow Interviewers AskRelated PagePriority
Transformer & Attention"Deep understanding of Transformer internals"Attention formula derivation, why divide by √d_k, positional encoding extrapolationTransformer Architecture Explained★★★
Pretraining & Data"Familiar with pretraining and data pipelines"Data mixing, scaling laws, Chinchilla conclusionsPretraining, Scaling Laws★★★
Fine-tuning (LoRA)"Proficient in SFT/LoRA fine-tuning"LoRA principles, how to choose rank r, memory comparisonFine-tuning, Fine-tuning Practice★★★
Alignment (RLHF/DPO)"Understanding of alignment techniques"RLHF three steps, PPO objectives, DPO vs RLHF differences, alignment taxAlignment★★
Evaluation"Capable of model evaluation"What benchmarks exist? Data contamination? LLM-as-a-judge bias?Evaluation & Benchmarks, Evaluation Practice★★★
RAG"Familiar with the RAG pipeline"Three stages, retriever comparison, chunk strategies, RAG vs fine-tuningRAG, RAG Practice★★★
Agent & Tool Calling"Agent experience"ReAct loop, function calling, memory design, prompt injectionLLM-based Agents★★
Inference Optimization"Familiar with KV cache and quantization"KV cache memory formula, quantization principles, batching, memory estimationInference Fundamentals, Deployment & Serving★★★
Engineering Skills"Engineering capability"Distributed training, framework selection, system designFramework & Tool Selection★★★

Priority guide: ★★★ = widely deep-dived in application/algorithm roles; ★★ = explaining principles + one example is sufficient. Detailed breakdowns follow.

2. Block One: Transformer & Attention (★★★★★) ​

Why it's the highest-frequency topic: No matter what role the JD says, Transformer is the underlying protocol. Interviewers use this block to quickly determine whether you've "read some blog posts" or "actually understand."

Must master:

  • The full scaled-dot-product attention formula and the QKV computation flow;
  • Why we divide by √d_k (variance control → softmax gradient problems);
  • The motivation for multi-head attention (different subspaces, complementary features);
  • Positional encoding evolution: absolute (sin/cos) → relative → RoPE / ALiBi, and extrapolation challenges;
  • Residual connections and LayerNorm (Pre-Norm vs Post-Norm), FFN and activation functions;
  • Decoder-only causal masking; self-attention's O(n²) complexity;
  • The concept of KV cache (why caching is needed at inference, what exactly gets cached).

How interviewers will ask: Write or explain the attention formula; "Why scale?"; "How does RoPE work, and what happens when extrapolating?"; "Why do multiple heads help?"; "What exactly does KV cache store during inference?"

Related pages on this site: Read the "Self-Attention," "Multi-Head," and "Positional Encoding" sections of Transformer Architecture Explained; see KV cache in Inference Fundamentals; see extrapolation solutions (YaRN/NTK) in Context Window & Long Context.

3. Block Two: Pretraining & Data (★★★★) ​

Why it matters: Mandatory for training roles; application roles are also frequently asked "why base models have these capabilities/deficiencies," and the answer lies in data and objectives.

Must master:

  • Pretraining objective = next-token prediction (autoregressive), and how it differs from BERT's masked language modeling;
  • The relationship between cross-entropy and perplexity;
  • The full data pipeline: collection → cleaning → deduplication (MinHash) → filtering → mixing;
  • How data mixing (multilingual/code/math/books) affects capabilities;
  • Scaling laws: Kaplan (loss decreases as a power law over parameters/data/compute) and Chinchilla (~1:20 parameter:token ratio for compute-optimal training);
  • The division of labor between pretraining, fine-tuning, and alignment.

How interviewers will ask: "What is the Chinchilla conclusion, and why shouldn't you just blindly apply it?"; "What happens if you add too much code data?"; "Why is data deduplication important?"; "How do you troubleshoot if loss isn't decreasing?"

Related pages on this site: Language Modeling (objectives and perplexity), Pretraining (full data pipeline), Scaling Laws (mixing and Chinchilla), Datasets & Benchmarks.

4. Block Three: Fine-tuning (SFT & LoRA) (★★★★★) ​

Why it matters: A core daily task for application roles and the most common direction for resume projects. Interviewers will follow your project descriptions deep into principles.

Must master:

  • What SFT is, what the data looks like (instruction-response pairs, conversation trees), and how SFT differs from pretraining;
  • Full-parameter fine-tuning vs. parameter-efficient fine-tuning (LoRA/QLoRA/Adapter/P-Tuning) — pros and cons;
  • LoRA principles: low-rank decomposition W = W₀ + BA, the role of rank r and scaling α, and why parameter efficiency works;
  • Estimating trainable parameters (roughly how many parameters are trained for a 7B model with LoRA r=16);
  • Overfitting and catastrophic forgetting; data quality > data quantity;
  • Whether alignment is still needed after SFT.

How interviewers will ask: "Why does LoRA work (intuition)?"; "What happens if r is too large?"; "Why can QLoRA save memory?"; "What to do if the model gets dumber after fine-tuning?"; "Can LoRA weights be merged back into the base?"

Related pages on this site: Deep-read Fine-tuning; see the full engineering flow in Fine-tuning Practice: Complete LoRA Pipeline; see how to prove fine-tuning effectiveness in Evaluation Practice.

5. Block Four: Alignment (RLHF & DPO) (★★★) ​

Why it matters: Mandatory for training roles; algorithm roles frequently use it as a deep-dive topic ("Do you understand RLHF?"). By 2025, the interview focus has shifted from RLHF internals to the DPO/RLHF tradeoffs.

Must master:

  • The alignment problem definition: helpful / honest / harmless;
  • InstructGPT's three steps: SFT → Reward Model (RM) training → PPO reinforcement learning;
  • The role of "importance sampling ratio + clip + KL penalty" in the PPO objective;
  • Why DPO doesn't need an explicit RM (implicit reward, direct preference optimization);
  • DPO vs RLHF comparison (training cost, stability, data requirements);
  • Alignment tax: the phenomenon of benchmark scores dropping after alignment.

How interviewers will ask: "How is the reward model trained, and where does the data come from?"; "Why clip in PPO?"; "What does the KL penalty do?"; "What are the pitfalls of DPO and RLHF, respectively?"; "Why does alignment make models 'dumber'?"

Related pages on this site: Deep-read Alignment: RLHF & DPO; see the relationship between alignment and safety in Safety & Risks; paper versions in Classic Paper Deep Dives (InstructGPT).

6. Block Five: Evaluation (★★★★★) ​

Why it matters: The "hidden main exam" for LLM role interviews — any effectiveness claim you make must be self-provable. Evaluation capability is what separates "can use" from "can deploy."

Must master:

  • Three evaluation categories: intrinsic metrics (perplexity), task benchmarks (MMLU/GSM8K/HumanEval), and human/model judgment (MT-Bench, LLM-as-a-judge);
  • What each common benchmark measures and its pitfalls (multiple-choice contamination, pass@k, format sensitivity);
  • Data contamination: what happens when benchmarks appear in pretraining corpora;
  • LLM-as-a-judge implementation and biases (position bias, length bias, self-preference);
  • Offline evaluation (regression testing, golden sets) and online evaluation (A/B, user feedback) closed loops.

How interviewers will ask: "Given a model, how would you score it?"; "What does an MMLU score of 90 actually mean, and what pitfalls might exist?"; "How do you prove your fine-tuning was effective?"; "What biases exist when using LLMs as judges?"

Related pages on this site: Deep-read Evaluation & Benchmarks; see engineering implementation in Evaluation Practice; benchmark profiles in Datasets & Benchmarks.

7. Block Six: RAG (★★★★★) ​

Why it matters: The most commonly appearing technology in application role projects. Interviews move from "the overall flow" to "specific details," then to "why not use an alternative."

Must master:

  • RAG motivation: knowledge cutoff, hallucination, private data, traceability;
  • Three-stage flow: indexing (parsing/chunking/embedding) → retrieval (vector + sparse + reranking) → generation;
  • The impact of chunk size and overlap; hybrid retrieval (BM25 + vector) and reranking;
  • RAG vs fine-tuning vs long context: comparison and composition;
  • RAG failure modes (insufficient recall, retrieval noise, faithfulness) and evaluation metrics;
  • Causes of hallucination and mitigation strategies (hallucination).

How interviewers will ask: "Walk through the full RAG flow"; "How do you chunk, and why?"; "How do you evaluate retrieval quality?"; "When should you use RAG vs fine-tuning?"; "What if everything retrieved is garbage?"

Related pages on this site: Deep-read RAG: Retrieval-Augmented Generation; see end-to-end flow and failure mode table in RAG Practice.

8. Block Seven: Agent & Tool Calling (★★★★) ​

Why it matters: The fastest-growing skill keyword in JDs. Interviews focus on "can you build a reliable, production-ready Agent, not just a demo."

Must master:

  • Agent architecture: LLM as controller + tools + memory + planning loop;
  • Why the ReAct pattern (Reason + Act interleaving) is stronger than pure Chain-of-Thought;
  • Tool calling (function calling) mechanisms and JSON schema constraints;
  • Memory design: short-term (context window) and long-term (vector stores / structured storage);
  • Multi-agent collaboration and role division;
  • Risks: prompt injection (direct/indirect), tool misuse, cost runaway.

How interviewers will ask: "Design an Agent that can check the weather and book a calendar event"; "What happens when an Agent freezes or loops?"; "How do you select tools when multiple are available?"; "How do you prevent prompt injection?"; "How do you choose between an Agent and a hardcoded workflow?"

Related pages on this site: Deep-read LLM-based Agents; see structured output in Prompt Engineering; see injection defense in Safety & Risks.

9. Block Eight: Inference Optimization (★★★★★) ​

Why it matters: Mandatory for training and inference roles; application roles increasingly ask about cost and latency ("can your solution actually run in production?").

Must master:

  • Autoregressive inference flow: prefill and decode phases;
  • KV cache: what gets stored, the memory formula (≈ 2 × batch × layers × heads × head_dim × seq_len × 2 bytes), linear growth with sequence length;
  • Quantization: INT8/INT4 principles, differences between GPTQ/AWQ/GGUF, and accuracy/speed tradeoffs;
  • What continuous batching and PagedAttention solve;
  • Memory estimation: weights (2 bytes/parameter under FP16) + KV cache + activations + optimizer states (~12 bytes/parameter for Adam during training);
  • Metrics: TTFT / TPOT / throughput (tokens/s), first-token latency, tail latency.

How interviewers will ask: "How much GPU memory does a 7B model's FP16 weights occupy during inference?"; "How do you calculate KV cache size?"; "How much quality drops after INT4 quantization?"; "Why is continuous batching necessary?"; "How do you troubleshoot slow serving?"

Related pages on this site: Deep-read Inference Fundamentals: Autoregressive Decoding & Sampling; see memory formulas, quantization, and metrics in Deployment & Serving; MoE deployment challenges in MoE Sparse Expert Models.

10. Block Nine: Engineering Skills (★★★★) ​

Why it matters: The "engineering requirements" block in JDs, distinguishing "lab-only candidates" from "production-ready candidates." Tested in both application and training roles.

Must master:

  • Python engineering fundamentals (typing, logging, error handling, testing);
  • Distributed training terms: data parallel / tensor parallel / pipeline parallel / ZeRO — what problems each solves and their communication overhead;
  • Framework tiering: model libraries (Transformers), training (DeepSpeed/Megatron-LM/TRL), inference (vLLM/SGLang), orchestration (LangChain/LlamaIndex), vector stores (FAISS/Milvus/pgvector);
  • The engineering toolkit: automated evaluation, observability (logs/metrics/alerts), version management (model and data versioning);
  • "Don't be enslaved by frameworks": understand the principles behind frameworks, not just how to call them.

How interviewers will ask: "What's the difference between data parallelism and model parallelism, and when do you use each?"; "What does each of ZeRO's three stages optimize?"; "How do you choose a vector database?"; "Production latency spikes — how do you debug?"

Related pages on this site: Deep-read Framework & Tool Selection; see serving in Deployment & Serving; see anti-patterns in Common Pitfalls & Anti-Patterns.

11. Role Priority Matrix for Knowledge Blocks ​

The same nine blocks carry completely different weights across roles. The table below assigns "exam weight" as ★/★★/★★★ (drawn from common-pattern observations in the JD List keyword radar) — use this to decide your study order:

RoleTransformerPretrainingFine-tuningAlignmentEvaluationRAGAgentInference OptimizationEngineering
Algorithm Engineer (LLM)★★★★★★★★★★★★★★★★★★★★★★★
NLP Algorithm Engineer★★★★★★★★★★★★★★★★★
LLM Training Engineer★★★★★★★★★★★★★★★★★★★★★
Inference Optimization Engineer★★★★★★★★★★★★★★★★
AI Application Engineer★★★★★★★★★★★★★★★★★★★★
Agent Engineer★★★★★★★★★★★★★★★★★★
Evaluation Engineer★★★★★★★★★★★★★★★★★
Data Engineer (LLM)★★★★★★★★★★★★★★

How to use this matrix:

  1. Read across your target role's row: ★★★ blocks are mandatory for interviews — prepare them to the level of "can explain principles + one example." ★★ blocks: at least know definitions and flows. ★ blocks: can be mentioned in passing with projects, no need to study specifically;
  2. Read down columns: you'll notice that "evaluation" and "engineering skills" are ★★★ across nearly all roles — they're the shared foundation for LLM roles in 2025, and should always have the highest priority;
  3. Reorder your study plan: multiply the gaps from Section 13's self-assessment × the weights, to get a clear order of "what to study first" — blocks rated ★★★ with self-assessment ≤2 go to the first tier.

Why "evaluation" is ★★★ across almost all roles

The black-box nature of LLM capabilities makes "proving effectiveness" a universal need for every role: algorithm roles prove model quality, application roles prove RAG ROI, training roles prove the value of new mixes or hyperparameters, and inference roles prove quantization didn't lose accuracy. Evaluation is the universal currency across roles, and also the area where you can most easily differentiate yourself with one or two projects.

12. High-Frequency Topic Cheat Sheet (Last Review Before the Interview) ​

Sweep through this card the night before (or while waiting for your interview). Each row is a one-liner version of a testable topic — if you can't answer it, go back to the linked page to study.

TopicOne-LinerSource
Attention formulasoftmax(QKᵀ/√d_k)V; dividing by √d_k stabilizes softmax gradientsTransformer
KV cacheCache historical token K/V during decoding to avoid recomputing prefixes; memory grows linearly with sequence lengthInference Fundamentals
Pretraining objectiveNext-token prediction (autoregressive); perplexity = exp(cross-entropy)Language Modeling
Scaling lawsCompute-optimal ratio: tokens ≈ 20 × parameters (Chinchilla)Scaling Laws
LoRAW = W₀ + BA, low-rank increments can be merged back into base weights; r controls expressivity, α controls scalingFine-tuning
RLHFSFT → Reward Model (Bradley-Terry) → PPO (clip + KL); alignment tax = benchmark score drop post-alignmentAlignment
DPOUses the log-ratio of policy and reference policy as implicit reward, eliminating RM training and RL loopsAlignment
RAGIndex (parse/chunk/embed) → Retrieve (hybrid + rerank) → Generate (with citations); complementary to fine-tuning, not a replacementRAG
AgentLLM as controller + tools + memory + planning loop; ReAct = reasoning and action interleaved; defend against injection via system boundaries, not promptsLLM-based Agents
Inference optimization7B FP16 weights ≈ 14GB; KV memory = 2×layers×heads×head_dim×seq×batch×bytes; always run regression after INT4 quantizationDeployment & Serving

How to use the cheat sheet correctly

For every row in the card, you should be able to expand it into a 3-minute explanation on the spot. If you can't explain "why" for any row, you haven't finished studying that topic — go back to the self-assessment table in Section 13, re-mark it as "conceptually fuzzy," and add it to your study plan.

13. Self-Assessment Table: Five-Level Scale ​

For each block, rate yourself on a five-level scale from "completely don't know" to "have hands-on experience and can be deep-dived" (0–4):

Block0 Completely don't know1 Conceptually fuzzy2 Can define3 Can explain principles + examples4 Hands-on, can be deep-divedTarget
Transformer & Attention☐☐☐☐☐4
Pretraining & Data☐☐☐☐☐3 (application) / 4 (training)
Fine-tuning (LoRA)☐☐☐☐☐4
Alignment (RLHF/DPO)☐☐☐☐☐3
Evaluation☐☐☐☐☐4
RAG☐☐☐☐☐4 (application)
Agent & Tool Calling☐☐☐☐☐3
Inference Optimization☐☐☐☐☐3
Engineering Skills☐☐☐☐☐4

The standard for level 3

"3 = Can explain principles + examples" means you can talk for 5 minutes without notes, and you can cite one real example you've done or read. If you can't do that, you'll come across in interviews as "seem to understand but can't explain thoroughly."

14. How to Generate Your Personal Study Plan ​

Step 1: Mark gaps. Note every block where your self-assessment is ≤2 — these are your study targets.

Step 2: Weight by role. Cross-reference with the JD List keyword radar: for application roles, study RAG/evaluation/fine-tuning first; for training roles, pretraining/distributed first; for inference roles, quantization/KV cache first.

Step 3: Apply the study-plan template:

Target role: ____________________
Study period: ____ weeks

| Block | Current | Target | This Week's Actions (read which page / do which exercise) | Verification Method |
|---|---|---|---|---|
| RAG | 1 | 3 | Read RAG practice + build a local demo | Can explain the flow and answer 5 follow-ups |
| Evaluation | 2 | 3 | Read evaluation practice + run a benchmark with lm-eval | Can explain scores and contamination |
| ... | | | | |

Step 4: Three execution rules.

  1. Only study to your target level — don't over-study: interviews only test high-frequency topics; the rest come through in projects;
  2. Use Interview Questions as your verification standard: you've only "finished" a block when you can answer the relevant questions;
  3. Review weekly: explain the "I know" blocks to someone (or something) outside the field. If you stumble, you don't truly know it yet.

15. Further Reading ​

Continue within the site

  • Interview Questions — every block on this page has full answers here
  • JD List — return to the keyword radar to confirm your study order
  • Resume Analysis — how turned completed knowledge into resume evidence
  • Learning Paths — when you have time, follow the systematic improvement path

References ​