Theme
Deconstructing JD Knowledge Points
You receive a JD from your dream company, with phrases like "proficient in machine learning," "familiar with deep learning," "strong mathematical foundation," "experience with large-scale distributed training"… you can read every line, but each one feels foggy: how proficient do I actually need to be? How will the interviewer test me?
What this article does is translate the "skill words" in a JD into a knowledge checklist you can review item by item—not the vague "you need to learn math," but precise check points like "can derive the chain rule for matrix differentiation" or "can draw the self-attention computation graph for Transformer." Paired with the five-level self-assessment table and study-gap-list generation method at the end, you can complete a "skills health check" of your own abilities in one afternoon, then get a learning roadmap that's uniquely yours.
Receive JD ──① Extract skill words──② Map to check points──③ Five-level self-assess──④ Generate study gap list──⑤ Weekly re-test
│ │ │ │
Proficient/Familiar/Understand Section 2-6 here Section 7 here Section 8 hereThis article suits: candidates applying for ML algorithm roles, LLM roles, ML engineering roles, and anyone who wants to turn "I think I know it all" into "I can explain it thoroughly."
1. What JD Skill Words Actually Test
A JD is not a course syllabus; it's the company's minimum expectation for "what you can independently produce." The subtext from hiring managers is:
"Proficient in machine learning" doesn't mean memorizing every algorithm. It means when faced with a new problem, you can independently judge which model to use, why, how to evaluate, and how to debug when things go wrong.
So the first step in deconstructing a JD is mapping skill words to "what the interviewer actually tests."
1. Real Test Points for Common JD Phrases
| Original JD Phrase | Literal Meaning | What Interviewers Actually Test (prove it within 5 minutes) |
|---|---|---|
| Proficient in machine learning | Familiar with common algorithms | Clarify boundaries between supervised / unsupervised / reinforcement learning; given a business scenario, complete the full loop of "problem definition → baseline → modeling → evaluation" |
| Familiar with deep learning | Can use PyTorch / TensorFlow | Derive backpropagation by hand; explain what CNN / RNN / Transformer each solve; clarify the cause of vanishing / exploding gradients |
| Strong mathematical foundation | Took advanced math | Derive the gradient of cross-entropy; explain "why L1 produces sparse solutions but L2 doesn't"; explain where Bayes' theorem fits in classification |
| Solid engineering skills | Can write Python | Write bug-free pandas data processing on the spot; explain offline/online feature consistency; know how to debug "runs in training but not in production" |
| Familiar with distributed training | Used GPU clusters | Clarify data parallelism vs. model parallelism, gradient synchronization methods, communication bottlenecks; explain why mixed precision accelerates training |
| Has LLM experience | Fine-tuned an LLM | Explain the pretrain → SFT → RLHF three-stage pipeline; compare LoRA with full-parameter fine-tuning; design a RAG system and troubleshoot retrieval failures |
| Understand MLOps | Has deployed models | Explain deployment option selection, what metrics to monitor post-deployment, what to do with data drift, and retraining frequency |
2. Why You Can't Take JDs at Face Value
Three common pitfalls:
- Pitfall 1: Treating the JD as a full checklist. Of the 10 requirements on a JD, the company usually has only 3 hard thresholds (what determines whether your resume passes the screen); the rest form an "ideal candidate profile." Before applying, distinguish threshold items from nice-to-haves, and prioritize filling threshold gaps first.
- Pitfall 2: Treating "familiar" as "heard of it." When an interviewer writes "familiar with X," they expect you to explain the principle, implement it, and discuss trade-offs—not just "have heard of X." In self-assessment, use "can explain thoroughly" as the passing bar for "familiar."
- Pitfall 3: Only reviewing individual check points, not combinations. Interviews almost never test in isolation: they often give you a scenario ("how do you handle cold start for a short-video app's recommender?") and expect you to simultaneously draw on math (why this model) + ML (how to evaluate) + engineering (how to deploy) + LLM (whether to use RAG). So these five blocks must be viewed together.
3. Full Pipeline from JD to Study Gap List
Receive JD
│
▼
① Extract skill words (proficient / familiar / understand / nice-to-have)
│
▼
② Map to knowledge point list (Sections 2–6 below)
│
▼
③ Five-level self-assessment per item (Section 7)
│
▼
④ Generate personal study gap list (method in Section 8)
│
▼
⑤ Execute weekly → re-test → update listThe next five sections are the five blocks of this map.
2. Knowledge Block 1: Mathematical Foundations
JD phrase: "Strong mathematical foundation." Why do interviewers test math? Because the ceiling of model capability is determined by data, but the floor is determined by math—if you can't read formulas, you can only "call libraries," and when things go wrong, you can't attribute the cause. The math required in interviews doesn't need graduate-level proofs, but must reach the level of "can derive formulas, has geometric intuition, can compute gradients."
1. Linear Algebra (the most underrated and most frequently tested)
| Check Point | How Interviewers Test | Why You Need It |
|---|---|---|
| Vectors, matrices, and basic operations | "What's the complexity of matrix multiplication?" | The foundation of all deep learning computation; Q·Kᵀ in attention is matrix multiplication |
| Transpose, rank, matrix inverse | "When is a matrix not invertible?" | Prerequisite for the least-squares solution w = (XᵀX)⁻¹Xᵀy |
| Norms and vector spaces | "What's the difference between L1 and L2 norms?" | Direct mathematical language for regularization and loss function design |
| Eigenvalues and eigenvectors | "What are the eigenvectors of the covariance matrix?" | The mathematical core of PCA |
| Singular value decomposition (SVD) | "What is SVD used for in recommendation systems?" | The common ancestor of matrix factorization, embeddings, and dimensionality reduction |
| Matrix differentiation | "Write the gradient of the loss w.r.t. the weight matrix" | Backpropagation is fundamentally matrix differentiation + the chain rule |
2. Probability and Statistics
| Check Point | How Interviewers Test | Why You Need It |
|---|---|---|
| Conditional probability and Bayes' theorem | "Why is Naive Bayes 'naive'?" | The foundation of generative classifiers and Bayesian optimization |
| Common distributions | "When to use Bernoulli / Gaussian / multinomial?" | Basis for choosing models and loss functions for your data |
| Expectation, variance, covariance | "What does data standardization actually do?" | Understanding normalization, PCA, and feature scaling |
| Maximum likelihood estimation (MLE) | "Is ordinary least squares MLE for linear regression?" | The bridge connecting statistics and loss functions |
| Cross-entropy and KL divergence | "Where does the cross-entropy loss come from?" | The origin of the de facto standard loss function for classification |
| Law of large numbers and CLT | "Why does larger sample size give more stability?" | Theoretical basis for understanding model stability and bagging |
| Hypothesis testing and confidence intervals | "How does A/B testing determine significance?" | The statistical foundation for production evaluation and business decisions |
3. Calculus
| Check Point | How Interviewers Test | Why You Need It |
|---|---|---|
| Derivatives, partial derivatives | "Why does the gradient point in the steepest ascent direction?" | Theoretical basis of gradient descent |
| Chain rule | "Which mathematical operation does backpropagation perform?" | The engine of training deep networks |
| Vanishing and exploding gradients | "Why do gradients vanish in deep networks?" | The root cause of the #1 training problem in deep learning |
| Taylor expansion | "What's the difference between Newton's method and gradient descent?" | Understanding second-order optimization, learning rates, and loss surface curvature |
4. Optimization
| Check Point | How Interviewers Test | Why You Need It |
|---|---|---|
| Convex vs. non-convex optimization | "Is the deep learning loss function convex?" | Understanding why "local optima" aren't the primary concern in deep learning |
| Gradient descent / SGD / Mini-batch | "What does batch size affect?" | The core training loop |
| Momentum, Adam | "What's better about Adam over SGD?" | Basis for optimizer selection |
| Learning rate | "What happens if the learning rate is too large / too small?" | The first thing to check when training doesn't converge |
5. Recommended Review Resources
Math review roadmap
Aim for "can compute, has intuition" first, not "can prove."
- 3Blue1Brown's Essence of Linear Algebra and Essence of Calculus — build geometric intuition with animations; watch these first for the big-picture view (official Chinese subtitles on Bilibili / YouTube);
- MIT 18.06 (Gilbert Strang's Linear Algebra) — widely regarded as the clearest linear algebra public course, free on OCW; pair with Strang's textbook Introduction to Linear Algebra;
- Andrew Ng's CS229 lecture notes (Stanford public course site) — the first several lectures cover convexity, optimization, and Bayesian topics that frequently appear in interviews; moderate mathematical density;
- StatQuest (Josh Starmer) — one ~10-minute video per statistical concept; great for plugging gaps on things you've "heard of but can't explain";
- Theoretical deepening: Goodfellow Deep Learning Chapters 2–4 (free online on the official site) and Bishop Pattern Recognition and Machine Learning.
The Math Primer on this site compresses the formulas above into a one-page printable cheat sheet—recommended to keep at your desk during review periods.
3. Knowledge Block 2: Machine Learning Foundations
JD phrase: "Proficient in machine learning." This is the mandatory main thread of interviews and the only block where "pure memorization gets you the base score"—but getting a high score requires being able to explain why. See the full framework on this site's Supervised Learning page.
1. Supervised Learning Algorithms
| Check Point | How Interviewers Test | Page on this Site |
|---|---|---|
| Linear regression & OLS | "Derive the normal equation by hand" | Supervised Learning |
| Logistic regression | "Why is it called regression but does classification? Where does sigmoid come from?" | Same |
| Decision trees | "How are information gain and Gini used to pick split points?" | Same |
| Random forests | "Why does bagging reduce variance?" | Same |
| GBDT / XGBoost | "Why does boosting reduce bias? What did XGBoost add?" | Same |
| SVM & kernel trick | "Where does the kernel map data to? Why does it work?" | Same |
| KNN, Naive Bayes | "What are their underlying assumptions?" | Same |
2. Unsupervised Learning
| Check Point | How Interviewers Test | Notes |
|---|---|---|
| K-Means | "How to pick K? What are K-Means' implicit assumptions?" | Must-test for clustering basics |
| Hierarchical clustering, DBSCAN | "How to handle arbitrarily shaped clusters?" | Answer by contrasting with K-Means |
| PCA & dimensionality reduction | "Difference between PCA and feature selection?" | Core dimensionality reduction topic |
| t-SNE / UMAP | "Why can't you use visualization reduction for production features?" | A frequently tested "trap question" |
| Association rules | "How are support, confidence, and lift defined?" | Low frequency but classic |
3. Evaluation (a major minefield in interviews)
Model evaluation is the block that most easily exposes "can only call libraries," because evaluation design directly reflects your understanding of the business. See Model Evaluation and Validation on this site.
| Check Point | How Interviewers Test | Page on this Site |
|---|---|---|
| Confusion matrix, accuracy / precision / recall / F1 | "What do you look at when positive/negative samples are severely imbalanced?" | Model Evaluation and Validation |
| ROC curve and AUC | "What is the physical meaning of AUC?" | Same |
| PR curve | "When does AUC mislead you?" | Same |
| Regression metrics MSE / MAE / R² | "Which metric does an outlier skew?" | Same |
| Cross-validation | "Why does K-fold reduce evaluation variance?" | Same |
| Overfitting / underfitting & bias-variance tradeoff | "Low training error, high validation error—what do you check first?" | Same, Overfitting & Regularization |
4. Regularization
| Check Point | How Interviewers Test | Page on this Site |
|---|---|---|
| L1 / L2 regularization | "Why does L1 produce sparse solutions and L2 doesn't? (geometric + mathematical explanation)" | Overfitting & Regularization |
| Dropout | "Randomly dropping neurons during training, not during inference—why?" | Same |
| Early stopping | "What metric determines early stopping?" | Same |
| Data augmentation | "Which tasks benefit from it?" | Same |
5. Feature Engineering
| Check Point | How Interviewers Test | Page on this Site |
|---|---|---|
| Missing value handling | "Delete, impute, or treat as a signal?" | Feature Engineering |
| Encoding: one-hot / label / target encoding | "What happens with one-hot when there are many categories?" | Same |
| Normalization & standardization | "Do tree models need normalization? Why?" | Same |
| Binning, interaction features | "Why do linear models need manual interaction terms?" | Same |
| Feature selection | "What are filter / wrapper / embedded methods?" | Same |
| Class imbalance | "Oversampling / undersampling / cost-sensitive / switching evaluation metrics—how to combine?" | Same |
A frequently tested reminder
"Consistency between training and production features" is a must-test at top companies in recent years: using future information during training, not having real-time features online, offline/online methodology mismatches… These questions test not the model, but whether you've truly done end-to-end projects. They span both Feature Engineering and MLOps, so review them together in self-assessment.
4. Knowledge Block 3: Deep Learning
JD phrase: "Familiar with deep learning." The interview weight of this block increases year by year, and the focus is shifting from "can call a framework" to "understands principles." See the overall framework on this site's Deep Learning Foundations.
1. Neural Network Essentials
| Check Point | How Interviewers Test | Page on this Site |
|---|---|---|
| Neurons & activation functions | "Why is ReLU better than sigmoid? What about dead ReLU?" | Deep Learning Foundations |
| Forward & backward propagation | "Derive backpropagation for a single layer by hand" | Same |
| Loss functions & output layer design | "What loss pairs with regression / binary classification / multi-class?" | Same |
| Optimizers | "What are the update rules for SGD / Momentum / Adam?" | Same |
| Weight initialization | "What happens if weights are initialized to 0? Why Xavier / He?" | Same |
2. Classic Architectures
| Check Point | How Interviewers Test | Page on this Site |
|---|---|---|
| CNN: convolution, pooling, receptive field | "How do you count CNN kernel parameters? What's a 1×1 conv for?" | Deep Learning Foundations |
| Classic CNN architectures | "What problem did ResNet solve? Why do residual connections work?" | Same |
| RNN / LSTM / GRU | "Why can't RNNs be trained deep? What do gating mechanisms solve?" | Same |
| Attention mechanism | "What are Q / K / V? Why scale by √dₖ?" | Transformer & NLP |
| Transformer overall structure | "Draw the Encoder's modular composition" | Same |
| Positional encoding | "Why does Transformer need positional encoding?" | Same |
3. Training Techniques
| Check Point | How Interviewers Test | Page on this Site |
|---|---|---|
| Learning rate scheduling | "Why is warmup + decay common?" | Tuning & Hyperparameter Optimization |
| Batch normalization | "Why does BN accelerate convergence? How does BN work at inference?" | Deep Learning Foundations |
| Gradient clipping | "When to use it?" | Same |
| Mixed precision training | "Why is FP16 faster? What's loss scaling for?" | Same |
| Overfitting troubleshooting workflow | "Check training error or validation error first?" | Tuning & Hyperparameter Optimization |
5. Knowledge Block 4: Engineering & Deployment
JD phrases: "Solid engineering skills," "Understand MLOps." A model that runs doesn't mean it's deployed—engineering and MLOps typically account for 30%+ of ML role interviews at top companies. See the overview on this site's MLOps & Model Deployment.
1. Programming & Data Structures
| Check Point | How Interviewers Test | Notes |
|---|---|---|
| Python basics | "Complexity of list / dict / set? What is the GIL?" | Coding environment for algorithm roles |
| Common libraries | "How to use numpy broadcasting and pandas groupby aggregation?" | Writing data processing on the spot is the most common coding question |
| Complexity analysis | "State the time complexity of a given algorithm" | Prerequisite for all coding questions |
| Common data structures | "Hash tables, heaps, union-find, graph traversal" | Question bank source for coding interviews |
2. Databases & Big Data
| Check Point | How Interviewers Test | Notes |
|---|---|---|
| SQL | "Window functions, differences between join types, how to troubleshoot slow queries?" | Algorithm roles often test basic SQL |
| Data warehouses & feature stores | "How to align offline and online features?" | Spans feature engineering and engineering |
| Distributed computing | "What is the shuffle in MapReduce / Spark?" | Bonus for big data roles |
3. MLOps
| Check Point | How Interviewers Test | Page on this Site |
|---|---|---|
| Experiment management & version control | "How to ensure every experiment is reproducible?" | MLOps & Model Deployment |
| Model registry & rollback | "Model performance degraded after deployment—how to quickly limit damage?" | Same |
| Deployment option selection | "Online / batch / streaming—which scenario fits each?" | Same |
| Inference serving | "How to estimate QPS, latency, GPU resources?" | Same |
| Monitoring & drift detection | "Which metrics to monitor? What to do with data drift?" | Same |
| Retraining loop | "How often to retrain? Who triggers it?" | Same |
| Model interpretability | "Principles and limitations of SHAP / LIME" | Same |
Review strategy for the engineering block
Engineering questions heavily favor practical experience. Rather than memorizing concepts, a more effective approach is: take one project through the full deployment pipeline (training → evaluation → wrap as API → monitoring). Once you've personally hit these pain points, every engineering answer will naturally carry the persuasiveness of "I've done this."
6. Knowledge Block 5: Large Language Models
JD phrase: "Has LLM experience"—nearly a standard bonus item for ML roles in the past two years. Note: the biggest trap of LLM questions is "you think you know it, but can't explain the mechanism." See the technical landscape on this site's Large Language Models (LLM) and Transformer & NLP.
1. Model Principles
| Check Point | How Interviewers Test | Page on this Site |
|---|---|---|
| Pre-training paradigm | "Why can next-token prediction learn language?" | Large Language Models (LLM) |
| Scaling laws | "What do Scaling Laws say? How to interpret emergence?" | Same |
| In-context learning | "What are the mechanistic explanations for few-shot / zero-shot?" | Same |
| Chain-of-thought (CoT) | "Why do we need intermediate reasoning steps?" | Same |
2. Fine-Tuning & Alignment
| Check Point | How Interviewers Test | Page on this Site |
|---|---|---|
| SFT instruction fine-tuning | "Where does SFT data come from? Why isn't a base model an usable model?" | Large Language Models (LLM) |
| LoRA / QLoRA | "Why does LoRA work with so few parameters? What is the low-rank assumption?" | Same |
| RLHF | "How is the reward model trained? Why is PPO expensive?" | Same |
| DPO | "What does DPO simplify compared to RLHF?" | Same |
3. RAG & Applications
| Check Point | How Interviewers Test | Page on this Site |
|---|---|---|
| Full RAG pipeline | "Retrieval → augmentation → generation—how to troubleshoot failure at each stage?" | Large Language Models (LLM) |
| Vector databases & embeddings | "How to chunk? How to compute similarity?" | Same |
| Prompt engineering | "How to design system prompts and few-shot templates?" | Same |
| Decoding strategies | "What do temperature, top-p, and beam search each affect?" | Same |
| Hallucination & evaluation | "How to evaluate generation quality? How to mitigate hallucination?" | Same |
4. Inference Optimization
| Check Point | How Interviewers Test | Page on this Site |
|---|---|---|
| Quantization | "Why can INT8 / INT4 accelerate? Where does precision loss come from?" | Large Language Models (LLM) |
| KV Cache | "Why is the generation phase faster than prefill?" | Same |
| Serving frameworks | "What problem does vLLM's PagedAttention solve?" | Same |
| Distillation | "How do small models learn from big models?" | Same |
7. Self-Assessment Table: Five Levels, Printable
Consolidate the check points from all five blocks into one printable self-assessment table. Scoring criteria:
| Level | Meaning | Judgment Criteria (rate yourself honestly in the mirror) |
|---|---|---|
| 1 Never heard of it | Completely unfamiliar | Haven't even seen this word |
| 2 Heard of it | Vaguely familiar | Seen/heard it, but can't give a definition |
| 3 Understand | Know the concept | Can give a definition and one example, but not principles or trade-offs |
| 4 Familiar | Can implement and use | Can independently implement / tune / troubleshoot; knows pros, cons, and applicable scenarios |
| 5 Can explain thoroughly | Can teach others | Can explain principles + derivation + trade-offs + counter-examples within 5 minutes; withstands deep follow-up |
How to print and check off
Print this section, go block by block, and check off the level that matches your current actual proficiency. Honesty is the premise—the goal is to expose every row below "can explain thoroughly," not to make the table look good. If unsure, check the lower level.
Block A: Mathematical Foundations
| Check Point | Never heard | Heard of | Understand | Familiar | Can explain thoroughly |
|---|---|---|---|---|---|
| Vector / matrix operations & matrix differentiation | □ | □ | □ | □ | □ |
| Eigenvalues / eigenvectors | □ | □ | □ | □ | □ |
| SVD & PCA mathematics | □ | □ | □ | □ | □ |
| Bayes' theorem & conditional probability | □ | □ | □ | □ | □ |
| Common distributions, expectation & variance | □ | □ | □ | □ | □ |
| MLE and its relationship to loss functions | □ | □ | □ | □ | □ |
| Cross-entropy & KL divergence | □ | □ | □ | □ | □ |
| Chain rule & backpropagation | □ | □ | □ | □ | □ |
| Gradient descent & SGD | □ | □ | □ | □ | □ |
| Convex optimization & local optima | □ | □ | □ | □ | □ |
Block B: Machine Learning Foundations
| Check Point | Never heard | Heard of | Understand | Familiar | Can explain thoroughly |
|---|---|---|---|---|---|
| Linear regression / logistic regression | □ | □ | □ | □ | □ |
| Decision trees / random forests / GBDT | □ | □ | □ | □ | □ |
| SVM & kernel trick | □ | □ | □ | □ | □ |
| K-Means / PCA & dimensionality reduction | □ | □ | □ | □ | □ |
| Confusion matrix / precision / recall / F1 | □ | □ | □ | □ | □ |
| ROC / AUC / PR curve | □ | □ | □ | □ | □ |
| Cross-validation & data splitting | □ | □ | □ | □ | □ |
| Overfitting & bias-variance tradeoff | □ | □ | □ | □ | □ |
| L1 / L2 / Dropout / early stopping | □ | □ | □ | □ | □ |
| Feature encoding / normalization / missing values | □ | □ | □ | □ | □ |
| Feature selection | □ | □ | □ | □ | □ |
| Class imbalance handling | □ | □ | □ | □ | □ |
Block C: Deep Learning
| Check Point | Never heard | Heard of | Understand | Familiar | Can explain thoroughly |
|---|---|---|---|---|---|
| Activation functions & weight initialization | □ | □ | □ | □ | □ |
| Hand-deriving backpropagation | □ | □ | □ | □ | □ |
| Loss functions & output layer design | □ | □ | □ | □ | □ |
| Optimizers (SGD / Adam) | □ | □ | □ | □ | □ |
| CNN: convolution / pooling / receptive field | □ | □ | □ | □ | □ |
| Classic architectures like ResNet | □ | □ | □ | □ | □ |
| RNN / LSTM / GRU | □ | □ | □ | □ | □ |
| Attention mechanism & QKV | □ | □ | □ | □ | □ |
| Transformer structure | □ | □ | □ | □ | □ |
| Training tricks (scheduling / BN / gradient clipping) | □ | □ | □ | □ | □ |
Block D: Engineering & Deployment
| Check Point | Never heard | Heard of | Understand | Familiar | Can explain thoroughly |
|---|---|---|---|---|---|
| Python / numpy / pandas hands-on | □ | □ | □ | □ | □ |
| Complexity & common data structures | □ | □ | □ | □ | □ |
| SQL basics & window functions | □ | □ | □ | □ | □ |
| Distributed computing (Spark, etc.) | □ | □ | □ | □ | □ |
| Experiment management & version control | □ | □ | □ | □ | □ |
| Deployment option selection | □ | □ | □ | □ | □ |
| Monitoring & drift detection | □ | □ | □ | □ | □ |
| Retraining loop | □ | □ | □ | □ | □ |
| Model interpretability (SHAP, etc.) | □ | □ | □ | □ | □ |
Block E: Large Language Models
| Check Point | Never heard | Heard of | Understand | Familiar | Can explain thoroughly |
|---|---|---|---|---|---|
| Pre-training & scaling laws | □ | □ | □ | □ | □ |
| In-context learning & chain-of-thought | □ | □ | □ | □ | □ |
| SFT instruction fine-tuning | □ | □ | □ | □ | □ |
| LoRA / QLoRA | □ | □ | □ | □ | □ |
| RLHF & DPO | □ | □ | □ | □ | □ |
| Full RAG pipeline | □ | □ | □ | □ | □ |
| Vector databases & embeddings | □ | □ | □ | □ | □ |
| Decoding strategies | □ | □ | □ | □ | □ |
| Quantization & inference acceleration | □ | □ | □ | □ | □ |
| Hallucination & model evaluation | □ | □ | □ | □ | □ |
8. Generating a Personal Study Gap List
After filling out the self-assessment table, use these four steps to turn it into an actionable weekly plan.
Step 1: Group by Level
Categorize each check point into three groups:
- Group A (Levels 1–2, completely unfamiliar): Highest priority. First solve "what is it." Goal: reach level 3 within a week.
- Group B (Level 3, understand): Missing principles and trade-offs. Goal: reach level 4 within two weeks, with the standard being "can independently implement + state applicable scenarios."
- Group C (Level 4, familiar): Missing the ability to "explain thoroughly." Goal: write a 5-minute verbal presentation for each check point. This maps directly to interview performance.
No need to spend time on Level 5 items, unless you're applying for teaching- or theory-oriented roles.
Step 2: Order by Block, Tally Up
Summarize Groups A, B, and C by block to get three "debt sheets." Suggested study order:
Priority 1 Block A: Math —— It's the foundation for all other blocks
Priority 2 Block B: ML Foundations —— The interview main thread, highest ROI
Priority 3 Block C: Deep Learning —— Prerequisite for LLMs
Priority 4 Block E: LLMs —— If your target role requires it
Priority 5 Block D: Engineering & Deployment —— Practice alongside projectsStep 3: Use the "Study Gap List" Template
| Week | Focus Block | Group A tasks (reach level 3) | Group B tasks (reach level 4) | Group C tasks (produce presentations) | Deliverable |
|---|---|---|---|---|---|
| Week 1 | Math + ML Foundations | 3 unfamiliar check points | 2 check points to fill principle gaps | 1 presentation (e.g., overfitting) | Notes + presentation |
| Week 2 | ML Foundations + DL | 2 unfamiliar check points | 3 check points to fill principle gaps | 2 presentations (e.g., AUC, regularization) | Hand-derived notes + presentations |
| Week 3 | Deep Learning | 1 unfamiliar check point | 2 check points to fill principle gaps | 2 presentations (e.g., backprop, attention) | Hand-derived notes |
| Week 4 | LLMs + Engineering | 1 unfamiliar check point | 2 check points to fill principle gaps | 2 presentations (e.g., LoRA, RAG) | Presentations + mini project |
| Week 5+ | Continue until A/B is cleared | Dynamic adjustment | Dynamic adjustment | Progress toward "combination" questions | Mock interview recordings |
Step 4: Three Execution Principles
- Principle 1: Can explain it = you know it. After reviewing each check point, record a 5-minute self-explanation on your phone. Listen back and note where you "stumble"—stumbling spots are where you don't truly understand.
- Principle 2: Hand-derive > watch derivations. For math check points, always derive them yourself (cross-entropy gradient, backpropagation, why L1 is sparse). "Getting it when watching" and "being able to derive it" are two different things in interviews.
- Principle 3: Weekly re-test. Re-score with the self-assessment table every Sunday evening. Cross off check points that "moved up a level," mark those that "stayed the same" and dig into why—usually the problem isn't lack of effort but wrong review methods (e.g., reading without practicing).
One last note
The goal of the study gap list is not to "fill the whole table," but to raise all threshold items from your target JD's requirements to level 4 and above. For blocks not required by the role (e.g., for a recommender role, LLMs don't need to be level 5), level 4 is sufficient. Invest time where it matters.
9. Further Reading
Cross-references on this site:
- Career Module Home · JD Interpretation & Role Selection · Resume Analysis · Interview Questions
- In-depth articles for the five blocks: Supervised Learning · Model Evaluation and Validation · Overfitting & Regularization · Feature Engineering · Deep Learning Foundations · MLOps & Model Deployment · Transformer & NLP · Large Language Models (LLM) · Tuning & Hyperparameter Optimization
- Quick reference: Glossary is a pocket book for plugging gaps; Math Primer is a wall poster for review periods
Reference materials (real sources):
- Andrew Ng. CS229 Lecture Notes (Stanford public course site) —— Classic ML lecture notes
- 3Blue1Brown: Essence of Linear Algebra, Essence of Calculus —— Animated math intuition
- MIT OCW 18.06 Linear Algebra (Gilbert Strang) —— Linear algebra public course
- StatQuest (Josh Starmer) —— Short videos on statistical concepts
- Goodfellow, Bengio, Courville. Deep Learning (2016) —— Deep learning textbook, free on the official site
- Bishop. Pattern Recognition and Machine Learning (2006) —— Classic textbook on statistical machine learning
- Li Hang. Statistical Learning Method (2nd ed., Tsinghua University Press) —— The benchmark Chinese algorithm textbook
- Zhou Zhihua. Machine Learning (Tsinghua University Press) —— Colloquially known as the "Watermelon Book," a classic ML intro in China
- scikit-learn official documentation and User Guide —— Algorithm APIs and official tutorials
- Hugging Face LLM Course —— Free hands-on course on LLM fine-tuning and RAG
- Vaswani et al. Attention Is All You Need (NeurIPS 2017) —— Original Transformer paper
- Touvron et al. LLaMA: Open and Efficient Foundation Language Models (2023) —— Milestone in open-source LLMs