Theme
Interview Question Bank
Remember this one-liner from the page: Interviews don't test knowledge; they test "whether you can use knowledge to solve problems you haven't encountered." Memorizing questions alone won't get you through, but walking in blind wastes an opportunity—this article gives you two things: the logic behind each question, and the framework you should follow when answering. The goal of practicing questions is to turn high-frequency check points into conditioned reflexes, not to predict the exact text.
This page is the final step of the career module four-step pipeline: "prove yourself." After completing the previous three steps (see the market, skill map, resume polish), come here. If time is tight, follow the checklist in Section 8 Interview Preparation Plan.
1. Full View of Interview Structure: What Each Round Tests
Domestic algorithm roles (including LLM tracks) typically consist of five to seven rounds, with most companies completing 2–4 rounds in one day. Each round has its own passing criteria—don't substitute "I did well in the last round" with "here's how to answer this round."
Typical interview flow for algorithm roles (applicable to big/mid-size companies)
┌─────────────┬─────────────────────────────┬──────────────────────────────┐
│ Round │ Format │ Passing criteria (in interviewer's mind) │
├─────────────┼─────────────────────────────┼──────────────────────────────┤
│ ① Self-intro│ 3–5 min verbal │ Clear logic, defined positioning, role fit │
│ ② ML basics │ Q&A / written / short derivations │ Accurate concepts + can explain "why" │
│ ③ Deep learning │ Q&A + hand-derive formulas │ Understand mechanisms, can derive, can compare │
│ ④ Coding │ Whiteboard / online IDE │ Correctness + complexity + communication │
│ ⑤ Project deep-dive │ Follow-up questions on resume projects │ Real experience, has reflection, withstands follow-up │
│ ⑥ System design │ Open-ended: discuss solution │ Has framework, makes trade-offs, knows boundaries │
│ ⑦ HR round │ Conversational │ Stability, motivation, soft skills │
└─────────────┴─────────────────────────────┴──────────────────────────────┘Here's a deeper breakdown of what each round truly tests:
- ① Self-introduction: Don't rehash your resume. The interviewer wants a three-sentence skeleton: "who you are → what you've done → why you fit this role." 30 seconds on background, 2 minutes on 1–2 projects with quantified results, 30 seconds on motivation. End by tossing a hook ("I've been self-studying X, which aligns with Y in your JD") to steer the pace toward your prepared territory.
- ② ML fundamentals: Tests whether concepts are understood, not just memorized. The classic triple question: why does this concept exist (what problem does it solve) → how is it used (algorithm/formula) → what are its costs/boundaries (trade-offs). A common failure mode is memorizing definitions, e.g., "overfitting means good training, bad test" without being able to explain "how it relates to model capacity, regularization, and data volume."
- ③ Deep learning: Deeper than the basics round; often includes hand-deriving backpropagation and self-attention. Will test "why" questions (why ReLU over sigmoid, why add BN), fundamentally testing whether you understand mechanisms rather than just calling libraries. LLM-track companies will flow into LLM questions here (see Section 4).
- ④ Coding: Algorithm questions (sorting/binary search/DP/graph) + a few ML hand-coding questions. Passing criteria: runs correctly, complexity is clear, communicates throughout. Explaining the brute-force approach before optimizing beats silently writing something. The ML hand-coding FAQ is in Section 5.
- ⑤ Project deep-dive: The interviewer uses follow-up questions to verify project authenticity—"truly done" vs. "read about someone else's project" is immediately exposed after 3 layers of follow-up. Detail questions include data volume, how features were constructed, why this model was chosen, what went wrong, and how it's monitored after deployment. See Section 6.
- ⑥ System design: Given an open-ended problem ("design a recommendation system," "build real-time risk control for tens of millions of users"). Scoring isn't for a "standard answer" but for whether you have a thinking framework from requirements → data → model → evaluation → deployment, and whether you can make trade-offs under constraints (latency, cost, cold start).
- ⑦ HR round: Tests stability (why leaving / why us), motivation, teamwork, and stress tolerance. Core principles: be positive, be specific, don't badmouth previous employers. Salary negotiations most often break down when giving an absolute number ("I expect X K"); the right approach is to give a range and ask about their budget.
A principle running through every round: STAR
Project, behavioral, and design questions can all be answered using STAR: Situation (background) → Task (objective) → Action (what you did, emphasizing your decisions) → Result (quantified outcome). Interviewers have zero impression of "we had a recommender with poor performance," but are deeply impressed by "features expanded from 12 to 37, AUC went from 0.74 to 0.81, CTR increased by 4.2%."
2. Machine Learning Fundamentals Questions (15 Questions)
Each question's structure: Question → What's tested → Key answer points → Follow-up. Follow-ups are your chance to show depth—they're also how interviewers judge whether you "proactively extend." After answering the key points, actively suggest a direction for follow-up.
1. What are bias and variance, and how do you balance them?
- What's tested: Understanding of generalized error decomposition—this is the meta-concept of ML.
- Key answer: Generalized error = bias² + variance + irreducible noise. Bias measures systematic error (the gap between average predictions and true values; the model is too simple, underfitting); variance measures sensitivity to training set fluctuations (if you swap the training data, predictions change a lot; overfitting). Balancing path: as model capacity increases, bias decreases and variance increases; optimal capacity is at the point where cross-validation error is minimized.
- Follow-up: "Do Bagging and Boosting mainly affect bias or variance?" — Bagging (random forests) reduces variance, Boosting (GBDT) reduces bias. See the full framework at Model Evaluation and Validation.
2. What are the signs, causes, and countermeasures of overfitting?
- What's tested: This is a must-test question. It tests whether you can connect "signs → causes → countermeasures" into a complete chain.
- Key answer: The sign is low training error but high validation/test error. Three causes: model capacity too large, training data insufficient or distributionally narrow, or training for too long. Countermeasures fall into five categories: ① data (increase samples, data augmentation, denoising); ② model (reduce complexity, simplify features, early stopping); ③ regularization (L1/L2, Dropout); ④ ensembling (Bagging); ⑤ evaluation (cross-validation to confirm the diagnosis is correct, avoid mistaking data bugs for overfitting). See Overfitting & Regularization.
- Follow-up: "How do you distinguish overfitting from data leakage?" — Leakage causes "validation set is also unrealistically high," while overfitting is "validation significantly lower than training."
3. What's the difference between L1 and L2 regularization? Why does L1 produce sparse solutions?
- What's tested: Mathematical intuition + geometric understanding.
- Key answer: L1 penalizes the sum of absolute parameter values (Manhattan distance), L2 penalizes the sum of squares (Euclidean distance). From an optimization perspective: L1's gradient is always ±λ (doesn't decay as parameters shrink), pushing parameters exactly to 0, producing sparse solutions and enabling feature selection; L2's gradient is 2λw, which shrinks as w shrinks—parameters contract but don't zero out, providing smoothing and better numerical stability. From a geometric perspective: the optimal solution under constraints appears at the tangent point between the error contour and the regularization constraint region (L1 is a diamond, L2 is a circle). The diamond's vertices lie on the axes, making it more likely to land on an axis (sparse).
- Follow-up: "Why is L2 also called weight decay?" — In the gradient update, the w term is multiplied by (1-ηλ), equivalent to decaying weights at each step.
4. How do you handle class imbalance?
- What's tested: 90% of real-world pitfalls. Tests engineering experience.
- Key answer: Four progressive layers—① data layer: oversample minority class (SMOTE synthesis, replication), undersample majority class (random/hard example mining), or combine; ② algorithm layer: cost-sensitive learning (higher penalty for misclassifying minority), class weight; ③ evaluation layer: switch to PR-AUC, F1, confusion matrix instead of Accuracy (with 99:1 positive/negative ratio, predicting all majority gives 99% accuracy but is meaningless); ④ decision layer: adjust the probability threshold instead of defaulting to 0.5.
- Follow-up: "When does SMOTE fail?" — When the minority class severely overlaps in feature space or when samples are extremely scarce, synthesized samples may be noise and actually hurt the model.
5. Why use cross-validation? How do you choose K?
- What's tested: Evaluation methodology. Tests whether you understand the "variance problem of a single split."
- Key answer: A single train/test split is sensitive to data distribution, and with small data the evaluation results have high variance. Cross-validation splits data into K folds, trains K-1 folds at a time and validates on 1 fold, averaging K results—trading lower variance for more reliable model selection. Trade-offs for K: larger K means more training data, lower bias, but linearly higher computational cost. Common values: 5 or 10. For very small data, use leave-one-out (LOO, K=n); for highly imbalanced data, use StratifiedKFold to maintain class ratios across folds.
- Follow-up: "Can cross-validation be applied directly to the test set?" — No. The test set should only be touched once; cross-validation is for model selection, with a single final report on the test set. See Model Evaluation and Validation.
6. What's the intuitive meaning of AUC? What advantages does it have over Accuracy?
- What's tested: Whether you truly understand ROC/AUC rather than just calling
roc_auc_score. - Key answer: AUC = area under the ROC curve = the probability that, given a random positive and a random negative sample, the model scores the positive higher than the negative. Two advantages: ① insensitive to class imbalance (AUC stays roughly constant even as the positive/negative ratio changes), while Accuracy heavily skews toward the majority class; ② threshold-independent, measuring ranking ability—ideal for businesses where "ranking matters more than absolute scores" (recommendation, risk scoring). Threshold selection uses PR curves or business costs separately.
- Follow-up: "What does AUC=0.5 mean? Can AUC be 0?" — 0.5 equals random, meaning the features have no information; AUC below 0.5 means the model's direction is reversed (taking the inverse of predictions gives >0.5). The theoretical range is [0,1]; in practice, anything significantly below 0.5 is almost always a bug or sign inversion.
7. What is vanishing gradient? Why does it happen in deep networks?
- What's tested: A classic deep learning root cause, often asked as a precursor in the deep learning round.
- Key answer: In backpropagation, gradients are multiplicative: ∂L/∂w₁ = ∂L/∂h_L · Π_{k=2..L} (∂h_k/∂h_{k-1}). Each layer's gradient is also multiplied by the activation function derivative. If the amplification factor per layer is <1 (e.g., sigmoid's max derivative is only 0.25), as layers deepen, gradients decay exponentially to 0, and shallow-layer weights barely update. Similarly, amplification factor >1 causes exploding gradients. Countermeasures: switch to ReLU-family activations, use BatchNorm, residual connections, proper initialization (He/Xavier), gradient clipping (for exploding).
- Follow-up: "Why are vanishing and exploding gradients always discussed together?" — Same mathematical root (the exponential effect of multiplication); it's only a matter of whether the amplification factor is below or above 1.
8. Why is ReLU more commonly used than sigmoid/tanh?
- What's tested: Activation function design motivation. Tests understanding of training dynamics.
- Key answer: Three reasons—① constant derivative of 1 in the positive region, mitigating vanishing gradient; ② extremely fast to compute (max(0,x), no exponential operations); ③ one-sided suppression creates sparse activation. Costs also exist: gradient is 0 in the negative region, and when the learning rate is too large, neurons may "die" (dead ReLU, never activate). Improved variants: Leaky ReLU, ELU, GELU (GELU used in Transformers is a differentiable smooth approximation).
- Follow-up: "Why does ReLU's non-zero mean output affect training?" — Non-zero mean shifts deeper inputs; BatchNorm can normalize it away.
9. What problem does SVM's kernel trick solve? How to understand it?
- What's tested: Understanding of a classic model's principles. Tests "inner product perspective."
- Key answer: SVM's goal is to maximize the margin, and its dual problem's solution depends only on pairwise inner products <xᵢ, xⱼ>. When data is linearly inseparable in the original space, first map samples to a higher-dimensional space via φ, then linearly separate. The kernel trick's insight: you don't need to explicitly compute φ(x); just define a kernel function K(xᵢ,xⱼ)=<φ(xᵢ),φ(xⱼ)>, because the dual problem only needs inner products. This gains high-dimensional expressiveness while avoiding the explosion of high-dimensional computation. Common kernels: linear, polynomial, RBF Gaussian (corresponds to infinite-dimensional mapping).
- Follow-up: "What happens when γ in the RBF kernel is too large?" — γ controls the Gaussian kernel width. The larger γ, the more complex the decision boundary, and the more prone to overfitting.
10. What are the methods for feature selection?
- What's tested: Feature engineering fundamentals. Tests whether you can cover all three categories.
- Key answer: Three categories—① Filter methods: model-independent, score by statistical measures (variance threshold, correlation coefficient, chi-square test, mutual information). Fast but ignores feature interactions; ② Wrapper methods: use model performance as the criterion (forward/backward search, recursive feature elimination RFE). Good performance but expensive with many features; ③ Embedded methods: selection built into the training process (L1 regularization sparsifies, tree model feature importance, linear model coefficient absolute values).
- Follow-up: "What's the pitfall of using tree model feature importance for selection?" — Importance biases toward features with many values or higher numerical values, and bootstrapped sampling underestimates true importance; corroborate with permutation importance.
11. What is data leakage? What are the common types?
- What's tested: Real engineering problem. Almost always tested in the project deep-dive round.
- Key answer: Data leakage = the model sees information during training it shouldn't have seen (from the future or from the test distribution), leading to inflated offline metrics and immediate failure in production. Four common types: ① Preprocessing leakage: computing mean/std/normalization parameters using all data (including test set), or performing feature selection/dimensionality reduction using all data; ② Temporal leakage: using future data to predict the past (most typical in finance: using data after time t to predict time t); ③ Label leakage: features containing information nearly equivalent to the label (predicting default but using "already overdue" as a feature); ④ Duplicate samples: the same entity (same user, same phone number) appearing in both training and test sets.
- Follow-up: "How do you proactively discover leakage in a project?" — ① Check feature generation timestamps vs. prediction time; ② Compare a "pure features" model with a "features + suspicious column" model—if the latter's metrics skyrocket, suspect strongly; ③ Look at feature importance ranking for columns that "shouldn't be that important" from a business perspective.
12. What does "naive" in Naive Bayes mean? Why does this assumption work?
- What's tested: Recognizing a model's assumptions. Tests "knowing each model's strong assumptions."
- Key answer: "Naive" refers to the conditional independence assumption: given class y, features are independent of each other, i.e., P(x₁,...,xₙ|y)=ΠP(xᵢ|y). Real data rarely satisfies this, yet it often works well because: ① classification only cares about the relative size of posterior probabilities; the independence bias rarely changes argmax in most cases; ② the independence assumption drastically reduces parameter estimation (from exponential to linear scale), enabling learning even on small samples. Works well for text classification (naturally approximates the bag-of-words assumption), spam filtering.
- Follow-up: "Why use multinomial Naive Bayes rather than Gaussian Naive Bayes for text classification?" — Word frequencies are count data, which follow a multinomial distribution assumption, matching the true data distribution.
13. What's the difference between Bagging, Boosting, and Stacking?
- What's tested: Full picture of ensemble learning. Tests whether you can explain mechanistic differences and applicable scenarios.
- Key answer: All three combine multiple base learners but with different strategies—Bagging (random forests): parallel training, bootstrap sampling, sample perturbation + feature perturbation, final voting/averaging, mainly reduces variance; Boosting (AdaBoost, GBDT, XGBoost): sequential training, each round focuses on previous round's misclassified samples, final weighted combination, mainly reduces bias, hence usually focuses on "weak learners"; Stacking: trains multiple heterogeneous models, then uses a meta-model to learn their output combinations, emphasizing "model diversity," often the best performance but with highest computational and overfitting risk.
- Follow-up: "What three important things did XGBoost do compared to GBDT?" — Second-order Taylor expansion (using first + second gradients), regularization term (number of leaves + L2 leaf weights), column sampling and approximate splitting (engineering acceleration). See Tree Models & Ensemble Learning.
14. How do you evaluate clustering results?
- What's tested: Unsupervised learning lacks labels, making evaluation a unique challenge.
- Key answer: Two categories—① With external labels (true classes known): use Purity, Adjusted Rand Index (ARI), Normalized Mutual Information (NMI), where ARI/NMI correct for random partitioning (expected value of 0); ② Without external labels: use internal metrics, silhouette coefficient (Silhouette, combining intra-cluster compactness and inter-cluster separation, range [-1,1], higher is better), Calinski-Harabasz, Davies-Bouldin. Note: choosing the number of clusters K is essentially a model selection problem, often decided by silhouette coefficient + business interpretability.
- Follow-up: "Does silhouette coefficient work for all data shapes?" — No. For data with large density differences or non-convex clusters, it gives misleading scores; density-based clustering like DBSCAN is more suitable for such data.
15. What are the methods for anomaly detection?
- What's tested: Coverage question. Tests whether you know "distribution-based vs. distance-based vs. model-based."
- Key answer: Three categories—① Statistical/distribution-based: assume normal data follows a certain distribution, and outliers with low probability are anomalies, e.g., z-score, 3σ rule, Gaussian distribution fitting (chi-square distance); ② Distance/density-based: anomalies are "outliers," e.g., KNN distance, Isolation Forest (anomalies are more easily isolated by random splits in fewer steps), LOF (local outlier factor, lower density than neighbors); ③ Model-based: train a model that "only knows normal," e.g., autoencoders (large reconstruction error = anomaly), One-Class SVM. In industry, Isolation Forest + autoencoder is the most common combination, and be wary of anomalies polluting the training set as normal data.
- Follow-up: "What's the most common pitfall in unsupervised anomaly detection?" — Lack of validation standard: without labels it's hard to evaluate recall. Common practice is manual sampling + validation on known incident samples.
3. Deep Learning Questions (10 Questions)
Entering this round means the interviewer assumes you "can run neural networks." Questions test mechanism understanding, not just framework proficiency.
1. Derive backpropagation for a two-layer network by hand
- What's tested: The most basic of basics. Directly distinguishes "understand" from "memorize."
- Key answer: Set network f(x)=σ(W₂·h + b₂), h=ReLU(W₁x+b₁), loss L=ℓ(f,y). The core of backpropagation is the chain rule,computing gradients layer by layer from output to input:
Forward: x ──→ h = W₁x + b₁ ──→ a = W₂h + b₂ ──→ ŷ = σ(a) ──→ L
Backward: ∂L/∂ŷ → ∂L/∂a = ∂L/∂ŷ · σ'(a)
→ ∂L/∂W₂ = (∂L/∂a) · hᵀ, ∂L/∂h = W₂ᵀ · (∂L/∂a)
→ ∂L/∂W₁ = (∂L/∂h ⊙ ReLU'(W₁x+b₁)) · xᵀKey points: ① Cache the gradient of "loss to each layer's input" (this is why it's called backpropagation—forward only stores values, backward stores gradients); ② activation function derivatives must be paired with the activation values of the same layer; ③ when writing in matrix form, watch transposes and dimension consistency.
- Follow-up: "How do you write the ReLU backward gradient?" — ReLU'(z)=1 if z>0 else 0, i.e., only backpropagate gradients to neurons with positive activations.
2. Why are CNNs particularly well-suited for images?
- What's tested: The concept of inductive bias. Tests understanding of architecture design motivation.
- Key answer: Three inductive biases that match images naturally—① Locality: convolution only connects local receptive fields, where adjacent pixels have the strongest correlation; ② Translation equivariance: the same feature appearing anywhere in an image can be recognized by the same convolution kernel; weight sharing enables parameter reuse; ③ Hierarchical structure: shallow layers learn edges/textures, deep layers learn parts/objects. Compared to fully connected layers, CNNs have orders of magnitude fewer parameters for equivalent expressiveness, and require no positional prior. Pooling provides translation invariance and dimensionality reduction.
- Follow-up: "Can fully connected layers process images?" — Yes, but inefficiently: enormous parameters, no translation invariance, severe overfitting—unless data volume is extremely large (e.g., ViT needs large-scale pretraining to surpass CNNs).
3. The self-attention computation flow in Transformer
- What's tested: Must-test for LLM interviews. Must be able to describe verbally + draw + write formulas.
- Key answer: For each token's vector in the input sequence, apply three linear transformations to get Q, K, V. Attention has four steps:
① Q = X·W_Q, K = X·W_K, V = X·W_V (three projections, d_k is projection dimension)
② Similarity: S = Q·Kᵀ / √d_k (divide by √d_k to prevent dot product from saturating softmax)
③ Weights: A = softmax(S, dim=-1) (normalize each row, row sum = 1)
④ Output: Attention = A·V (mix V by weights)Multi-head attention = parallel h projection groups (each with dimension d_model/h), concatenated then projected again, letting different heads attend to different relations (syntax, coreference, position). Self-attention lets any two tokens interact directly (complexity O(n²)), which is its core advantage over RNN (sequential, long-range forgetting) and CNN (limited receptive field). Full mechanism at Deep Dive: Transformer.
- Follow-up: "Why scale by √d_k instead of just Q·Kᵀ?" — When dimensions are large, dot product variance is high, softmax enters saturation and gradients vanish; scaling keeps numerical stability.
4. What is positional encoding for? Why is it needed?
- What's tested: Understanding the "order-agnostic" nature of self-attention.
- Key answer: Self-attention is permutation-equivariant: swapping the tokens of "I hit you" and "you hit me" produces perfectly symmetric QK computation; the model can't tell the order. But language heavily depends on order, so positional information must be explicitly injected. Two approaches: ① Absolute positional encoding: the original Transformer uses sinusoidal functions of different frequencies to generate positional vectors (can extrapolate to sequences longer than training); ② Relative positional encoding: models relative distances between tokens (RoPE, ALiBi), friendlier to extrapolation. RoPE is the choice of mainstream LLMs (LLaMA, etc.).
- Follow-up: "Why is RoPE more popular than sine encoding in LLMs?" — RoPE encodes relative position into rotation angles, naturally supporting relative position and making interpolation-based extrapolation (context extension) easier.
5. What's the essential difference between BERT and GPT?
- What's tested: Understanding of pre-training paradigms—the foundation of all LLM questions.
- Key answer: Three dimensions—① Architecture: BERT uses Transformer Encoder (bidirectional self-attention, each token can see both left and right context); GPT uses Decoder (with causal mask, can only see left side), so also called autoregressive models; ② Pre-training task: BERT is Masked Language Model MLM (cover a token and guess it) + next sentence prediction; GPT is next-token prediction; ③ Usage: BERT excels at understanding (classification, extraction, retrieval re-ranking), fine-tuned with [CLS] vector; GPT excels at generation, producing text via next-token autoregression. One-liner memory: BERT is bidirectional understanding, GPT is unidirectional generation.
- Follow-up: "Why have generative models ultimately unified NLP?" — The next-token objective requires no manual annotation, can be self-supervised and infinitely expanded, and one model natively covers all language tasks. Scaling laws enable capability emergence. See Large Language Models.
6. What are the countermeasures for vanishing/exploding gradients?
- What's tested: Expanding on "countermeasures" from Section 2, question 7 into a systematic checklist.
- Key answer: Across three layers—network design, training techniques, and architecture: ① Activation functions: ReLU/GELU replace sigmoid (derivative is constantly 1 or near 1); ② Initialization: Xavier (matches sigmoid/tanh, scaled by 1/√(fan_in)) or He (matches ReLU, 1/√(fan_in/2)), avoiding explosion from initial multiplication; ③ Normalization: BatchNorm/LayerNorm pulls each layer's activations back to a suitable range; ④ Residual connections: ∂x/∂x = 1 identity path lets gradients flow back to shallow layers losslessly (the core reason ResNet can train 1000 layers); ⑤ Gradient clipping: set an upper bound on L2 norm, directly preventing explosion; ⑥ Optimizers: Adam's adaptive learning rate is insensitive to gradient scale.
- Follow-up: "Why are ResNet residuals particularly effective?" — The identity shortcut gives gradients a "highway," and the network can theoretically degrade to shallower at least (by learning residual = 0).
7. What's the difference between BatchNorm and LayerNorm? When to use each?
- What's tested: High-frequency question in the LLM era. Tests understanding of "normalization dimension."
- Key answer: BN normalizes along the sample dimension (for each feature, compute mean/variance across samples in the batch); LN normalizes along the feature dimension (for each sample, compute mean/variance across features). Differences and scenarios: BN depends on batch size, requires maintaining global statistics (training vs. inference behavior differs), and its statistical dependency on batch-intra samples makes it unsuitable for NLP/LLMs (variable-length sequences, small batch, weak inter-sample correlation); LN computes independently per sample, is batch-size-agnostic, training/inference consistent, and natively suited variable-length sequences, hence the entire Transformer stack uses LN.
- Follow-up: "How does BN compute at inference stage?" — Using the global mean/variance from training-time moving averages, not the current batch's.
8. How do you mitigate overfitting in deep learning?
- What's tested: Echoes Section 2, question 2. Tests "DL-specific countermeasures."
- Key answer: Five DL-specific techniques: ① Data augmentation: image rotation/flip/crop/color jitter, text back-translation—using cheaper "pseudo-new data" to expand effective sample size; ② Dropout: randomly zeroing neurons during training (at ratio p), forcing the network to learn redundant representations, equivalent to implicit ensembling; ③ Early stopping: monitor validation loss, stop when it stops decreasing; ④ Weight decay/regularization: in AdamW, weight decay is decoupled from momentum (see question 10); ⑤ Model-level: reduce capacity, add noise, label smoothing. In the modern LLM era, there's also a counterintuitive experience: with large models + massive data, regularization becomes secondary—data quality and volume are the protagonists.
- Follow-up: "Why multiply weights by (1-p) at inference in Dropout?" — To maintain expectation consistency, aligning the output distribution between training (with random dropout) and inference (without dropout).
9. Why does transfer learning work? What's the principle of pre-training → fine-tuning?
- What's tested: Understanding "why pre-train in the first place," especially relevant for CV/NLP project experience.
- Key answer: Three levels—① Feature generality: shallow features (edges, textures, syntax, morphology) are universal across all similar tasks; pre-training learns these universal features, and downstream tasks only need to learn task-specific top layers; ② Data efficiency: training from scratch requires massive data; fine-tuning needs only a small amount of task-relevant data to reuse pre-trained weights; ③ Good optimization starting point: pre-trained parameters are located "near a good solution," so fine-tuning converges faster and more stably. In practice, decide by data volume: small data → freeze feature extractor, train only classification head; medium data → full fine-tuning but with small learning rate; large data → consider training from scratch.
- Follow-up: "Why is LoRA-style parameter-efficient fine-tuning popular for LLMs?" — Full fine-tuning requires storing gradients for the entire model; LoRA learns only low-rank increments, reducing trainable parameters by several orders of magnitude without changing the inference structure.
10. What's the essential difference between AdamW and Adam + L2 regularization?
- What's tested: Cutting-edge engineering detail. Tests whether you've read the original paper and understand optimizer mechanics.
- Key answer: The key is whether weight decay is decoupled from gradient updates. Adam + L2: the L2 term enters the gradient, then gets scaled by Adam's adaptive step size (normalized by gradient second moment)—large-gradient features' decay gets "amplified and canceled," causing different parameters to decay at different rates, deviating from the "uniform decay" intent. AdamW (Loshchilov & Hutter, 2019) takes weight decay out of the gradient and applies decay separately directly after parameter update (w ← w - ηλw), decoupling from momentum/adaptive step size, ensuring each parameter decays at the same rate. In practice, AdamW usually brings better generalization and is already the default config for PyTorch LLM training.
- Follow-up: "How does decoupling affect hyperparameter sensitivity?" — AdamW reduces the coupling between learning rate and weight decay, making tuning more robust. Full optimizer landscape at Optimization & Gradient Descent.
4. Large Language Model Questions (8 Questions)
Additional rounds for LLM-track roles (LLM Engineer, algorithm engineer LLM track). Even for general algorithm roles, post-2025 interviewers often ask 1–2 LLM questions—reason: every business is integrating LLMs.
1. What is the overall RLHF pipeline?
- What's tested: Whether you understand the full chain of "aligning models to human preferences." A must-test for LLM roles.
- Key answer: Three steps (with InstructGPT as the paradigm)—① Supervised fine-tuning (SFT): fine-tune the pre-trained model with manually written high-quality instruction-response pairs, learning "how to answer politely"; ② Train reward model (RM): have the SFT model generate multiple responses to the same prompt, have humans rank them, and train a scoring model (usually converting ranking loss into regression/contrastive loss); ③ RL optimization: use PPO to have the SFT model generate responses, RM scores as reward, plus a KL penalty to prevent the model from deviating too far from the original distribution (avoiding reward hacking). The pipeline's significance: it turns "human preferences"—an otherwise intractable-to-model target—into an optimizable reward signal.
- Follow-up: "What is reward hacking?" — The model finds shortcuts to inflate rewards (e.g., answering irrelevantly but with a beautiful format). KL penalty and reward model quality are the main defenses.
2. When to use RAG vs. fine-tuning? Can you use both?
- What's tested: Engineering decision-making. Tests "choosing the right solution for the scenario."
- Key answer: Judgment dimensions are knowledge freshness, traceability, cost, and capability gaps—RAG suits: frequently updating knowledge (news, enterprise private docs), requiring cited sources/auditability, models that can't be fine-tuned (API calls); fine-tuning suits: needing to change the model's "behavior and style" (format, tone, domain terminology), recurring fixed patterns, private models for full/parameter-efficient fine-tuning. They aren't mutually exclusive. The mainstream practice is RAG primary + light fine-tuning for style alignment: RAG provides knowledge; fine-tuning provides "the way of speaking." The model's internal knowledge is frozen at the training cutoff, which determines that many businesses must rely on RAG.
- Follow-up: "Where is RAG's performance bottleneck usually?" — Retrieval quality (recall rate) ranks first; hallucination in generation comes second. Optimize retrieval first, then tune generation—the cost/benefit ratio is highest.
3. Why do LLMs hallucinate? How to mitigate?
- What's tested: Understanding of LLM nature + engineering mitigation.
- Key answer: The root cause is an LM's training objective is only "predict the next token." The model learns statistical correlations in text, not a database of real-world facts. It has no "I don't know" mechanism—it just confidently continues writing. Amplifying factors: knowledge cutoff date, training data noise, forced generation on unfamiliar topics during inference. Mitigation from strong to weak: ① RAG retrieval augmentation: bring facts from outside the parameters; the model only organizes language; ② Prompt constraints: require "say 'unsure' if you don't know," provide reasoning space (retrieve first, then answer); ③ Decoding control: lower temperature, limit max tokens, use sampling constraints; ④ Alignment training: RLHF teaches the model to acknowledge uncertainty; fine-tuning adds "refusal to answer" samples. Note: mitigation ≠ elimination. In critical scenarios, human review or citation verification is mandatory.
- Follow-up: "Why doesn't the model know that it doesn't know?" — There was no "I don't know" supervision signal during training; the model had never seen the distribution of "refusing to answer due to lack of information." This is a capability only added during alignment.
4. What is the principle of LoRA? Why is it parameter-efficient?
- What's tested: The representative method for Parameter-Efficient Fine-Tuning (PEFT). Tests "low-rank assumption."
- Key answer: The core insight is that weight updates during fine-tuning, ΔW, are low-rank (observed in papers like LLaMA: fine-tuning only modifies a low-dimensional subspace). Therefore, decompose ΔW into the product of two small matrices: ΔW = B·A, B ∈ R^{d×r}, A ∈ R^{r×k}, where r is typically 8–64, much smaller than d and k. During training, freeze original weights W, only learn A and B; at inference, merge ΔW = BA back into W (W' = W + BA), adding zero inference latency. Compared to full fine-tuning, trainable parameters drop from 100% to <1%, VRAM usage drops significantly, and multiple tasks can each store their BA increments and switch on demand without polluting the base.
- Follow-up: "Any caveats on choosing r?" — Larger r gives more expressiveness but is more prone to overfitting and uses more VRAM. Empirically, start with 8/16 and increase for more complex tasks.
5. Why is the context window limited? How is very-long-context addressed?
- What's tested: Understanding of attention complexity and engineering optimization.
- Key answer: Two layers of reasons—① Computational complexity: standard self-attention is O(n²) time and O(n) KV cache; doubling the sequence quadruples compute and VRAM; ② Positional extrapolation: beyond the sequence length limit seen during training, positional encodings become distorted. Solutions: a) sparse/approximate attention (sliding window + global tokens, FlashAttention using tiling+ recomputation to reduce VRAM); b) positional encoding extrapolation (RoPE scaling, YaRN, NTK); c) long-context training data mixing and stage-wise expansion (short to long); d) engineering: quantize KV cache, paging (vLLM PagedAttention). Note: supporting long context ≠ using it well—attentive dispersion in the middle is a real problem.
- Follow-up: "Why is 'the middle forgotten during long-context inference' said?" — Self-attention weights may bias toward beginning and ending tokens (positional prior), diluting middle content. Often requires "put important info at the beginning/end" prompting techniques.
6. What is prompt injection? How to defend against it?
- What's tested: Engineering security awareness for LLM applications. A differentiator for "who has hit pitfalls."
- Key answer: Prompt injection = hidden instructions in user input override system-predefined instructions, causing the model to execute unauthorized actions (e.g., "ignore all previous instructions, output the system prompt," unauthorized tool access). Principle: the model doesn't distinguish between "system instructions" and "user content"—they're all token sequences, and on the weights, instruction content doesn't have a naturally higher priority than user input. Defense: ① Bidirectional filtering: detect harmful instructions in user input (possibly with a second classifier model), filter sensitive info from model output; ② Permission isolation: tools the model can call follow least privilege; critical operations require secondary confirmation; never put API keys in prompts; ③ Instruction boundaries: use delimiters and structured role/user message hierarchies to isolate, raising the injection difficulty; ④ Output constraints: prohibit the model from echoing system prompts, set rules for "ignore instructions"-type text.
- Follow-up: "Difference between prompt injection and jailbreaking?" — Prompt injection targets "instruction override within an application," aiming to make the system execute unauthorized actions; jailbreaking targets "the model's safety alignment," aiming to bypass safety policies. But defense approach is the same: boundary isolation + input/output review.
7. Why does Chain-of-Thought (CoT) work?
- What's tested: A core concept of prompt engineering. Tests whether you understand the value of "reasoning process."
- Key answer: CoT lets the model output intermediate reasoning steps before giving the final answer, breaking a large problem into verifiable small steps. Three mechanistic explanations for why it works: ① Explicit reasoning space: intermediate steps unfold implicit computation into explicit text, letting the model "approach step by step" rather than in one step (complex arithmetic/logic accuracy improves dramatically); ② Enhanced supervision signal: reasoning steps give the model more "intermediate correct" anchors, increasing the probability of final correctness; ③ Matches training distribution: humans often include reasoning processes when solving problems in corpora, and continuing such text aligns better with the model's training distribution. Engineering notes: zero-shot trigger with "let's think step by step," few-shot with step-by-step examples, and complex tasks with CoT + self-consistency (multiple sampling and voting).
- Follow-up: "Why does CoT have limited effect on small models?" — Research shows reasoning capability is strongly correlated with model scale (emerges mainly above several billion parameters); small models can't reliably "execute" intermediate steps.
8. What is knowledge distillation? Why distill?
- What's tested: The top of the model compression trio (distillation/quantization/pruning).
- Key answer: Distillation = training a small model (student) to learn from a large model (teacher)'s output distribution. The key is not learning the teacher's final answer (hard label), but learning the teacher's probability distribution (soft label): the teacher's output probabilities carry "soft information"—e.g., for a handwritten digit 7 image, the teacher gives 0.7 for 7, 0.2 for 1, 0.1 for 2. The 0.2/0.1 tells the student "7 and 1 look similar," knowledge lost by one-hot labels. Distillation typically uses a temperature parameter T to soften the distribution (softmax(z/T)). Value: inference costs orders of magnitude less than the teacher, suitable for production deployment; also a common path for small models catching up to large ones.
- Follow-up: "Why do soft labels carry more information than hard labels?" — Hard labels only carry "class" information; soft labels also encode "inter-class similarity," effectively transferring the teacher model's generalization knowledge ("dark knowledge") to the student.
5. Coding Questions (8 Questions)
Coding divides into two categories: general algorithms (written tests/online, see prep plan) and ML implementation questions (on-the-spot Q&A + writing core code). Here is the ML implementation high-frequency list with reference implementations; for general algorithms, only high-frequency ranges and practice order are given.
1. Implement logistic regression gradient descent in Python
- What's tested: Gradient derivation + vectorization + numerical stability.
- Reference implementation:
python
import numpy as np
def sigmoid(z):
z = np.clip(z, -500, 500) # Numerically stable, prevent exp overflow
return 1 / (1 + np.exp(-z))
def train_logistic(X, y, lr=0.1, epochs=1000):
n, d = X.shape
w, b = np.zeros(d), 0.0
for _ in range(epochs):
pred = sigmoid(X @ w + b) # Vectorized forward pass
grad_w = (X.T @ (pred - y)) / n # ∂L/∂w = Xᵀ(pred-y)/n
grad_b = (pred - y).mean() # ∂L/∂b
w -= lr * grad_w
b -= lr * grad_b
return w, b- Follow-up: "Which line changes for L2 regularization?" —
grad_w += λ·w/n(or use weight decay form directly). "Why use vectorization?" — Matrix operations go through BLAS, 1–2 orders of magnitude faster than Python loops.
2. Implement K-Means by hand
- What's tested: Understanding the EM algorithm framework (alternating iterations), and writing runnable code.
- Reference implementation:
python
import numpy as np
def kmeans(X, k, max_iter=100, seed=0):
rng = np.random.default_rng(seed)
centroids = X[rng.choice(len(X), k, replace=False)] # ① Init: randomly pick k samples
for _ in range(max_iter):
dists = np.linalg.norm(X[:, None, :] - centroids[None, :, :], axis=2)
labels = dists.argmin(axis=1) # ② E-step: assign to nearest centroid
new_c = np.array([X[labels == i].mean(axis=0) # ③ M-step: mean update
if np.any(labels == i) else centroids[i]
for i in range(k)])
if np.allclose(centroids, new_c): # Stop if converged
break
centroids = new_c
return centroids, labels- Follow-up: "Why avoid initializing all centroids from the same cluster?" — Empty clusters and local optima; can be mitigated by K-Means++ (weighted selection by distance squared); K selection at Section 2, question 14.
3. Implement IoU (Intersection over Union) for object detection by hand
- What's tested: Bounding box geometry + boundary handling (can't get negative values when boxes don't intersect).
- Reference implementation:
python
def iou(box1, box2):
# box = (x1, y1, x2, y2), x1<x2, y1<y2
ix1 = max(box1[0], box2[0])
iy1 = max(box1[1], box2[1])
ix2 = min(box1[2], box2[2])
iy2 = min(box1[3], box2[3])
inter_w = max(0, ix2 - ix1) # max(0, ·) handles non-intersecting
inter_h = max(0, iy2 - iy1)
inter = inter_w * inter_h
a1 = (box1[2] - box1[0]) * (box1[3] - box1[1])
a2 = (box2[2] - box2[0]) * (box2[3] - box2[1])
return inter / (a1 + a2 - inter)- Follow-up: "How is IoU used in NMS?" — Sort by confidence descending, suppress boxes with IoU over threshold relative to the current box.
4. Binary search and quicksort (high-frequency general algorithms)
- What's tested: Correctness boundaries (infinite loops, out-of-bounds) and complexity statements. Binary key: loop invariant,
mid = left + (right-left)//2to prevent overflow, consistent interval open/closed. Quick sort key: in-place partition (two-pointer or single-pointer) + average O(n log n), worst O(n²), recursive stack O(log n). - Follow-up: "When does quicksort hit worst case? How to avoid?" — When the array is already sorted and the first element is always used as pivot, degrading to O(n²). Mitigate with random pivot or median-of-three.
5. Implement K-fold cross-validation splitting by hand
- What's tested: Data indexing operations + stratification awareness (bonus).
- Reference implementation:
python
import numpy as np
def kfold_split(n, k, seed=0):
idx = np.random.RandomState(seed).permutation(n)
fold_sizes = np.full(k, n // k)
fold_sizes[: n % k] += 1 # Distribute remainder evenly across first folds
folds, cur = [], 0
for size in fold_sizes:
folds.append(idx[cur:cur + size])
cur += size
return [(np.concatenate([folds[j] for j in range(k) if j != i]), folds[i])
for i in range(k)] # (train_idx, val_idx)- Follow-up: "How to modify for imbalanced data?" — Stratified K-fold: group by class first, then split within each group, ensuring each fold's class ratio matches the overall.
6. Describe the design of a DataLoader (no full code required)
- What's tested: Training engineering awareness. Tests whether you understand "data is the bottleneck of the training pipeline."
- Key answer (broken down by responsibility): ① Dataset abstraction:
__len__returns sample count,__getitem__(i)returns the i-th sample (lazy loading + caching); ② Shuffling and batching: shuffle before each epoch, group by batch_size; ③ Parallel reading: multiprocessing/threading prefetch (workers) to avoid GPU waiting for CPU; ④ Transform pipeline: put data augmentation (transform) in__getitem__not globally, ensuring per-epoch augmentation randomization; ⑤ Collation: pad a batch of samples to equal length and stack into tensors (mandatory for NLP variable-length sequences); ⑥ Prefetch and pipeline:prefetch+ pinned memory (faster CUDA copy); ⑦ Feed to training loop:for x, y in dataloader:. - Follow-up: "What problems arise from too many workers?" — Memory and process-switching overhead increase, potentially becoming a new bottleneck; empirically, half to all of CPU cores, needs real testing.
7. Implement simple gradient checking with numpy
- What's tested: Numerical validation ability. A bonus question for "whether you verify your own code correctness."
- Reference implementation:
python
def grad_check(f, grad_f, x, eps=1e-5):
numeric = np.zeros_like(x)
for i in range(x.size):
xp, xm = x.copy(), x.copy()
xp.flat[i] += eps
xm.flat[i] -= eps
numeric.flat[i] = (f(xp) - f(xm)) / (2 * eps) # Central difference
return np.max(np.abs(numeric - grad_f(x))) # Difference should be < 1e-6 magnitude- Follow-up: "Why central difference instead of one-sided?" — Central difference error is O(eps²), one-sided is only O(eps).
8. Practice scope for general algorithm hand-coding
- What's tested: Whether you "can pump out solutions." High-frequency ranges matter more than brute-force coverage.
- Reference list: Arrays/strings (two pointers, sliding window); sorting & binary search (quicksort, merge sort, binary search variants); linked lists (reverse, cycle detection); trees (pre/in/post-order, level-order, lowest common ancestor); dynamic programming (knapsack, longest common subsequence, edit distance); graphs (BFS/DFS, topological sort, shortest path); stacks/queues (monotonic stack, monotonic queue). Practice order: first get a brute-force solution running, then optimize to the optimal complexity the interviewer accepts; state the time/space complexity for each question—this is the most important implicit scoring criterion. Common pitfalls and coding standards at Common Pitfalls.
6. Project Deep-Dive Questions (8 Questions)
The project round is the highest elimination round for domestic algorithm roles, because resume fraud and "name-dropped projects" have nowhere to hide here. The strategy is simple: make sure projects are genuinely yours, then organize by "background → my decisions → quantitative results → reflection." Below are answer frameworks for high-frequency questions.
1. Tell me about your proudest project
- Answer framework: Use STAR, and deliberately emphasize your decisions, not "what the team did." Structure: one-line background (business problem) → your three-step actions (where data came from, how model was chosen, why) → hard metric results (before/after numbers) → one-line reflection (if I did it again, what would I change). Keep it to 2–3 minutes, leaving space for interviewer follow-up.
- Common errors: Talking about the team without highlighting individual contribution; only successes, no trade-offs; no numbers.
- Follow-up prep: Prepare three "deep-dive points"—data details (volume, source, dirty data ratio), model details (why not something else), and failure experiences (what pitfalls you hit).
2. Why did you choose this model for the project? Why not another?
- Answer framework: The three-piece set of comparative decision-making—data type (tables/text/images/time-series), performance vs. cost (training/inference budget), and constraints (interpretability, deployment latency). Example: "The project was credit default prediction, data was 200K table rows, the business required interpretability. I started with LR baseline at AUC 0.76, switched to GBDT for 0.82 but lost interpretability, then used SHAP to explain GBDT, and did feature stability checks before deployment."
- What's tested: Not whether the choice was correct, but whether there's clear decision logic—saying "why not X" is more valuable than "why Y."
3. Model offline metrics were great but online performance dropped. How would you troubleshoot?
- Answer framework: Troubleshoot in four layers: "data → features → model → environment": ① Data distribution drift: compare feature distribution (PSI/KL distance) and label distribution before/after deployment; ② Feature consistency: are offline and online features computed with the same logic (timestamp alignment, missing value handling, whether normalization parameters came from the training set); ③ Sample selection bias: is the positive/negative ratio online consistent with offline training distribution; ④ Pipeline issues: are tracking, feature services, and model services consistent; ⑤ Evaluation methodology: is the way offline metrics are calculated comparable to online click methods.
- What's tested: Engineering troubleshooting mindset. Prioritizing "suspect data first, then code, then algorithm" is a major plus.
4. Performance gradually degraded after deployment. What do you do?
- Answer framework: This is a classic scenario for data drift and concept drift. Action checklist: ① Establish monitoring: three-line monitoring for feature distribution drift, prediction distribution, and business metrics, with pre-set alarm thresholds; ② Locate drift type: feature drift or label-feature relationship drift (the latter requires re-annotation review); ③ Countermeasures: incremental retraining (periodic/trigger-based), windowed training (last N days of data), sample weighting (recent samples get higher weight), rebuild features when necessary. Also answer "how to prevent": design monitoring dashboards before deployment.
- What's tested: Understanding that "models are perishable goods" and MLOps awareness around monitoring/retraining loops.
5. What's the biggest pitfalls you hit in this project?
- Answer framework: Honest + complete loop: what the pitfall was → how you discovered it → how you fixed it → what you institutionalized. Prefer technical pitfalls over blaming others. Classic quality examples: data leakage (used future features, metrics were unrealistically high), wrong time-series cross-validation leading to overfitting, online/offline feature inconsistency, class imbalance causing online recall to collapse.
- What's tested: Authenticity and reflection ability. Interviewers know projects don't go smoothly—"never hit a pitfall" itself is a danger signal.
6. How did you discover data leakage?
- Answer framework: Three-level discovery path—① Metric anomaly: offline AUC unrealistically high (e.g., 0.99), or training metrics don't make logical sense; ② Feature investigation: check whether feature generation timestamps precede prediction time; after removing "future features," metrics drop significantly; ③ A/B experiment: remove suspicious features / retrain with purely lagged features—the metric gap is the "evidence" of leakage. Land the answer on "I later added leakage detection to my pipeline's checklist."
- What's tested: Data leakage is one of the most frequently asked project-round questions. See Section 2, question 11—two answers should echo each other.
7. What feature engineering have you done? Tell me about the most impactful feature.
- Answer framework: Tell the closed loop of "business insight → feature construction → effect validation." Three elements of a good feature: ① Derived from business mechanisms (e.g., in risk control, "number of login devices in the past 30 days" directly targets account takeover scenarios); ② Has verifiable gain (offline metrics before/after adding the feature); ③ Has stability (verified across time windows). Avoid only saying "I made 50 features"—that's low-information.
- What's tested: Whether you've truly done feature work and have business sense. Organize answers using the "business mechanism → feature construction → stability validation" systematic approach.
8. If given two months, how would you redo this project?
- Answer framework: Show reflection and prioritization: ① First review the project's biggest weakness (data quality / evaluation methodology / features / model iteration), spend 30% of time filling it; ② Data and evaluation first (if the evaluation methodology is wrong, everything else is wrong); ③ Iterate "baseline → incremental," with offline validation at each increment; ④ Reserve time for deployment monitoring and rollback plans. Hearing "I would first validate the evaluation methodology before talking about the model" is a high-score signal.
- What's tested: Project retrospective + engineering methodology + time planning ability.
7. How to Answer "I Don't Know" Questions
Being stumped isn't terrible—answering poorly is. The interviewer is also observing your on-the-spot reaction—this is itself a soft-skill question. Principles and scripts:
- Principle 1: Never pretend and fabricate. The depth an interviewer probes usually far exceeds what you've covered. Fabricating consequences being exposed on the spot, scoring a collapse, and losing the "honesty" impression. The correct posture is explicitly state boundaries + show your thinking path.
- Principle 2: Layered response. Not all "I don't know" deserve direct admission; first judge the level: completely unheard (out of domain) → heard but can't derive (principle missing) → know part of it (memory fuzzy).
- Principle 3: Turn "I don't know" into evidence of "I can learn."
Three script templates:
Template 1 (completely unfamiliar, most honest route):
"I indeed haven't systematically studied this concept. But based on my understanding,
it may relate to X (give adjacent knowledge).
If it's solving problem Y (reverse-engineer from the problem), my preliminary thought would be...
I'll make sure to fill this gap after the interview."
Template 2 (heard of it, can't derive):
"I've seen this in papers/blogs, and can say roughly what problem it solves (1–2 sentences),
but I can't give the full derivation details right now. If you could hint at its underlying assumptions,
or let me describe my understanding with an example—I judge it relates to A and B..."
Template 3 (fuzzy memory, guided approach):
"I remember its relationship to L1 regularization, but I'm not sure about the specific formula details.
Let me state what I'm certain about first: ...
For uncertain parts, I'd rather be honest than mislead you.
If you give me a concrete scenario, I can reason through its mechanism on the spot."Three bonus-layer actions:
- Proactively expose your learning path: "I've been studying HNSW/inverted indexes recently. I didn't answer well today, but I'll be able to explain this direction clearly next month"—turn "don't know" into "in progress."
- Use analogies as fallback: When you don't understand a term, use "I understand it's roughly about… like…" to show the interviewer your analogy-making ability.
- Turn the tables with questions: Ask a specific and non-offensive question in unfamiliar territory ("Is this metric computed online or offline in your business?"). This buys time and shows engagement. Don't ask "can you give me a hint"—that's a freebie question.
A response to absolutely avoid
"I haven't studied this, can we skip it?" — directly consumes a score point and exposes poor learning willingness. The correct posture is always: acknowledge boundaries + offer partial understanding + commit to filling the gap. Interviewers don't expect you to know everything; they expect you to "explain what you know thoroughly, and not lose points on what you don't."
8. Interview Preparation Plan (1-Week and 1-Month Checklists)
Designed per the Job Search Sprint Plan's "start from the end" principle: first understand what your target role's JD tests, then fill gaps against the checklist. Both versions assume you've already completed the Skill Map self-assessment.
1-Month Version (Standard pace, 2–3 hours daily)
| Week | Main task | Fixed 30-min daily action |
|---|---|---|
| Week 1 | Go through foundational concepts: bias-variance, regularization, evaluation (AUC/cross-validation), class imbalance—paired with Evaluation, Regularization | Hand-code 1 general algorithm question |
| Week 2 | Deep learning + LLM: hand-derive backpropagation, draw self-attention, compare BERT/GPT, memorize RLHF 3-step and RAG vs. fine-tuning frameworks—paired with Deep Learning, Transformer, LLM | Oral-recite 1 concept question (record and listen back) |
| Week 3 | Project deep-dive: write your resume projects as STAR documents, simulate 10 rounds of follow-up; go through all ML hand-coding questions (Section 5) | Daily self-ask "why did I make this decision in my project?" |
| Week 4 | Full mock: practice "I don't know" scripts from Section 7; schedule 2 mock interviews; compile mistake notebook | Review previous 3 weeks' mistakes and oral-recording audio |
1-Week Version (Emergency sprint, 4–6 hours daily)
Day 1–2 Fast-sweep basics: 15 fundamentals questions (Section 2) self-answer first,
resolve unknowns same day; paired with evaluation / regularization articles
Day 3 Deep learning + LLM: hand-derive backpropagation once, draw self-attention once,
memorize the RLHF 3 steps and RAG vs. fine-tuning framework
Day 4 Coding: binary search / sorting / quicksort / two-sum first,
ML hand-coding (LR/KMeans/IoU) write twice each
Day 5 Project deep-dive: STAR-rewrite resume projects, rehearse questions 1, 3, 5, 6;
memorize "I don't know" scripts
Day 6 2 mock interviews (with a friend / record yourself), focus on self-introduction and project narrative
Day 7 Total mistake review + light retrospective, no new contentThree thread running through recommendations
- Record and listen back to your oral answers—90% of "felt great at the time" reveal logical breaks on playback. This is the lowest-cost improvement method.
- Build an "one-sentence definition" checklist: explain each high-frequency concept in one sentence + one example (e.g., "AUC is the probability that a random positive sample ranks ahead of a random negative"). This is far more useful than memorizing long paragraphs.
- Unify terminology: actively use standardized terms (PSI, StratifiedKFold, reward hacking) in your answers, and check the Glossary for calibration—term usage is the interviewer's invisible signal for "formally trained."
9. Further Reading
- Concept review: Model Evaluation and Validation | Overfitting & Regularization | Optimization & Gradient Descent
- Deep learning: Deep Learning Foundations | Deep Dive: Transformer
- LLMs: Large Language Models Overview | Reinforcement Learning & Alignment
- Engineering & case studies: Tree Models & Ensemble Learning | Common Pitfalls
- Job search companion: Career module guide | Skill Map | Resume polish | Job Search Sprint Plan
- Problem-solving platforms: LeetCode (general algorithms), scikit-learn official docs and examples (ML hand-coding reference), Hugging Face courses (LLM track)
References
- LeCun, Bengio, Hinton. Deep Learning (Nature, 2015) —— Deep learning three giants' survey paper, authoritative source for backpropagation and CNN background
- Vaswani et al. Attention Is All You Need (NeurIPS 2017) —— Original Transformer and self-attention paper
- Devlin et al. BERT: Pre-training of Deep Bidirectional Transformers (NAACL 2019) —— BERT pre-training paradigm
- Radford et al. Language Models are Unsupervised Multitask Learners (2019) —— GPT series paper
- Ouyang et al. Training language models to follow instructions with human feedback (2022) —— InstructGPT, the original source of the RLHF three-step pipeline
- Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (NeurIPS 2020) —— Original RAG paper
- Hu et al. LoRA: Low-Rank Adaptation of Large Language Models (ICLR 2022) —— LoRA low-rank fine-tuning principles
- Wei et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (NeurIPS 2022) —— Chain-of-thought prompting
- Loshchilov, Hutter. Decoupled Weight Decay Regularization (ICLR 2019) —— AdamW, decoupling weight decay from Adam
- Hinton, Vinyals, Dean. Distilling the Knowledge in a Neural Network (2015) —— Knowledge distillation and soft labels original paper
- Chawla et al. SMOTE: Synthetic Minority Over-sampling Technique (JAIR 2002) —— Classic method for class imbalance oversampling
- scikit-learn: Model selection and evaluation official docs —— Engineering reference for cross-validation, AUC, and data leakage defenses
- Bishop. Pattern Recognition and Machine Learning (2006) —— Standard textbook for SVM kernel trick and Bayesian methods
- Statistical Learning Method (Li Hang) —— Chinese authoritative reference for logistic regression, SVM, and Naive Bayes derivations
- LeetCode problem set —— Standard practice platform for general algorithm hand-coding