Theme
Interview Question Bank
In one sentence: a categorized deep learning interview question bank organized across seven dimensions — each question provides "answer key points + linked pages," serving as both a pre-interview self-test checklist and an exam paper to verify the study outcomes from the JD Knowledge Map.
Usage tips: Don't memorize questions in order. First pick weak dimensions from your study plan, then for each question, talk through it yourself first (record audio or practice in front of a mirror) before checking the key points. Each question should take 3 minutes to answer fully and withstand one follow-up to pass.
1. Fundamentals (must-know, failing means automatic rejection)
1. What is backpropagation? Why is it efficient?
Answer key points: Backpropagation uses the chain rule to reuse intermediate gradients in an "output → input" order. One forward pass + one backward pass computes gradients for all parameters. The complexity scales with the number of parameters, not exponentially with the number of layers. Compare this to numerical differentiation (which requires O(n) forward passes) to explain the efficiency gap.
Linked page: Backpropagation & Automatic Differentiation.
2. How do you choose activation functions? Why is ReLU so popular?
Answer key points: ReLU is fast to compute, mitigates vanishing gradients (derivative is 1 in the positive region), and introduces sparsity. Its weakness is that the derivative is constantly 0 in the negative region, causing "dead neurons" — can be paired with Leaky ReLU. Sigmoid is only used for binary classification output layers. Tanh outputs zero-centered, making it better than Sigmoid for hidden layers. The output layer activation is task-dependent: linear for regression, Sigmoid for binary classification, Softmax for multi-class. See Loss Functions & Output Layers.
Linked page: Neural Network Fundamentals.
3. What do you do when your model is overfitting?
Answer key points: Answer in priority order — add data / data augmentation, reduce model capacity, add regularization (weight decay, Dropout), early stopping, ensemble. Explain why each works (increases generalization, reduces variance). Ideally, walk through a real overfitting debug session from your own project. See Overfitting & Regularization.
4. What's the problem when training loss is low but validation loss is high, versus when both are high?
Answer key points: The former is overfitting (high variance), the latter is underfitting / insufficient model capacity (high bias). Diagnose first, then treat: overfitting → add regularization / data; underfitting → increase capacity / more training epochs. See Deep Learning Evaluation & Experimentation for evaluation methodology.
5. What is a typical ML workflow? Why normalize features during preprocessing?
Answer key points: Data → preprocessing → modeling → evaluation → iteration. Normalization brings all features to a consistent scale, preventing gradient updates from being dominated by large-scale features and accelerating convergence — tree models are unaffected by this. See Data & Data Engineering for more data-side details.
2. Architectures
6. Write the attention formula in Transformer and explain each part. (must-know)
Answer key points: Attention(Q, K, V) = softmax(QKᵀ/√d_k)V. Q, K, V are derived from inputs via linear projections. QKᵀ measures similarity. Dividing by √d_k prevents dot products from growing with dimension (which would cause softmax saturation). Softmax normalizes into weights. Weighted sum over V. Deriving why dividing by √d_k is a bonus point.
Linked pages: Transformer Architecture, Attention Mechanism.
7. What is a receptive field? How do you increase it?
Answer key points: The receptive field is the size of the input region that corresponds to a given output neuron. Ways to increase it: deepen the network, use pooling / strided convolutions, dilated convolutions. Stacking deep small conv kernels can approximate large kernel effects with fewer parameters. See CNNs & Computer Vision.
8. What's the difference between RNNs and Transformers in processing sequences? What does positional encoding do?
Answer key points: RNNs process sequentially, can't parallelize, and suffer from vanishing gradients on long-range dependencies (LSTM/GRU alleviate). Transformers use self-attention for global modeling and can parallelize, but lose sequential information — positional encoding compensates. The difference between absolute positional encoding (sinusoidal) and relative positional encoding (RoPE, etc.) is an advanced bonus point. See RNN & Sequence Modeling.
9. Why are convolutions better than fully connected layers for images?
Answer key points: Local connectivity + weight sharing (translation equivariance) + fewer parameters, matching the prior of local correlation in images — the concept of inductive bias. A highlight is explaining "this is why ViT needs more data to catch up to CNNs" — the inductive bias gap.
3. Training
10. What's the difference between Adam and SGD? When to use each?
Answer key points: Adam has adaptive learning rates (first and second moment estimates), converges faster, and is less sensitive to initial learning rate. SGD + Momentum converges slower but generalizes better and is more sensitive to hyperparameters. In practice: AdamW is commonly used for LLM fine-tuning; switch to SGD later in tuning or when generalization is the priority. Explaining why AdamW corrects Adam's weight decay implementation is a bonus point. See Optimization & Gradient Descent.
11. What happens if the learning rate is too large or too small? What are common LR schedules?
Answer key points: Too large → non-convergence / oscillation. Too small → slow convergence or getting stuck in local minima. Common schedules: step decay, cosine annealing, warmup (essential for LLM training), ReduceLROnPlateau. Explaining why warmup is important for Transformers (large gradient variance in early training) is a bonus point.
Linked page: Training Recipes & Hyperparameter Tuning.
12. How do vanishing/exploding gradients occur, and how do you fix them?
Answer key points: Deep networks involve multiplying gradients during backpropagation. If activation derivatives are < 1 or weight matrix spectral radius > 1, gradients decay/amplify exponentially. Fixes: proper initialization (Xavier/Kaiming), normalization (BN/LN), residual connections, ReLU-family activations, gradient clipping (for exploding). See Initialization & Normalization and Debugging & Diagnosis.
13. What's the difference between BatchNorm and LayerNorm? Why does Transformer use LN?
Answer key points: BN normalizes across the batch dimension, depends on batch statistics, behaves differently in training vs inference, and is sensitive to batch size. LN normalizes across the feature dimension, independent of batch. Transformer uses LN because sequence lengths vary and batches are often small due to VRAM constraints. Bonus points for discussing why "BN fails on NLP but works on images."
4. Math
14. Matrix calculus: Outline the derivation of dL/dW in general.
Answer key points: Use the chain rule for scalars w.r.t. matrices, unfolding layer by layer from the scalar loss. Being able to write dL/dW = dL/dy · xᵀ for y = Wx, L = f(y) is sufficient. The key is dimension alignment (forward pass shape matches gradient shape). Systematic review: Math Primer.
15. Why does cross-entropy loss pair with softmax?
Answer key points: Softmax converts logits into a probability distribution and is differentiable. The derivative of cross-entropy w.r.t. softmax inputs yields the clean (p - y) form (prediction minus true one-hot), with simple and stable gradients. This avoids the gradient saturation problem of MSE + softmax. Derivation: Loss Functions & Output Layers.
16. What are conditional probability and Bayes' theorem? How do you interpret posterior probability in classification?
Answer key points: P(A|B) = P(B|A)P(A)/P(B). Classification outputs can be viewed as posterior probability P(y|x). Explaining why Naive Bayes is "naive" (the conditional independence assumption of features) is a bonus point. See Math Primer.
17. What is KL divergence? What role does it play when training GANs?
Answer key points: KL divergence measures the difference between two distributions — it is asymmetric, ≥ 0 (equals 0 only when distributions are identical). The GAN generator minimizes some divergence between the generated and true distributions (the original GAN is equivalent to minimizing a variant of JS divergence). The variational lower bound in diffusion models is also related to KL. See VAE & GAN.
5. Engineering
18. What are common data augmentation methods? When can't you use them?
Answer key points: Images — flip, crop, rotate, color jitter, MixUp/CutMix. Text — back-translation, synonym replacement, EDA. But watch task semantics: OCR can't have horizontally flipped text, medical images can't be arbitrarily cropped. See Data & Data Engineering.
19. What is the principle behind mixed-precision training (FP16/BF16), and what are the pitfalls?
Answer key points: Using low precision for storage/computation accelerates training and saves VRAM. The key component is loss scaling to prevent underflow. BF16 and FP16 differ in exponent/bits allocation — BF16 is more stable for LLMs. Pitfalls: some operators are precision-sensitive and must remain in FP32 (e.g., BN statistics, certain normalization layers). See Training Recipes & Hyperparameter Tuning.
20. How do you optimize when model VRAM is insufficient?
Answer key points: Reduce batch size / gradient accumulation, mixed precision, gradient checkpointing (trading compute for VRAM), reduce model size (distillation/pruning/quantization), optimize data loading. Giving concrete numbers (e.g., gradient checkpointing saves 60% VRAM but slows by 30%) is a bonus point.
21. What's the difference between training and inference? What are inference optimization techniques?
Answer key points: Training involves backpropagation, needs gradients, often distributed; inference only needs forward pass, pursuing low latency and high throughput. Inference optimizations: quantization (INT8), distillation, pruning, operator fusion, KV cache, batch scheduling. See MLOps & Model Deployment.
6. Large Language Models
22. What's the difference between fine-tuning and RAG? How do you choose?
Answer key points: Fine-tuning changes model parameters (knowledge/style injection). RAG doesn't change parameters — it relies on retrieving external knowledge (strong timeliness, interpretable, cheap to update). Selection logic: knowledge timeliness / privacy → RAG; behavioral style / task format → fine-tuning; often used in combination. See Large Language Models (LLMs).
23. Why do LLMs hallucinate? How do you mitigate it?
Answer key points: Fundamentally, language models generate text that is "probabilistically reasonable" rather than factually retrieved. Mitigations — RAG introduces evidence, fine-tuning for alignment, decoding constraints, evaluation and fallback strategies. Distinguishing between "factual hallucination" and "reasoning errors" is a bonus point.
24. What is KV cache? Why is it important?
Answer key points: During autoregressive decoding, the Key/Value of already-generated tokens can be cached and reused, avoiding redundant computation in exchange for VRAM overhead. In long-context scenarios, KV cache becomes a VRAM bottleneck, leading to optimizations like MQA/GQA and quantized KV. This is a high-frequency deep-dive topic in "LLM inference interviews."
25. Why does LoRA save VRAM?
Answer key points: It freezes the original weights and only trains a low-rank increment ΔW = BA (B ∈ R^{d×r}, A ∈ R^{r×k}, r ≪ d), reducing trainable parameters by several orders of magnitude, with plug-and-play task switching. Explaining the relationship between rank r selection and task is a bonus point. See Large Language Models (LLMs).
7. Open-Ended Questions (testing system design)
26. How would you design an image classification system?
Answer key points: Walk through the full pipeline — requirements and metrics (accuracy/latency/cost trade-offs) → data (collection / annotation / quality check / class balance) → model selection (CNN vs ViT, scale) → training (transfer learning, tuning) → evaluation (layered metrics, badcase analysis) → deployment (serving / edge, monitoring and rollback). The testing point is "whether you have complete closed-loop thinking," not whether any single step is done to perfection. Reference Building a Deep Learning Project from Scratch.
27. The model performs worse online than offline. How do you troubleshoot?
Answer key points: Data distribution drift (feature missing / value changes) → online implementation bugs (preprocessing inconsistency) → evaluation metric inconsistency (offline vs online label definitions) → business environment changes (user behavior shifts). A "troubleshooting checklist sorted by probability" is a high-scoring answer. See Debugging & Diagnosis.
28. How would you design an evaluation plan for your model?
Answer key points: Clarify evaluation objectives (capability / robustness / fairness) → choose metrics and benchmark sets → design adversarial and boundary tests → automated evaluation pipeline → periodic regression. Referencing the Evaluation in Practice principle of "choose metrics before experiments" is a bonus point.
8. Interview Etiquette and Answering Rhythm
- Give the conclusion first, then expand: One-sentence conclusion + 2–3 reasons + one example;
- Be honest when you don't know: "I haven't worked in this area, but my understanding is…" beats making things up;
- Prepare 2 projects you can discuss deeply: Each project should sustain 15 minutes of continuous discussion withstanding 5 rounds of follow-ups — more valuable than 5 projects you can't discuss deeply;
- Prepare 2–3 questions for the interviewer's turn: Ask about team structure, evaluation criteria, tech stack — demonstrating genuine interest in the role.
Further Reading
- JD Knowledge Map — Locate weak dimensions based on your study plan
- Capability Match: What to Highlight in Your Resume — Self-test every word on your resume before an interview
- Training Recipes & Hyperparameter Tuning — System foundation for training-related questions
- Math Primer — A refresher tool for math questions
- Large Language Models (LLMs) — Knowledge base for LLM-related questions
- Evaluation in Practice — Methodology for open-ended and evaluation questions