Theme
Interview Question Bank
Remember this one-liner from this page: interviews don't test how much you've memorized — they test how deep you can explain "why" under pressure. So every question here provides "reference answer key points" rather than script-like answers. For the underlying principles behind each point, go back to the linked page for deep reading. This bank organizes 40 questions by topic, covering the seven blocks of LLM role interviews.
How to use this bank
① Self-assess first using the Knowledge Breakdown, then come back to practice; ② Cover the answers and answer yourself first, then compare against the key points — anything you can't say is a gap; ③ You can only pass a question if you can explain it for a full 3 minutes without this page (including one example).
1. Interview Structure Overview: What Each Round Tests
The common structure for LLM role interviews (varies widely by company, but the skeleton is consistent):
| Round | What's Tested | Related Section | Bank Recommendation |
|---|---|---|---|
| Written test / Coding | Coding problems (algorithms + LLM-related implementations) | VII. Coding Questions | Practice high-frequency problems repeatedly |
| Fundamentals round | Transformer, training, inference principles | II–V | Explain "why," don't memorize |
| Project deep-dive round | Tradeoffs, pitfalls, quantified results from your resume projects | VIII. Open-ended + Resume Analysis | Prepare a 20-minute deep explanation |
| System design round | End-to-end RAG/Agent design | VIII. Open-ended | Explain tradeoffs and metrics clearly |
| HR round | Motivation, stability, salary expectations | — | Honest + clear decision logic |
2. Fundamentals & Transformer (8 Questions)
1. Write the scaled dot-product attention formula and explain each step
Key points:
Attention(Q, K, V) = softmax(Q·Kᵀ / √d_k) · V- Q (query) determines "what I'm looking for," K (key) determines "what I have that can be found," V (value) determines "what to give once found";
Q·Kᵀcomputes similarity scores between all key-value pairs (larger dot product = more relevant);- Dividing by
√d_kscales the values (see next question); softmax normalizes to weights; weighted sum·Vproduces the output. - Related: Transformer Architecture Explained.
2. Why divide attention by √d_k?
Key points:
- The variance of the dot product
q·kgrows with dimension d_k (≈ d_k); as values get large, softmax enters its saturation zone and gradients approach 0, slowing or preventing convergence; - Dividing by √d_k restores variance to O(1), keeping softmax in the high-gradient region;
- One-liner: control variance, preserve gradients.
3. What's the purpose of multi-head attention?
Key points:
- A single attention head has only one attention pattern; multi-head projects Q/K/V into multiple lower-dimensional subspaces, so each head can attend to different relationships (syntax, coreference, long-range dependencies);
- Head outputs are concatenated and linearly projected, achieving feature complementarity;
- Intuition similar to multi-channel CNNs; multi-head parallel computation costs roughly the same as a single head (total compute is unchanged).
4. What positional encodings exist for Transformer? What if RoPE can't extrapolate?
Key points:
- Attention itself is permutation-equivariant — positional information must be injected: absolute positional encoding (sin/cos), learned position embeddings, relative positional encoding, RoPE (Rotary Positional Encoding), ALiBi (attention with linear biases);
- RoPE encodes position into Q/K using rotation matrices; during extrapolation (beyond trained length), performance drops sharply because rotation frequencies exceed the learned distribution;
- Solutions: position interpolation (linear/NTK-aware interpolation, YaRN) compresses position coordinates back into the training range at inference time; up-sample long text during training to extend the window; or simply use RAG instead of forcing long context (see Context Window & Long Context).
5. What is KV cache? Why is it needed at inference?
Key points:
- During autoregressive decoding, generating token t+1 requires attention over the previous t tokens; each token's K and V values depend only on itself and its position, not on future queries, so they can be cached and reused;
- Without caching, you'd recompute Q/K/V for the entire prefix at every step, causing computation to grow quadratically with sequence length;
- KV cache trades memory for compute: 2 × batch × layers × heads × head_dim × seq_len × 2 bytes (FP16);
- Cost: linear memory inflation for long sequences — requires PagedAttention, quantization, upper-bound constraints, etc. (see Inference Fundamentals).
6. Decoder-only / Encoder-only / Encoder-Decoder: what's the difference?
Key points:
- Encoder-only (BERT): bidirectional attention, excels at understanding/classification/retrieval, weak at generation;
- Decoder-only (GPT/Llama/Qwen): causal masked autoregressive decoding, can both understand and generate — the contemporary mainstream;
- Encoder-Decoder (T5): encode then decode, strong for sequence-to-sequence tasks like translation/summarization;
- Why Decoder-only won: unified objective, simplicity, and excellent scaling (detailed discussion in Language Modeling).
7. What is the complexity of self-attention? Why is long context expensive?
Key points:
- Time O(n²·d) (n = sequence length), memory O(n²) — every token interacts with every other token;
- With long context: computation, KV cache memory, and the quadratic attention matrix all expand linearly or quadratically;
- Mitigations: FlashAttention (I/O optimization, doesn't store the full matrix), sparse/sliding-window attention, long-context position interpolation, and using RAG in business instead of force-feeding the entire document.
8. Pre-Norm vs Post-Norm: what's the difference?
Key points:
- Post-Norm (original Transformer): residual add then LayerNorm — gradients are unstable in deep layers, requires warmup;
- Pre-Norm (modern mainstream, GPT/Llama): LayerNorm then residual — more stable training, can eliminate warmup;
- Tradeoff: Pre-Norm has a slightly lower theoretical upper bound than Post-Norm but is much easier in practice, making it the de facto standard.
3. Training (6 Questions)
9. What is the objective of pretraining? How does it differ from fine-tuning?
Key points:
- Pretraining objective = next-token prediction (autoregressive language modeling), maximizing conditional probability over massive unlabeled text;
- Pretraining learns general language capabilities and knowledge; fine-tuning (SFT) adapts the model to specific tasks/styles using instruction-response data;
- Key differences: data (unlabeled vs. labeled instructions), objective (language modeling vs. instruction following), and massive scale/cost differences;
- Pretraining is often compared to "reading ten thousand books," while fine-tuning is "learning a specific skill" (see Fine-tuning).
10. What is perplexity? What's the relationship with cross-entropy?
Key points:
- Perplexity = exp(cross-entropy) = 2^cross-entropy; it measures the model's average "uncertainty" over a text;
- Intuition: perplexity approximates how many candidate words the model is uncertain about at each position on average;
- Only comparable under the same tokenizer and corpus distribution; perplexity is not always positively correlated with downstream task performance — it's a "chimney metric" (internal indicator), not a "treatment metric" (effectiveness indicator).
11. How do you determine data mixing ratios? What happens if you add too much code/math data?
Key points:
- Mixing ratios determine capability structure: general text gives language ability, code gives reasoning/structured thinking, math gives rigorous reasoning, multilingual gives language coverage;
- Upsampling code/math data → improved reasoning ability, but may damage conversational naturalness and language style;
- Mixing ratios require empirical validation: use ablation experiments on small models first, then scale up; excessive domain upsampling leads to domain overfitting and general capability degradation;
- Resources: Datasets & Benchmarks.
12. Explain scaling laws and the Chinchilla conclusions
Key points:
- Kaplan (2020): test loss decreases as a power law over model parameters N, data volume D, and compute C;
- Chinchilla (2022): at compute-optimal ratios, token count ≈ 20 × parameter count (e.g., a 70B model needs ~1.4T tokens);
- Implication: given fixed compute, data and parameters should grow together — don't just stack parameters; in practice, engineers often "over-train" small models or "under-train" large ones to accommodate compute and inference cost constraints;
- Emergent abilities remain controversial: are they sudden capability jumps or smooth metric transitions (see Scaling Laws)?
13. What problems do data parallelism, tensor parallelism, pipeline parallelism, and ZeRO each solve?
Key points:
- Data Parallelism (DP/FSDP): each GPU holds a complete model copy, splits data — solves "data too large to process on one machine," but ineffective when a single GPU can't fit the model;
- Tensor Parallelism (TP): splits single-layer weights across GPUs by dimension, dense intra-layer communication, best for single-machine multi-GPU (NVLink);
- Pipeline Parallelism (PP): splits by layer, each GPU handles a segment, reducing communication but introducing pipeline bubbles;
- ZeRO: sharded optimizer states/gradients/parameters — can fit massive models under data parallelism, communication increases with shard level;
- Combo: PP + TP handle "model won't fit," DP handles "data too much," ZeRO handles "memory" (see Framework & Tool Selection).
14. How do you troubleshoot training that doesn't converge / has abnormal loss?
Key points:
- Check data first: any contamination/NaN in batches, label alignment issues, tokenization anomalies;
- Check numerical stability: is loss NaN (learning rate too high, mixed-precision overflow → lower LR / switch to BF16, add grad clip);
- Check dynamics: loss plateau (LR schedule, data repetition, insufficient model capacity), oscillation (LR too large, batch too small);
- Engineering checks: gradient checking, checkpoint recovery consistency, silent overflows;
- Iron rule: reproduce before attributing; change only one variable at a time.
4. Alignment (5 Questions)
15. What is the full RLHF pipeline?
Key points:
- Three steps: ① SFT — fine-tune the base model with high-quality instruction-response data; ② Train a reward model (RM) — collect human preference rankings across multiple responses to the same prompt, train a model that "scores responses"; ③ PPO reinforcement learning — have the SFT model continue optimizing in the direction of "getting high RM scores," while using KL penalty to constrain it from straying too far from the SFT model;
- Key details: preference data scale (typically much smaller than SFT data), RM and policy model share architecture, PPO needs a ref model to compute KL;
- See Alignment: RLHF & DPO.
16. How is the reward model trained? Where does the data come from?
Key points:
- Data: humans rank multiple responses to the same prompt (e.g., A > B > C), produced by a labeling team;
- Training: convert rankings to pairwise comparisons, fit with a Bradley-Terry model (pairwise softmax probability), loss is the negative log-likelihood of ranking accuracy;
- Pitfall: reward models can over-optimize (reward hacking), and human scores themselves have noise and bias; reward model scale is often comparable to the policy model.
17. What do clip and KL penalty do in the PPO objective?
Key points:
- Policy updates use importance sampling ratio
r = π_θ(a|s) / π_old(a|s); - clip: clamps the ratio to [1-ε, 1+ε], preventing the policy from updating too far in one step and collapsing training (the core of PPO stability);
- KL penalty: keeps the new policy from straying too far from the reference policy (SFT/ref model), preventing language collapse (where the model stops sounding human);
- Together: move toward higher reward, but take small steps and don't get lost.
18. What's the difference between DPO and RLHF? Why doesn't DPO need a reward model?
Key points:
- DPO (Direct Preference Optimization) optimizes the policy directly from preference pairs (chosen/rejected), implicitly using a "reward function" (the log-ratio of policy to reference policy) to replace an explicit RM;
- Mathematical intuition: under the optimal policy, reward can be written as
r(x,y) ∝ log(π_θ/π_ref), substituting this reward into the Bradley-Terry loss yields the policy loss directly; - Difference: DPO skips RM training and RL loops, making it simpler, cheaper, and more stable; but explicit RMs allow offline iteration and are more interpretable, and RLHF's upper bound may not be lower than DPO's;
- Practice: DPO is sensitive to preference data quality and requires overfitting prevention (see Alignment).
19. What is "alignment tax"? Why does alignment make models dumber?
Key points:
- Alignment tax: the phenomenon where models' scores on some benchmarks (especially reasoning and factual tasks) drop after alignment (SFT/RLHF);
- Causes: alignment data has a narrower distribution than pretraining data, compressing the original pretraining distribution; the optimization target is "please humans" rather than "maximize correctness"; KL constraints squeeze out capabilities;
- Mitigations: use stronger base models for alignment, control alignment data ratios, use lower-interference methods like DPO, run post-alignment evaluation to prevent over-optimization.
5. Inference (5 Questions)
20. How do you estimate KV cache memory?
Key points:
- Formula:
KV memory ≈ 2 (K and V) × num_layers × num_kv_heads × head_dim × seq_len × batch × bytes_per_element - Example: 7B model (32 layers, 32 heads, head_dim 128, FP16), sequence 2048, batch 1:
2×32×32×128×2048×2 B ≈ 1.07 GB; - Note: grouped attention (GQA, e.g., Llama's KV heads < query heads) can significantly compress; for long sequences, KV cache often uses more memory than weights.
21. What is quantization? What's the difference between GPTQ / AWQ / GGUF?
Key points:
- Quantization: reducing weights/activations from FP16 to INT8/INT4, trading lower precision for less memory and faster inference, at the cost of accuracy loss;
- GPTQ: post-training quantization (PTQ), uses Hessian-weighted per-layer minimization of quantization error, suitable for GPU deployment;
- AWQ: protects important weight channels based on activation distribution, balanced quality/speed, also GPU-friendly;
- GGUF: llama.cpp ecosystem's quantization format (Q4_K_M, etc.), targeting CPU/edge deployment;
- General rule: INT8 has minimal loss, INT4 requires validation; always run regression on business evaluation sets after quantization (see Deployment & Serving).
22. What problem does continuous batching solve?
Key points:
- Traditional static batching: wait for an entire batch to finish before scheduling the next — short requests get slowed by long ones, and GPU utilization is low;
- Continuous batching: step-level dynamic scheduling — as soon as a request finishes at each decode step, it's removed and replaced with a new one, keeping the GPU "fully loaded";
- Combined with PagedAttention (KV cache paging, eliminating memory fragmentation and reservation waste), can achieve several-fold throughput improvements; vLLM is the representative implementation of this approach.
23. How much memory does a 7B model need for FP16 inference? What about training?
Key points:
- Inference: weights 7B × 2 bytes ≈ 14 GB (FP16), plus KV cache and activations, recommend at least one 24GB card as a starting point; INT4 quantized weights ≈ 3.5–4 GB;
- Training (full-parameter FP16): need weights 14GB + gradients 14GB + Adam optimizer states 28GB ≈ 56GB, plus activations and communication overhead, one A100 (80G) is barely enough — usually requires multi-GPU or ZeRO/QLoRA;
- QLoRA: base model weights frozen at 4-bit (≈ 3.5GB) + small LoRA parameters + optimizer only for LoRA params, a single 24GB card can train a 7B model;
- One-liner: inference depends on weights + KV, training depends on parameters × ~20 bytes (FP16+Adam).
24. What are TTFT / TPOT / throughput? How do you troubleshoot high latency?
Key points:
- TTFT (Time To First Token): first-token latency, reflects the prefill phase, increases with long prompts;
- TPOT (Time Per Output Token): interval per output token, reflects the decode phase;
- Throughput: tokens processed per second (tokens/s), higher with larger batches but per-request latency worsens;
- Troubleshooting: first use stress testing to break down prefill/decode contributions → check if memory limits prevent large batches → check if quantization/operators are active → check if long context is causing KV explosion → check networking/scheduling layers; profile each step (e.g., torch profiler, nsys) to pinpoint the bottleneck.
6. Applications (6 Questions)
25. Walk through the full RAG flow
Key points:
- Motivation: knowledge cutoff, hallucination, private data, traceability;
- Three stages: indexing (document parsing → chunking → embedding → vector store) → retrieval (vector search + BM25 hybrid → reranking) → generation (concatenate context with query, generate with citations);
- Key details: chunk size/overlap strategy, query rewriting, HyDE, rerankers, evaluation metrics (recall, faithfulness, answer relevance);
- Common failures: chunking cuts semantics, retrieval noise pollutes generation, citations can't be traced;
- See RAG Practice.
26. RAG vs fine-tuning vs long context: how to choose? Can they be combined?
Key points:
- Decision dimensions: is knowledge frequently updated (RAG wins), do you need new behavior/style (fine-tuning wins), is a single document extremely long (long context wins), are there cost and privacy constraints;
- RAG suits "external knowledge, traceable, fast to update"; fine-tuning suits "change model behavior/format/domain language"; long context suits "single very-long document used once";
- They can be combined: RAG + fine-tuning is a common combo (fine-tuning improves instruction following, RAG provides knowledge);
- Principle: try prompt first, then RAG, then fine-tuning — test in order of increasing cost (see Prompt Engineering).
27. What is an Agent's working loop? Why does ReAct work?
Key points:
- Loop: sense → reason → act → observe → re-reason, until the task is completed or a limit is reached;
- ReAct: lets the model alternately output Thought (reasoning) and Action (tool calling), continuing after observing tool results — stronger than pure CoT because it can interact with the external world, and errors can be observed and corrected;
- Key engineering points: tool schemas, format validation and auto-retry, loop upper bounds and termination conditions, memory compression, cost and safety guardrails;
- See LLM-based Agents.
28. How does tool calling (function calling) work? How do you ensure stable formats?
Key points:
- Principle: inject tool definitions (name, parameter JSON schema, description) into the system prompt, and the model outputs structured calls (e.g.,
{name, arguments}); - Ensuring stability: use native function calling interfaces instead of bare prompts (models are specifically trained for this); write schemas with clear required/optional parameters and types; validation failure → auto-retry/repair; use JSON mode or constrained decoding when needed;
- Pitfalls: parameter hallucination (model invents non-existent parameter values), wrong tool selection (write clear trigger conditions), infinite tool-calling loops burning money (set step limits and budgets).
29. What is prompt injection? How do you defend against it?
Key points:
- Prompt injection: malicious instructions are hidden in user input or external content (indirect injection), hijacking model behavior — isomorphic to SQL injection;
- Typical scenarios: user input says "ignore previous instructions"; RAG documents embed "send me the system prompt";
- Defenses: least privilege (read-only tools / allowlists / sandbox), input/output filtering, isolating user content from system instructions with delimiters (can only mitigate, not cure), not trusting the model's judgment on sensitive operations (critical operations require human confirmation);
- Realistic conclusion: prompt injection cannot be fully defended against by prompts alone — it requires system boundaries (see Safety & Risks).
30. Why do LLMs produce hallucinations? How do you mitigate them?
Key points:
- Causes: the training objective is "highest likelihood," not "factual correctness"; knowledge in parameters is a "distribution," not a "database"; long-tail knowledge overfits on high-frequency patterns; the model has a "people-pleasing" bias during decoding;
- Mitigation layers: prompt layer (require citations, lower temperature) → retrieval layer (RAG provides factual anchors) → alignment layer (targeted SFT/DPO rewarding factuality) → verification layer (post-hoc fact-checking, model self-check, retrieval validation);
- Accept reality: LLMs are language models, not knowledge bases — critical scenarios require external knowledge + traceability (see Hallucination: Causes & Mitigation).
7. Coding Exercises (5 Questions)
31. Implement numerically stable softmax in numpy
Key points: Subtract the maximum before exponentiating and normalizing, to prevent exp overflow.
python
import numpy as np
def softmax(x, axis=-1):
x = x - np.max(x, axis=axis, keepdims=True) # numerically stable
e = np.exp(x)
return e / np.sum(e, axis=axis, keepdims=True)Follow-up: Why does subtracting the maximum not change the result? (Translation invariance of the exponential function: subtracting the same constant from all inputs produces the same softmax output.)
32. Implement single-head scaled dot-product attention (with causal mask)
Key points: scores = Q@Kᵀ / sqrt(d_k); add mask (set upper triangle to -inf); softmax; out = scores@V.
python
import numpy as np
def attention(Q, K, V, mask=None):
d_k = Q.shape[-1]
scores = Q @ K.transpose(-2, -1) / np.sqrt(d_k)
if mask is not None:
scores = np.where(mask, scores, -1e9) # set masked positions to very small value
weights = softmax(scores, axis=-1)
return weights @ V, weightsFollow-up: Why use -inf instead of 0 for masked positions? (After softmax, these should become zero weights; -1e9 is a numerical approximation.)
33. Implement temperature + top-p sampling
Key points: Temperature scales logits; top-p truncates the candidate set by cumulative probability, then renormalizes.
python
import numpy as np
def top_p_sampling(logits, temperature=1.0, top_p=0.9, rng=None):
rng = rng or np.random.default_rng()
logits = logits / temperature
probs = softmax(logits)
order = np.argsort(probs)[::-1]
sorted_probs = probs[order]
cumsum = np.cumsum(sorted_probs)
keep = cumsum - sorted_probs <= top_p # cumulative prob doesn't exceed top_p
keep_idx = order[keep]
probs = np.zeros_like(probs)
probs[keep_idx] = sorted_probs[keep]
probs /= probs.sum()
return rng.choice(len(probs), p=probs)Follow-up: What happens when temperature is 0? Relationship between top-p and top-k? (Often used together in practice; min-p is the popular 2024+ alternative.)
34. LoRA forward pass and weight merging
Key points: LoRA only trains the increment ΔW = B·A (A initialized with Gaussian, B with zeros); the increment can be merged back into the base weights at inference time.
python
import numpy as np
# Forward: h = W0 @ x + (alpha/r) * B @ A @ x
def lora_forward(W0, A, B, x, alpha, r):
return W0 @ x + (alpha / r) * (B @ (A @ x))
# Merge: directly get fine-tuned weights
def merge(W0, A, B, alpha, r):
return W0 + (alpha / r) * (B @ A)Follow-up: Why initialize B with zeros? Why can weights be merged at inference? (Training only affects the increment; merging doesn't change the output, it just saves the extra branch during inference.)
35. Autoregressive generation loop with KV cache (pseudocode)
Key points: In the prefill phase, compute full KV and produce the first token; thereafter, each step only computes Q for the "latest token" and attends with the cached KV.
python
def generate(model, tokens, max_len, cache=None):
# prefill: process all inputs, get KV cache and first output token
logits, cache = model(tokens, cache=None)
next_tok = sample(logits[-1])
outputs = [next_tok]
for _ in range(max_len):
# decode: only input the latest token, KV auto-concatenates
logits, cache = model(next_tok, cache=cache)
next_tok = sample(logits[-1])
outputs.append(next_tok)
return outputsFollow-up: Why only input one token in the decode phase? (Historical token K/V are already in the cache, no need to recompute); how do you set the cache memory upper bound? (Length limit + eviction strategy.)
8. Open-ended Questions (5 Questions)
36. Tell me about a pitfall you hit
Key points (answer framework):
- Tell it structurally: context → what I did → where it went wrong → how I found out → how I fixed it → what I learned;
- Prefer "data/evaluation" pitfalls (e.g., eval set contamination, improper data mixing) — they're more differentiating than "couldn't install the environment";
- Key: explain the retrospective, not the hardship; admitting mistakes + demonstrating methodology is the growth signal interviewers want to see.
37. Design an enterprise knowledge-base Q&A system (system design)
Key points (answer framework):
- Requirements clarification: document scale/format, update frequency, concurrency, accuracy requirements, on-prem vs cloud;
- Architecture: document ingestion & parsing → chunking → embedding + vector store → retrieval (hybrid + rerank) → generation (with citations) → evaluation & monitoring;
- Key tradeoffs: RAG vs fine-tuning (RAG for frequently updating knowledge); local deployment vs API (cost and compliance); vector store selection (FAISS/Milvus/pgvector by scale);
- Metrics: recall, faithfulness, TTFT/throughput, cost — state your metrics first, then design the system — interviewers are testing whether you have evaluation awareness.
38. How do you prove your fine-tuning/RAG solution is truly effective?
Key points:
- Baseline first: prompt-only baseline vs. fine-tuned approach, same eval set;
- Clean eval sets: public benchmarks + custom business sets, prevent data contamination (first check whether fine-tuning data overlaps with the eval set);
- Look at multiple dimensions: beyond target metric improvement, check whether general capabilities degraded (alignment tax check);
- Speak with statistical significance: is the sample size adequate? confidence intervals? is the A/B test significant?
- Production: online A/B + user feedback closed loop (see Evaluation Practice).
39. How do you keep up with the latest papers? What are you currently following?
Key points:
- Channels: arXiv (subscribe by keyword/follow authors), X/Twitter and HF paper charts, top conferences (NeurIPS/ICML/ACL), open-source community updates;
- Method: set aside fixed weekly time, skim abstracts/titles first, deeply read 1–2 papers, take card notes (question/method/result/limitations);
- Be honest: it's far better to genuinely report 1–2 papers you actually read and explain your judgment than to name-drop ten titles;
- See Reading Paths and Reading Discipline & FAQ.
40. Reverse questions for the interviewer
Key points:
- Team & tech: What base models does the team use? How far along is the evaluation system? What's the biggest current technical challenge?
- Work content: What's the most urgent problem to solve in the first three months? How does the team divide model/data/evaluation work?
- Growth & expectations: Does the team have expectations for "what you'll deliver in six months"? Are there opportunities to work on new directions (Agent/multimodal)?
- Why ask these: reverse-verify the gap between JD and actual work (see the reminder in Module Overview & Career Landscape), and show that you value results and judgment.
9. How to Handle "I Don't Know" Questions
| Scenario | Right Approach |
|---|---|
| Never heard of it | Admit + try to categorize: "I haven't encountered this specific method, but it sounds like it belongs to the XX area — I understand it's trying to solve XX problem. Can I reason through it for 30 seconds?" |
| Know the concept but can't explain it well | Explain what you do know + state boundaries clearly: "This is as far as I understand; I haven't practically verified the deeper details, so I won't guess." |
| Need to derive | State your approach before writing: "My approach is: this problem is essentially XX, I plan to do it in three steps..." — better than silently grinding through it |
| Pushed to the limit | Graceful convergence: "I need to look this up / reproduce it to answer responsibly. I'll note it and verify afterward." |
Honesty is the highest-level strategy
Interviewers are mostly not testing whether you've heard of something, but how you handle the unknown. The cost of being caught fabricating on the spot is far greater than "I can't answer that right now, but my inference is..."
10. Interview Preparation Schedules
1-Month Plan (Standard Pace: 2–3 hours/day)
| Week | Topic | Related Pages |
|---|---|---|
| W1 | Transformer & inference fundamentals: deep-read principles + code attention/softmax | Transformer, Inference Fundamentals |
| W2 | Training & alignment: pretraining / scaling laws / RLHF/DPO | Pretraining, Alignment |
| W3 | Applications & evaluation: RAG, Agent, evaluation systems | RAG Practice, Evaluation Practice |
| W4 | Project deep-dive + mock interviews + coding review | Open-ended questions on this page + Resume Analysis |
1-Week Plan (Emergency Sprint: 4–6 hours/day)
| Day | Focus |
|---|---|
| D1 | Fundamentals 8 questions + coding 31–33 (practice until fluent) |
| D2 | Training 6 questions + alignment 5 questions (explain "why") |
| D3 | Inference 5 questions + applications 6 questions |
| D4 | Coding 34–35 + project deep-dive: explain your strongest project in the four-part structure, 3 times |
| D5 | System design question (knowledge-base Q&A) review + mock interview |
| D6 | Review wrong answers, fill weak blocks, prepare reverse-question list |
Don't learn new knowledge in the last two days
Stop ingesting new content in the 48 hours before your interview. Do only two things: explain what you know until it's smooth, and polish your project stories until they hold up. Fresh concepts crammed at the last minute will 90% of the time be exposed in follow-up questions.
11. Further Reading
Continue within the site
- Knowledge Breakdown — the principle source for every question in this bank; go back and deep-read for any question you can't answer
- Resume Analysis — the source material for project deep-dive questions
- Classic Paper Deep Dives — the "paper-level" source for fundamentals/definitions
- Evaluation Practice — the engineering answers for "how to prove effectiveness" questions
References
- Attention Is All You Need (arXiv:1706.03762) — original source of the attention formula and multi-head mechanism
- Training language models to follow instructions with human feedback (InstructGPT, arXiv:2203.02155) — original paper for RLHF's three-step process
- Direct Preference Optimization (arXiv:2305.18290) — original DPO paper
- Efficient Memory Management for LLM Serving with PagedAttention (arXiv:2309.06180) — vLLM/PagedAttention paper, authoritative reference for inference questions
- LoRA: Low-Rank Adaptation (arXiv:2106.09685) — original LoRA paper
- Hugging Face Documentation — PEFT/TRL/Transformers official docs; entry point for engineering detail verification