A Brief History
From information theory in 1948 to multimodal and agents in 2025, nearly eighty years of large language model evolution: N-gram, neural language models, Transformers, GPT/BERT, scaling laws, RLHF, and ChatGPT's three paradigm shifts.
LARGE LANGUAGE MODEL HANDBOOK
A systematic guide from next-token prediction to artificial general intelligence — language modeling · Transformer architecture · Pre-training & alignment · Evaluation · RAG & Agents · Engineering & deployment · Career paths
50+ pages isn't a library you read cover to cover — it's a map you assemble as you go
Preparing for interviews within 1–3 months? Go straight to JD gap analysis, resume rewrites, and interview question banks—every page targets LLM roles.
View path →Have the time to build a solid foundation? Progress through the material week by week, from primer to practice, with clear milestones each week.
View path →Already building in production? Skip the deep read—jump from problem to page using the index. Treat it like a reference manual.
View path →If you only read ten pages, start here
Large Language Models (LLMs) are neural network language models pre-trained on massive text corpora, with "predict the next word" as their fundamental task, typically having billions of parameters or more. Their capabilities emerge from thecombination of scale, data, and alignment.
Getting StartedTransformer is the backbone of all modern LLMs. Starting from Attention Is All You Need, this article systematically breaks down self-attention QKV computation, scaled dot-product, multi-head attention, positional encoding (absolute/relative/RoPE/ALiBi), residual connections and LayerNorm (Pre-Norm), FFN (GELU/SwiGLU), causal masking and KV Cache, along with complexity analysis and a comparison of three architectural forms.
Core KnowledgeChatGPT productized RLHF alignment technology into a consumer conversational assistant, serving as the breakout inflection point for large models. This article breaks down its background, dialogue system tech stack, reasons for going mainstream, competitor timelines, dialogue evaluation methods, and product evolution.
Case StudiesMeta's Llama series opened the "open-source window" for large models, catalyzing a complete open-source ecosystem chain for fine-tuning, quantization, and inference deployment. This article breaks down Llama 1/2/3 evolution, the open-source ecosystem chain, Chinese open-source models (Qwen/DeepSeek/Baichuan/ChatGLM/Yi), and the open-source vs. closed-source debate.
Case StudiesLLM-based Agents upgrade large models from 'answering questions' to 'completing tasks': brain + tools + memory + planning loop. This article breaks down ReAct patterns, function calling and MCP, short-term/long-term memory, multi-agent collaboration, representative cases, and risk boundaries.
Case StudiesMixture of Experts (MoE) breaks the strong coupling between parameter count and compute via "sparse activation," making trillions of total parameters feasible. This article breaks down MoE routing/load-balancing mechanisms, traces milestones from GShard → Switch → GLaM → Mixtral → DeepSeek-V3, and compares activated parameters and deployment challenges.
Case StudiesMultimodal LLMs enable LLMs to simultaneously "see, hear, and speak." This article compares modular vs. native multimodal approaches, traces the evolution of CLIP, Flamingo, LLaVA, Qwen-VL, GPT-4V/4o, Gemini, and Sora, and provides multimodal evaluation guidance and boundary explanations.
Case StudiesRARAG connects external knowledge to large models via "retrieve first, generate after," solving three major problems: knowledge cutoff, hallucination, and private data. This article breaks down RAG motivation, the index/retrieve/generate three-stage flow, Naive/Advanced/Modular evolution, retrievers and evaluation, and selection comparison with fine-tuning and long context.
Case StudiesThe GPT series defines the "generative pretraining + scale + alignment" trajectory for large models. This article breaks down GPT-1 through GPT-4o by generation, covering parameters, data, key innovations, and historical significance, plus lessons and an overview table for each milestone.
Case StudiesBrowse by section, or search directly
From information theory in 1948 to multimodal and agents in 2025, nearly eighty years of large language model evolution: N-gram, neural language models, Transformers, GPT/BERT, scaling laws, RLHF, and ChatGPT's three paradigm shifts.
This handbook offers three learning paths — Career Sprint, Systematic Deep Dive, and Desk Reference — each tailored to different audiences and time budgets, with weekly schedules, page checklists, and verification criteria. It serves as the master index map for the entire site.
undary clarification between large language models and NLP, deep learning, statistical language models, foundation models, agents, RAG systems, and AGI: who is a subset of whom, who is an application form of whom, and which "general intelligence" expectations actually fall short.
A complete lifecycle map of large model systems: data engineering → pretraining → post-training → evaluation → deployment & inference → applications → feedback loop, dissecting inputs, outputs, and key questions at each stage, and mapping to pages across the site.
Large Language Models (LLMs) are neural network language models pre-trained on massive text corpora, with "predict the next word" as their fundamental task, typically having billions of parameters or more. Their capabilities emerge from thecombination of scale, data, and alignment.
Alignment is the post-training stage that makes models "both smart and reliable." This article breaks down InstructGPT's SFT→Reward Model→PPO three steps, the KL penalty mechanism, explains why DPO doesn't need a reward model, and compares the trade-offs of RLHF/DPO/RLAIF/Constitutional AI and other methods, concluding with alignment tax and weak-to-strong generalization.
The context window determines how far a model can "see" at once. This article clarifies the composition of context windows, the essence of positional encoding extrapolation (RoPE/ALiBi/YaRN/NTK-aware), the engineering challenges of O(n²) attention and KV Cache (FlashAttention), long-context training mix and test-time extension, and uses needle-in-a-haystack and other evaluations to map the real boundaries of long context and trade-offs with RAG.
Evaluation is the only way to answer 'whether the model really works.' This article explains why LLM evaluation is hard, three categories of evaluation systems (intrinsic metrics/task benchmarks/human & model judges), capability-layered evaluation, benchmark contamination issues and how to read leaderboards correctly, and how offline and online evaluation form a closed loop.
Fine-tuning continues training on a pre-trained foundation with supervised data, transforming "can generate" into "can do things by instruction." This article systematically covers SFT data and objectives, compares full-parameter fine-tuning with LoRA/QLoRA/Adapter/P-Tuning, explains LoRA low-rank decomposition principles, and addresses critical engineering issues like catastrophic forgetting and data quality.
Hallucination is the phenomenon where LLMs "confidently talk nonsense" — the core trust issue for generative AI. This article systematically covers hallucination taxonomy (factuality/faithfulness, knowledge/reasoning), five categories of causes, a layered mitigation system from data to evaluation, RAG's boundary, and the "AI lying vs hallucination" intent distinction.
The inference stage determines how the model "speaks." This article breaks down the autoregressive generation loop, the prefill/decode two phases of KV Cache, compares greedy decoding, temperature sampling, top-k, top-p, min-p, and beam search decoding strategies, and gives practical advice on temperature/top_p/penalties for different scenarios.
Language modeling is the foundation of all large language models: treating 'predict the next token' as a proxy objective to estimate the probability distribution of text sequences. This article clarifies the inheritance chain from N-gram to neural language models to pretraining paradigms, conditional probability decomposition, perplexity and cross-entropy, and 'why predicting the next token learns knowledge.'
MoMoE (Mixture-of-Experts) uses "router + a set of expert FFNs" to achieve activation sparsity: large total parameters but only a small subset activated per token, trading less compute for more capacity. This article covers the motivation, routing and load balancing, milestones from GShard to DeepSeek-V3, expert parallel communication, and the memory and bandwidth challenges of inference deployment.
PrPretraining is Phase 1 of an LLM: next-token prediction on trillions of tokens of corpus, writing "language and world knowledge" into parameters. This article covers training objectives and loss, the full data pipeline (collection/cleaning/deduplication/filtering/mixing), data quality equals model quality, training dynamics like learning rate and batch size, and sets up scaling laws and post-training.
Prompting is the programming interface for interacting with large language models. This article covers the constituent elements of prompts (role/instruction/context/examples/output format), zero-shot and few-shot, chain-of-thought and self-consistency, structured output, gives a systematic prompt design methodology and safety boundaries, and answers "when is prompting enough, and when to fine-tune."
LLM safety is the engineering proposition of "with greater capability comes greater responsibility." This article covers safety alignment objectives, seven risk categories, the attack surface of jailbreak and prompt injection, the red team and defense system, the open-source vs closed-source safety debate, and a map of AI safety research (alignment, interpretability, regulation).
Scaling laws reveal that "loss decreases as a power law with parameter, data, and compute," turning "stacking scale" into a predictable engineering decision. This article covers Kaplan 2020's three power laws, Chinchilla 2022's 1:20 compute-optimal ratio, the emergent abilities debate, inference-time compute as the "fourth dimension," and practical guidance.
Tokenization chops continuous text into the smallest unit a model processes — tokens — and defines the true "input language" for language modeling. This article systematically compares character/word/BPE/WordPiece/Unigram/SentencePiece approaches, walks through the BPE algorithm step by step, and explains vocabulary size, special tokens, multilingual/code/number effects, and token-to-cost conversions.
Transformer is the backbone of all modern LLMs. Starting from Attention Is All You Need, this article systematically breaks down self-attention QKV computation, scaled dot-product, multi-head attention, positional encoding (absolute/relative/RoPE/ALiBi), residual connections and LayerNorm (Pre-Norm), FFN (GELU/SwiGLU), causal masking and KV Cache, along with complexity analysis and a comparison of three architectural forms.
BERT redefined the pretraining paradigm of NLP with "masked language modeling + bidirectional encoding." This article breaks down BERT's MLM/NSP mechanisms, its divergence from GPT, the RoBERTa/ALBERT/DistilBERT family evolution, T5's text-to-text framework, and BERT's legacy in the LLM era.
ChatGPT productized RLHF alignment technology into a consumer conversational assistant, serving as the breakout inflection point for large models. This article breaks down its background, dialogue system tech stack, reasons for going mainstream, competitor timelines, dialogue evaluation methods, and product evolution.
Meta's Llama series opened the "open-source window" for large models, catalyzing a complete open-source ecosystem chain for fine-tuning, quantization, and inference deployment. This article breaks down Llama 1/2/3 evolution, the open-source ecosystem chain, Chinese open-source models (Qwen/DeepSeek/Baichuan/ChatGLM/Yi), and the open-source vs. closed-source debate.
LLM-based Agents upgrade large models from 'answering questions' to 'completing tasks': brain + tools + memory + planning loop. This article breaks down ReAct patterns, function calling and MCP, short-term/long-term memory, multi-agent collaboration, representative cases, and risk boundaries.
Mixture of Experts (MoE) breaks the strong coupling between parameter count and compute via "sparse activation," making trillions of total parameters feasible. This article breaks down MoE routing/load-balancing mechanisms, traces milestones from GShard → Switch → GLaM → Mixtral → DeepSeek-V3, and compares activated parameters and deployment challenges.
Multimodal LLMs enable LLMs to simultaneously "see, hear, and speak." This article compares modular vs. native multimodal approaches, traces the evolution of CLIP, Flamingo, LLaVA, Qwen-VL, GPT-4V/4o, Gemini, and Sora, and provides multimodal evaluation guidance and boundary explanations.
RARAG connects external knowledge to large models via "retrieve first, generate after," solving three major problems: knowledge cutoff, hallucination, and private data. This article breaks down RAG motivation, the index/retrieve/generate three-stage flow, Naive/Advanced/Modular evolution, retrievers and evaluation, and selection comparison with fine-tuning and long context.
The GPT series defines the "generative pretraining + scale + alignment" trajectory for large models. This article breaks down GPT-1 through GPT-4o by generation, covering parameters, data, key innovations, and historical significance, plus lessons and an overview table for each milestone.
In-depth readings of 11 classic papers that shaped the LLM landscape: Attention Is All You Need, GPT-1/2, BERT, Scaling Laws, Chinchilla, InstructGPT, LoRA, FlashAttention, CoT, RAG — each with background problem, core method, key experimental numbers, limitations, and extensions.
A 2023–2025 LLM frontier trends overview: inference-time scaling (o1), long context and million-token windows, MoE at massive scale, native multimodal, Agents and tool use, Mamba state space models, quantization and efficient inference, open-source catching up to closed-source, and new evaluation paradigms — each with evolution timeline, representative papers/models, and a one-line takeaway.
Why LLM practitioners must read papers. This page clarifies the three layers of value in paper reading, outlines three paths (2-hour intro, engineering implementation, and research), introduces the six-page paper section map, lists common pitfalls and time schedules, and shares one crucial reading recommendation.
A map of key LLM papers over decades, arranged by timeline and topic: pre-Transformer era, architecture revolution, scale era, alignment era, applications and systems, and the frontier. Each paper gets one row with year, one-line contribution, and significance.
A complete methodology for reading LLM papers: the three-pass reading method with checklists for each pass, troubleshooting for when you can't understand, criteria for judging paper quality (including LLM-specific pitfalls), card note method, whether blogs alone suffice, tools and techniques for following arXiv, reproduction advice, and troubleshooting.
curated essential reading list from thousands of yearly arXiv papers: ~12 core papers, three paths (2-hour intro / engineering / research) with specific paper checklists and reading order, a "which sections to read" guide for each paper, and a quick-reference table for selecting papers by goal scenario.
llowing the nanoGPT approach, walk through the complete pipeline of "data → tokenization → model → training → sampling" using minimal runnable PyTorch code: hand-train a small GPT that can produce Shakespeare-like text, and get an extension checklist from 1M to 1B parameters with four acceptance criteria.
Ten costly, recurring pitfalls in LLM engineering — leaderboard obsession, data contamination, evaluation overfitting, treating hallucination as a bug, endlessly bloating prompts, blind fine-tuning, ignoring costs, context-window stuffing, neglecting safety, and using LLMs as databases — each with symptoms, root causes, and correct approaches.
From memory estimation formulas (weights + KV cache + activations), quantization (GPTQ/AWQ/GGUF), continuous batching, to TTFT/TPOT metrics, GPU selection, and a vLLM deployment example — covering the complete pipeline and decision framework for "taking a model from your laptop to production."
Build a layered evaluation system combining 'public benchmarks + custom golden sets + LLM-as-a-judge + regression testing + online evaluation': with batch evaluation code, judge implementation and bias control, and cost management — turning 'is the model good?' into quantifiable, regressible engineering metrics.
Walk through the complete fine-tuning pipeline with LoRA — data preparation, base model selection, LoRA configuration, training monitoring, merging and export, evaluation comparison — including QLoRA VRAM estimation, common tool comparison, failure troubleshooting, and criteria for "fine-tuning vs prompting."
A six-layer breakdown of the LLM toolchain — model libraries → training → inference → orchestration → vector DBs → evaluation — with scenario-driven selection decision tables and the "don't let frameworks hijack you" principles plus three self-build criteria.
om prompt template design patterns to structured output, from systematic prompt-tuning methodology to version management and regression testing, to the "prompt vs. fine-tuning vs. RAG" decision tree: turning prompt engineering from "black magic" into reproducible, measurable engineering capability.
Build a production-ready RAG system with an end-to-end pipeline of "document parsing → chunking → embedding → vector DB → hybrid search → reranking → generation → evaluation," and turn retrieval quality into a measurable, optimizable engineering discipline using chunk experiments, HyDE, failure mode tables, and evaluation metrics.
Complete collection of high-frequency LLM role interview questions: 40 questions across fundamentals & Transformer, training, alignment, inference, applications, coding exercises, and open-ended questions, each with reference answer key points. Includes guidance on answering 'I don't know' and two-tier interview prep schedules (1 week / 1 month).
JD templates and keyword radar organized by role type: responsibilities and common requirements across eight categories including Algorithm Engineer (LLM), NLP, Training, Inference Optimization, AI Applications, Agent, Evaluation, and Data Engineering. Includes a generic JD skeleton and a master list of high-frequency skill keywords.
Maps every high-frequency JD requirement (Transformer/attention, pretraining, RLHF/DPO, LoRA, evaluation, RAG, Agent, inference optimization, engineering skills) to knowledge points and site pages. Each point includes "how interviewers will ask about it," plus a five-level self-assessment table and study-plan generator.
A comprehensive guide to the job market for large language models — roles, core skills, and salary ranges for seven key positions including LLM Algorithm Engineer, NLP Engineer, Training/Inference Optimization, AI Applications, Agent, Evaluation, and Data Engineering. Includes a skill radar chart, decision framework, and a map of all five pages in the career module.
The golden structure for LLM resumes (project context → technical challenge → solution → quantified results) and evidence-chain writing; how to write five project types: fine-tuning, RAG, Agent, evaluation, deployment; what you have done vs how deep you can explain; common resume mistakes and a self-assessment checklist.
A high-quality resource map for learning and engineering with large models: introductory courses, official documentation, classic visual guides, open-source projects, model repositories, and community leaderboards, with links, ratings, and recommended combinations by learning objective.
A comprehensive reference of data across the entire LLM pipeline: pretraining corpora, post-training instruction data, and mainstream evaluation benchmarks — including scale, purpose, how to access, and licensing considerations, from Common Crawl to MT-Bench and C-Eval.
A quick-reference guide to core LLM terminology: covering model paradigms, architecture mechanisms, training and alignment, inference and deployment, applications, and evaluation with 50+ terms, each with a one-line explanation and cross-references.
quick-reference guide to mainstream large models: spec, licensing, and highlights comparison across GPT, Claude, Gemini, Llama, Qwen, DeepSeek, Mistral, Gemma, GLM, Phi, and more — including a "how to choose a model" decision table and selection case studies.