Skip to content

LARGE LANGUAGE MODEL HANDBOOK

LLM Handbook

A systematic guide from next-token prediction to artificial general intelligence — language modeling · Transformer architecture · Pre-training & alignment · Evaluation · RAG & Agents · Engineering & deployment · Career paths

If you think of massive text corpora as raw ore, an LLM is the entire pipeline that refines knowledge and patterns from language — data engineering, pre-training, alignment, evaluation & iteration, inference, deployment, and application building. Learn more →

Which path should you start with?

50+ pages isn't a library you read cover to cover — it's a map you assemble as you go

✨ Editor's Picks

If you only read ten pages, start here

All Content

Browse by section, or search directly

🧠 Context Window & Long Context

The context window determines how far a model can "see" at once. This article clarifies the composition of context windows, the essence of positional encoding extrapolation (RoPE/ALiBi/YaRN/NTK-aware), the engineering challenges of O(n²) attention and KV Cache (FlashAttention), long-context training mix and test-time extension, and uses needle-in-a-haystack and other evaluations to map the real boundaries of long context and trade-offs with RAG.

context-windowlong-contextpositional-encoding-extrapolationflashattentionneedle-in-a-haystack

🧠 Fine-Tuning: SFT and Parameter-Efficient Fine-Tuning

Fine-tuning continues training on a pre-trained foundation with supervised data, transforming "can generate" into "can do things by instruction." This article systematically covers SFT data and objectives, compares full-parameter fine-tuning with LoRA/QLoRA/Adapter/P-Tuning, explains LoRA low-rank decomposition principles, and addresses critical engineering issues like catastrophic forgetting and data quality.

fine-tuningsftloraparameter-efficient-fine-tuningcatastrophic-forgetting

🧠 Language Modeling: The Next-Token Prediction Paradigm

Language modeling is the foundation of all large language models: treating 'predict the next token' as a proxy objective to estimate the probability distribution of text sequences. This article clarifies the inheritance chain from N-gram to neural language models to pretraining paradigms, conditional probability decomposition, perplexity and cross-entropy, and 'why predicting the next token learns knowledge.'

language-modelingnext-token-predictionperplexityn-gramcross-entropy

🧠 Pretraining: Data and Objectives

PrPretraining is Phase 1 of an LLM: next-token prediction on trillions of tokens of corpus, writing "language and world knowledge" into parameters. This article covers training objectives and loss, the full data pipeline (collection/cleaning/deduplication/filtering/mixing), data quality equals model quality, training dynamics like learning rate and batch size, and sets up scaling laws and post-training.

pretrainingtraining-datadata-engineeringlearning-ratenext-token-prediction

🧠 Prompt Engineering

Prompting is the programming interface for interacting with large language models. This article covers the constituent elements of prompts (role/instruction/context/examples/output format), zero-shot and few-shot, chain-of-thought and self-consistency, structured output, gives a systematic prompt design methodology and safety boundaries, and answers "when is prompting enough, and when to fine-tune."

prompt-engineeringchain-of-thoughtfew-shotstructured-outputprompt-injection
✨ Recommended

🧠 Transformer Architecture Deep Dive

Transformer is the backbone of all modern LLMs. Starting from Attention Is All You Need, this article systematically breaks down self-attention QKV computation, scaled dot-product, multi-head attention, positional encoding (absolute/relative/RoPE/ALiBi), residual connections and LayerNorm (Pre-Norm), FFN (GELU/SwiGLU), causal masking and KV Cache, along with complexity analysis and a comparison of three architectural forms.

transformerself-attentionpositional-encodingqkvmulti-head-attention
✨ Recommended

🔍 RAG: Retrieval-Augmented Generation

RARAG connects external knowledge to large models via "retrieve first, generate after," solving three major problems: knowledge cutoff, hallucination, and private data. This article breaks down RAG motivation, the index/retrieve/generate three-stage flow, Naive/Advanced/Modular evolution, retrievers and evaluation, and selection comparison with fine-tuning and long context.

RAGRetrieval-Augmented GenerationVector RetrievalKnowledge BaseHallucination