Theme
Large Language Models (LLMs)
In one sentence: A Large Language Model (LLM) is an ultra-large-scale Transformer decoder trained on massive text with the goal of "predicting the next token." Through the pretrain-fine-tune-align pipeline, it elevates "statistical co-occurrence" into "general capabilities"—it is the next-generation intelligence foundation that emerged from scaling Transformer architecture.
1. Pretraining: One Sentence, The Entire World
The pretraining objective of an LLM is deceptively simple: given context, predict the next token. Chop internet-scale corpora into token sequences of length 2048/8192, batch them, and feed them to the decoder for cross-entropy loss. Is it really that simple? The keys are scale and data:
- Data composition: Real-world corpora mix web data (Common Crawl), books, code (GitHub), papers, and encyclopedias. Composition determines the capability spectrum: models with higher code proportions have stronger reasoning, while those with more textbooks have cleaner knowledge. Deduplication, cleaning, and quality filtering are prerequisite engineering steps for pretraining—methods in Data and Data Engineering.
- Tokenization: Algorithms like BPE split words into subwords, balancing vocabulary size and sequence efficiency—this determines the "granularity of the world" the model sees.
- Scale and emergence: Small models learn "word collocations," but when parameter count crosses a certain threshold (billions to tens of billions), capabilities like few-shot learning, instruction following, and arithmetic reasoning "emerge." There's fierce debate in academia over "whether emergence is just an artifact of evaluation metrics"—but the existence of scale effects is undisputed.
2. Scaling Laws: The Triangle of Compute, Parameters, and Data
Scaling laws quantify the empirical relationship between "scale → capability":
- Kaplan's law (2020, OpenAI): Model loss decreases as a power law with respect to parameter count, data size, and compute budget—the more you spend, the better the model.
- Chinchilla's law (2022, DeepMind): Model scale and data size should be scaled proportionally. The optimal data size is roughly parameters × 20 (i.e., 70B parameters paired with 1.4T tokens). The previous practice of "only making models bigger without regard for data" was shown to be suboptimal.
- Lesson: With a fixed compute budget, the parameter-to-data ratio isn't arbitrary. "Data is scarcer than models" has been the consensus since 2023, spawning approaches like synthetic data and data reuse.
Engineering Reality
Scaling laws describe "trends," not a "personal shopping list." In practice, compute budgets are limited. What actually matters is: use smaller models + more data + better data quality, or use open-source foundations + fine-tuning/distillation. See MLOps and Model Deployment for specific strategies.
3. SFT: Instruction Fine-Tuning
A pretrained model only knows how to "continue writing," not how to "answer questions." Supervised Fine-Tuning (SFT) turns the model into an "assistant" using (instruction, response) pairs: concatenate the instruction into the context and let the model learn to generate the standard response. Key points:
- Data is the "core ingredient": hundreds of thousands of high-quality instructions can dramatically improve conversational ability. Quality far outweighs quantity.
- Use cross-entropy loss only on the response portion (mask the instruction part), preventing the model from merely learning to "repeat instructions."
- SFT is a mandatory first step for subsequent alignment (RLHF/DPO)—see Representation Learning and Pretraining for the general framework of pretraining/fine-tuning.
4. Alignment: RLHF and DPO
SFT lets the model "speak human language"; alignment lets it "say what humans want to hear." RLHF (Reinforcement Learning from Human Feedback) has three steps:
- Use human-labeled preference pairs (response A vs. B, which is better) to train a reward model (RM).
- Have the policy model generate responses, scored by the reward model.
- Use RL algorithms like PPO to maximize the reward, while adding a KL penalty to prevent drifting too far from the reference model.
DPO (2023) skips the explicit reward model, directly optimizing classification loss on preference pair data—simpler and more stable, making it the default choice in the open-source community. RLHF/DPO is essentially "injecting human value preferences into the model"—it's both the RL's biggest industrial application in Deep Reinforcement Learning Applications and the source of side effects like "over-alignment leads to conservatism/obsequiousness," with ongoing academic debate.
5. In-Context Learning (ICL) and Few-Shot
ICL (In-Context Learning) is a phenomenon discovered by GPT-3 (2020): without changing weights, just place a few examples in the prompt, and the model can learn to perform new tasks—essentially "runtime programming." Few-shot (a few examples) and zero-shot (zero examples, just describe the task) have become the default usage patterns for LLMs. They let us "change behavior without training," at the cost of making prompt engineering a craft—and the effect is heavily influenced by example order and format.
6. RAG: Retrieval-Augmented Generation
An LLM's knowledge is frozen at its training data cutoff, and it hallucinates. RAG (Retrieval-Augmented Generation, 2020) attaches a retriever in front of the generator:
- Chunk documents into pieces, embed them into vectors.
- When a user asks a question, first retrieve the top-k most relevant chunks.
- Concatenate the retrieved results into the prompt, letting the model "answer with reference material in hand."
RAG enables access to private data, real-time knowledge, traceability, and updatability—making it the mainstream approach for enterprise LLM deployment (far cheaper than retraining). Selection criteria versus fine-tuning: use RAG for knowledge-type questions, fine-tuning for style/behavior-type questions. Vector retrieval infrastructure and evaluation methods are in Evaluation in Practice.
7. Inference Optimization: KV Cache, Quantization, Speculative Sampling
Deploying LLMs is costly primarily at inference time. Three key techniques:
- KV cache: Cache historical K and V during decoding to avoid recomputation (see the Transformer Architecture article for principles). Memory grows linearly with sequence length, hence solutions like PagedAttention (vLLM), KV quantization, and context compression.
- Quantization: Compress weights from FP16 to INT8/INT4 (GPTQ, AWQ, GGUF), significantly reducing memory and latency at the cost of minor accuracy loss—note that "quantized models + low-power devices" is the prerequisite for edge deployment.
- Speculative decoding: A small model drafts a batch of tokens first, and the large model verifies them all at once—trading compute for latency, achieving lossless 2–3× speedup.
Industrial inference frameworks (vLLM, TGI, SGLang) have productized these capabilities—see MLOps and Model Deployment.
8. Hallucination and Evaluation: Benchmark Critique
- Hallucination: The model "confidently makes things up." Causes include: the training objective only optimizes for "text-like" output, no factual supervision, and knowledge cutoff. Mitigation relies on RAG, retrieval-based fact-checking, and decoding constraints—but complete elimination remains an open problem.
- Evaluation: MMLU (57 academic disciplines, multiple choice), HELM (Stanford, multi-dimensional standardized evaluation), HumanEval (code) are common benchmarks. Benchmark critique is equally important: questions leak into training sets, multiple choice is easily "gamed," and static leaderboards can't reflect reliability on real tasks. The truly effective approach is task-specific evaluation sets + human blind evaluation. Methodology is in Deep Learning Evaluation and Experimentation.
Look at Leaderboards Rationally
Model scores are "evidence of a lower bound on capability," not "a promise of an upper bound." A single model's performance varies wildly across different domains and question phrasings—evaluation design itself is a research topic. Don't let a single score dictate your technology choices.
9. From LLMs to Agents
The next paradigm leap for LLMs is letting models use tools: combining chain-of-thought (CoT), function calling, multi-step execution, memory, and environmental feedback into an "agent loop." The model is no longer "answering questions" but "achieving goals."
- Tools: search, code interpreters, browsers, databases, external APIs.
- Frameworks: ReAct (reason-act-observe loop), planner + executor architectures.
- Practical constraints: Agent success rates are affected by the multiplicative effect of per-step accuracy; hallucinations get amplified. The current academic consensus is "reliability > cleverness." Evaluation and fallback design (see Evaluation in Practice) matter more than stacking features.
When combined with multimodal capabilities, agents can also operate on images, video, and physical devices—this intersects with Multimodal Models and is a frequent topic in Frontier Progress.
10. Trade-offs
- Foundation vs. fine-tuning vs. RAG: Data-sensitive? Use RAG. Behavioral shaping? Fine-tuning. General capabilities? Rely on the foundation model. Most scenarios call for the combination of "RAG + lightweight fine-tuning."
- Open-source vs. closed-source: Closed-source saves ops overhead but sacrifices data sovereignty and is cost-controlled by the provider. Open-source (LLaMA, Qwen, DeepSeek) offers control but requires self-maintenance and alignment—trade-offs also involve compliance and interpretability and fairness.
- Scale vs. efficiency: Bigger models are generally smarter, but inference cost, latency, and energy consumption all scale up. For budget-constrained scenarios, distillation and MoE (Mixture of Experts, e.g., Mixtral, DeepSeek-V3) offer better "capability/cost" ratios.
- Capability vs. alignment: Over-alignment dulls capability and diversity; under-alignment is dangerous. This tension has no perfect solution—only continuous engineering and governance.
Further Reading
- Transformer Architecture—The physical foundation of LLMs
- Multimodal Models—LLMs' "see, hear, speak" extensions
- Deep Reinforcement Learning Applications—The RL perspective on RLHF/RLVR
- Deep Learning Evaluation and Experimentation—The foundation of evaluation methodology
- Data and Data Engineering—Behind the scenes of pretraining corpora
- MLOps and Model Deployment—Inference optimization and engineering deployment
References
- Brown et al. Language Models are Few-Shot Learners (GPT-3) (NeurIPS 2020)
- Kaplan et al. Scaling Laws for Neural Language Models (2020)
- Hoffmann et al. Training Compute-Optimal Large Language Models (Chinchilla) (2022)
- Ouyang et al. Training language models to follow instructions with human feedback (InstructGPT) (NeurIPS 2022)
- Rafailov et al. Direct Preference Optimization: Your Language Model is Secretly a Reward Model (NeurIPS 2023)
- Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (NeurIPS 2020)
- Liang et al. Holistic Evaluation of Language Models (HELM) (TMLR 2023)
- Yao et al. ReAct: Synergizing Reasoning and Acting in Language Models (ICLR 2023)