Theme
Large Language Models (LLM)
Concept Definition: The Scaling Miracle of Predicting the Next Token
Large Language Models (LLMs) are generative models based on Transformer decoders, pre-trained on ultra-large-scale corpora (architecture see Transformer and NLP). The training objective is astonishingly simple — predict the next token:
Input: "The capital of China is" → Model output: "Beijing" (probability distribution of the next token)But "predicting the next word" under the influence of scale spurs translation, summarization, reasoning, programming, and conversational abilities — this is the scaling law: model capability grows as a power law with parameters, data, and compute, and many abilities emerge suddenly after a certain scale threshold. GPT-3 (175 billion parameters) proved this path in 2020; ChatGPT turned it into a product in 2022.
2. How LLMs Are Made: Three Stages
Stage 1: Pretraining
"Next-word prediction" on trillions of tokens of internet corpora (web pages, books, code, papers).
- Cost: hundreds of millions to billions of dollars in GPU compute (GPT-4 level); only a few major companies can afford it;
- Data: quality > quantity (post-2024 "data curse" — too much low-quality data drags the model down, see Data Engineering);
- Output: a base model — "can speak but isn't usable": can continue text, but won't follow instructions, may output harmful content.
Stage 2: Supervised Fine-Tuning (SFT)
Fine-tune with "instruction-answer" pairs (manually written or distilled) to teach the model to follow instructions:
User: "Explain what machine learning is in one sentence"
Assistant: "Machine learning is the technology of letting computers automatically learn patterns from data..."Stage 3: Alignment (RLHF / DPO)
Make the model's output align with human values and preferences. RLHF (Reinforcement Learning from Human Feedback) in three steps:
① Collect human preference rankings for multiple answers
② Train a reward model: learn "which answer is better"
③ Optimize the language model with PPO RL to maximize rewardDPO (Direct Preference Optimization) is the simplified alternative post-2023: no need to train a separate reward model; fine-tune directly with preference pairs — simpler and more stable, becoming the open-source community's mainstream. The goal of alignment is to make models "helpful, honest, and harmless" — this is the dividing line between ChatGPT and the base model.
3. In-Context Learning and Chain of Thought: Using Without Training
LLM's killer feature is inference-time capability — no weight updates, relying only on input content:
| Technique | Approach | Suitable for |
|---|---|---|
| Zero-shot prompting | Ask directly ("Translate to English: hello") | Simple tasks |
| Few-shot prompting | Give a few examples, then ask | Complex format/style tasks |
| Chain of Thought (CoT) | "Please think step by step before answering" | Math/reasoning tasks, significant improvement |
| Self-Consistency | Multiple sampling + voting | High-reliability reasoning |
| Tool Use | Model outputs structured calls, connects to external tools | Needs real-time/computation/external data |
RAG (Retrieval-Augmented Generation): for facts the LLM doesn't know (real-time news, private documents), retrieve first, then generate:
User question → Vector retrieval (document library) → Insert into context → LLM generates cited answerRAG solves three LLM shortcomings: outdated knowledge, hallucination, and no private data. It's enterprises' first choice for LLM deployment (cheaper than fine-tuning, updatable, traceable). See RAG Practice.
4. Fine-Tuning: Specializing the Model
RAG solves "knowledge"; fine-tuning solves "style/format/domain behavior":
| Method | Principle | Cost | Suitable for |
|---|---|---|---|
| LoRA | Train only low-rank adaptation matrices, freeze original weights | Low (can run on single GPU) | First choice for most fine-tuning scenarios |
| QLoRA | 4-bit quantization + LoRA | Very low (consumer-grade GPU) | Small-team fine-tuning |
| Full-parameter fine-tuning | Train all weights | High (multi-GPU cluster) | Extremely large domain gaps |
| Continued pretraining | Continue training on domain corpora | High | Vertical domains (legal/medical) |
Fine-tuning vs RAG choice: knowledge-type problems → RAG; behavior/format-type → fine-tuning; the two are often combined (RAG gives knowledge + fine-tuning gives style). The most common misconception in fine-tuning is "using fine-tuning to feed knowledge to the model" — that's RAG's job; fine-tuning can't learn new facts.
5. Inference and Deployment: Making Large Models Run
Key engineering points for LLM inference (see MLOps):
- KV Cache: cache attention keys/values for historical tokens, avoid re-computation during generation — several-fold speedup for inference;
- Continuous Batching: dynamically merge multiple requests, GPU utilization from single digits to 90%+;
- Quantization: FP16 → INT8/INT4 (AWQ, GPTQ), halve VRAM, double throughput;
- Speculative Decoding: small model drafts + large model verifies, 2-3x speedup;
- Inference frameworks: vLLM (open-source de facto standard), TensorRT-LLM, SGLang; API services: OpenAI/Anthropic/ByteDance Volcengine/DeepSeek, etc.;
- Distillation: large model teaches small model (GPT-4's answers distilled to a 7B model), cost drops 100x, 80% of effects retained.
6. Model Family Landscape (2025–2026 Perspective)
| Category | Representatives | Characteristics |
|---|---|---|
| Closed-source flagship | GPT-4o/o-series, Claude, Gemini | Strongest capability, requires paid API access, mature ecosystem |
| Open-source catch-up | Llama, Mistral, Qwen, DeepSeek | Open weights, deployable privately, closing the gap with closed-source |
| Reasoning models | o1/o3, DeepSeek-R1, Claude thinking | Chain-of-thought + RL-enhanced reasoning, strong in math/code |
| Multi-modal | GPT-4V, Gemini, Qwen-VL, Claude | Unified understanding of text+image+audio+video |
| Domain small models | Distilled/fine-tuned 1–32B models | Private deployment, low cost, focused on vertical scenarios |
Key trends:
- Rise of reasoning models (from late 2024): train the model to "think long" with RL during training, reasoning ability leaps, but inference-time token cost is large ("slow thinking");
- Open-source catching up: DeepSeek-R1 proved that "algorithmic innovation can make up for compute gaps," narrowing the open-source vs closed-source gap to within 1–2 generations;
- From "chatting" to "working": LLMs as Agent brains (tool use, code execution, autonomous tasks) are the main battleground post-2025 — beyond the model, you need a harness (tools, loops, permissions, memory), which is exactly what the sister site Agent Harness Manual studies.
7. Limitations and Risks of LLMs
- Hallucination: confidently fabricating facts. Mitigation: RAG retrieval + citations, ask the model to "say I don't know," evaluation fallbacks;
- Limited context window: although it grew from 2K to 200K+, "long context ≠ long memory" (lost-in-the-middle problem); ultra-long inputs still need retrieval/summarization/compression;
- Inference cost: one LLM inference can be several orders of magnitude more expensive than traditional ML — use routing (simple problems go to small models) to control costs;
- Safety alignment: jailbreaks, prompt injection, harmful content — need guardrails and evaluation (see Interpretability and Fairness);
- Copyright and compliance: copyright disputes over training data and generated content are still evolving; need compliance assessment before commercial use.
Using LLMs as search engines is the biggest misuse
LLMs are "generators from similar distributions," not "fact databases." For knowledge-type, real-time, and verifiable needs, the correct approach is RAG/tool use + human review, not asking the model naked.
8. A Landing Path for Practitioners
For using LLMs in an enterprise, the recommended gradual path:
① Use API + prompt engineering for PoC (verify business value, within days)
② Add RAG (connect to private knowledge, solve fact-type needs)
③ Tool use / Agent (connect to systems, solve "working" needs)
④ LoRA fine-tuning (lock in style/format, solve behavior-type needs)
⑤ Distillation + quantization + private deployment (control costs, ensure data security)
Evaluate at each step; don't jump straight to the heaviest.Team invariant: every LLM application needs a evaluation set (dozens of representative questions + scoring criteria); otherwise "feels better" can't be judged. Evaluation methods are in Building a Model Evaluation System from Scratch.
Further Reading
- Transformer and NLP — The architectural foundation of LLMs
- Generative Models — Autoregressive generation mechanisms
- Reinforcement Learning — The RL foundation of RLHF
- MLOps and Model Deployment — Inference optimization and monitoring
- Interpretability and Fairness — Alignment and safety
- Classic Paper Deep Dives — GPT/BERT original papers
- Frontier Progress — 2025–2026 research threads
References
- Vaswani et al. Attention Is All You Need (2017)
- Brown et al. Language Models are Few-Shot Learners (GPT-3, 2020)
- Kaplan et al. Scaling Laws for Neural Language Models (2020)
- Hoffmann et al. Training Compute-Optimal LLMs (Chinchilla, 2022)
- Ouyang et al. Training language models to follow instructions with human feedback (RLHF, 2022)
- Rafailov et al. Direct Preference Optimization (DPO, 2023)
- Hu et al. LoRA: Low-Rank Adaptation of Large Language Models (ICLR 2022)
- Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (RAG, 2020)
- Wei et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (NeurIPS 2022)
- DeepSeek-AI. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL (2025)