Skip to content

Glossary

At a glance A quick-reference guide to core LLM terminology: covering model paradigms, architecture mechanisms, training and alignment, inference and deployment, applications, and evaluation with 50+ terms, each with a one-line explanation and cross-references.

Glossary ​

This page collects the core terminology of the large language model field, organized into eight themed groups: fundamentals and paradigms, architecture and mechanisms, tokenization and representations, training and alignment, inference and sampling, inference optimization and deployment, applications, and evaluation and safety. Each entry provides the English name + a one-sentence explanation and, where applicable, a cross-reference to a related page in this guide. A "commonly confused terms" section follows at the end.

How to use this glossary

There's no need to read terms in order. Start with "Fundamentals and Paradigms" to build your vocabulary backbone, then look up specific terms as needed. Bolded links point to in-depth discussions of that term; if you encounter an acronym, read the full name in parentheses first.

I. Fundamentals and Paradigms ​

Large Language Model (LLM) A neural network language model pretrained on massive text corpora with "next-token prediction" as its core task, typically with billions of parameters or more. Its capabilities emerge from the combination of scale, data, and alignment. See What is a Large Language Model.

Language Model A model that learns the probability distribution P(w₁…wₙ) over text sequences, using "next-token prediction" as its proxy objective; from N-gram models to neural language models to LLMs, the lineage is continuous. See Language Modeling: The Next-Token Prediction Paradigm.

Next-Token Prediction An objective that decomposes sequence probability into a product of per-token conditional probabilities, maximizing the likelihood of the true next token during training. The same mechanism is shared between pretraining and inference-time generation. See Language Modeling.

Pretraining The phase of training on massive unlabeled corpora to learn foundational language abilities. Self-supervised, label-free, requiring only "next-token prediction." See Pretraining: Data and Objectives.

Posttraining An umbrella term for all stages after pretraining: supervised fine-tuning (SFT), preference alignment (RLHF/DPO), etc. It transforms a model that "can continue text" into an assistant that "follows instructions." See Alignment.

Fine-tuning Continuing training on top of pretrained weights to adapt a model to a specific task, style, or domain. Includes full-parameter fine-tuning and parameter-efficient fine-tuning (PEFT). See Fine-tuning: SFT and PEFT.

Alignment The process of making model outputs consistent with human intent (helpfulness, honesty, safety). The dominant methods are RLHF and DPO. See Alignment: RLHF and DPO.

In-Context Learning (ICL) The ability to complete tasks on the fly by providing instructions and examples in a prompt, without updating model weights. This is the core mechanism behind "few-shot" learning. See Prompt Engineering.

Zero-shot / Few-shot Task completion without examples (zero-shot) or with a small number of examples (few-shot). GPT-3 was the first to systematically demonstrate few-shot capability in large models. See Prompt Engineering.

Scaling Laws An empirical law showing that model loss decreases as a power law with respect to parameter count, data volume, and compute budget. Chinchilla further established the compute-optimal ratio of "approximately 1:20 for parameters to data tokens." See Scaling Laws.

Emergent Abilities Capabilities that small models lack but appear suddenly once a scale threshold is crossed (arithmetic, code, chain-of-thought, etc.). Whether these are "true" emergent properties is debated and may be tied to evaluation metrics. See Scaling Laws.

Hallucination When a model generates content that seems plausible but is factually incorrect. Models learn "sequences that look like human language" rather than "fact databases." Mitigation approaches include RAG, citation, and human review. See Hallucination: Causes and Mitigation.

II. Architecture and Mechanisms ​

Transformer A fully attention-based sequence architecture proposed in 2017 that eliminates recurrence. It is highly parallelizable and models long-range dependencies well, serving as the foundation for BERT, GPT, and all LLMs. See Transformer Architecture Explained.

Attention A mechanism that aggregates sequence information by "relevance weights": for a query q, it computes softmax over the dot product of q with each key k (scaled by √d), then takes a weighted sum over values v. See Transformer Architecture Explained.

Self-Attention Attention where queries, keys, and values all come from the same sequence, enabling each token to attend to other tokens in its context. Multi-head attention captures relationships in different subspaces in parallel. See Transformer.

QKV The three vector roles in attention computation: query ("what am I looking for"), key ("what can I match"), and value ("what to provide when matched"). See Transformer.

Positional Encoding A technique to supplement the order-agnostic parallel attention with token position information: sinusoidal, learnable vectors, relative positional encodings, etc. RoPE is the current mainstream approach. See Transformer.

RoPE (Rotary Positional Encoding) A positional encoding that encodes position information as a rotation matrix applied to Q and K. It has good extrapolation properties for long text and is standard in mainstream models like Llama, Qwen, and DeepSeek. See Context and Long Context.

KV Cache (Key-Value Cache) An inference optimization that caches key-value pairs of historical tokens during autoregressive decoding, avoiding redundant attention computations. It enables "one attention pass per generated token" but is also the primary bottleneck for memory and long-context inference. See Inference Fundamentals: Autoregression and Sampling.

MoE (Mixture of Experts) A sparse architecture that replaces the FFN with multiple "experts" plus a router. Each token activates only a subset of experts, significantly reducing computation while keeping the total parameter count the same. See MoE Sparse Expert Models.

Context Window The maximum number of input + output tokens a model can process in one pass. From GPT-2's 1,024 to Gemini's 1 million tokens, this has been one of the fastest-growing specifications from 2023–2025. See Context and Long Context.

III. Tokenization and Representations ​

Tokenizer The component that splits text into a sequence of tokens (and performs the inverse operation). Mainstream algorithms include BPE, WordPiece, Unigram, and SentencePiece. See Tokenization and Vocabulary.

Token The basic unit of text processing for a model, approximately 0.7 English words or 1 Chinese character. Billing and context windows are measured in tokens. See Tokenization and Vocabulary.

BPE (Byte-Pair Encoding) A subword tokenization algorithm that starts from characters/bytes and iteratively merges the most frequent adjacent pair. The standard tokenization method for the GPT series, with vocabulary sizes ranging from a few thousand to about 100K. See Tokenization and Vocabulary.

Embedding A representation layer that maps discrete tokens (or any objects) into dense vectors. Objects with similar semantics have similar vector representations. This is the foundation of vector search and "vector databases." See Tokenization and Vocabulary.

Vocabulary The complete set of tokens a tokenizer can produce, determining the model's smallest language unit and vocabulary overhead. For reference: GPT-2 has ~50K, Llama 3.2 has ~32K, and GPT-4 has ~100K. See Tokenization and Vocabulary.

IV. Training and Alignment ​

SFT (Supervised Fine-tuning) Supervised learning on "instruction–response" pairs to teach the model to follow instructions and match target output formats. This is the first step of the RLHF pipeline and can also be used independently. See Fine-tuning.

RLHF (Reinforcement Learning from Human Feedback) A three-step workflow pioneered by InstructGPT: SFT → train a reward model on human preferences → optimize the policy with PPO reinforcement learning, aligning model behavior to human preferences. See Alignment.

Reward Model (RM) A scoring model trained on human preference data that provides reward signals for the policy's outputs during RLHF. Its scoring quality determines the upper bound of alignment effectiveness. See Alignment.

PPO (Proximal Policy Optimization) The reinforcement learning algorithm commonly used in RLHF. It clips policy update magnitudes to ensure training stability and pairs with KL penalties to prevent the model from straying too far from the SFT starting point. See Alignment.

DPO (Direct Preference Optimization) Directly fine-tunes the model on "good/bad response" preference pairs, combining RLHF's reward modeling and reinforcement learning steps into a single classification-style loss. No reward model needed; it is the mainstream approach in the open-source community. See Alignment.

PEFT (Parameter-Efficient Fine-Tuning) A family of fine-tuning methods that train only a small number of added/extra parameters while freezing the original weights: LoRA, QLoRA, Adapters, Prompt Tuning, etc. These significantly reduce memory and storage costs. See Fine-tuning.

LoRA (Low-Rank Adaptation) A PEFT method that constrains weight updates to a low-rank decomposition W = W₀ + BA, training only two small matrices. It can run on a single GPU and has become the de facto standard for open-source fine-tuning. See Fine-tuning.

QLoRA A variant that applies 4-bit quantization to the base model before training LoRA, bringing the memory requirement for fine-tuning a 65B model down to about 48 GB, allowing users with limited GPUs to fine-tune large models. See Fine-tuning.

Catastrophic Forgetting The phenomenon where a model forgets its original capabilities and knowledge when fine-tuning on new tasks or knowledge. Mitigation strategies include data mixing, lower learning rates, regularization, and parameterized constraints like LoRA. See Fine-tuning.

V. Inference and Sampling ​

Autoregressive Generation A generation method that predicts one token at a time, appending each new token back to the input before predicting the next. This is the core mechanism for the GPT series and the primary source of inference latency. See Inference Fundamentals.

Greedy Decoding A deterministic decoding strategy that picks the highest-probability token at every step. Results are stable but tend to be repetitive and bland, making it suitable for fact-oriented tasks. See Inference Fundamentals.

Temperature A parameter that divides logits before softmax: low temperature produces more deterministic output (facts/code), high temperature produces more diverse output (creative/brainstorming). Typically set between 0.1 and 1.0. See Inference Fundamentals.

top-k A truncation strategy that samples only from the top-k highest-probability tokens, often combined with top-p to control generation quality and diversity. See Inference Fundamentals.

top-p (Nucleus Sampling) A sampling strategy that draws from the smallest candidate set whose cumulative probability reaches p. The candidate count adapts dynamically, making it more aligned with the distribution shape than top-k. This is ChatGPT's default style. See Inference Fundamentals.

Perplexity A metric for how well a language model fits text, equal to the exponential of the average negative log-likelihood per token. Lower values mean the model is "less surprised" by the sequence. See Language Modeling.

Chain-of-Thought (CoT) A technique that prompts the model to "reason step by step before answering," significantly improving performance on math, logic, and other reasoning tasks. Paired with self-consistency sampling, it becomes even stronger. See Prompt Engineering.

Prompt Engineering The practice of designing instructions, roles, examples, and output constraints to guide model outputs. Includes patterns such as zero-shot prompting, few-shot prompting, and CoT. See Prompt Engineering.

System Prompt The fixed top-level instruction in a conversation structure, used to set the role, behavioral guidelines, and global constraints. It typically takes higher priority than user messages. See Prompt Engineering.

VI. Inference Optimization and Deployment ​

Quantization A compression technique that reduces weights/activations from FP16/FP32 to INT8/INT4, roughly halving memory and speeding up inference, with typically manageable precision loss. This is the first choice for deployment optimization. See Deployment and Serving.

GPTQ / AWQ / GGUF Three mainstream quantization approaches: GPTQ and AWQ target GPU serving inference; GGUF is llama.cpp's quantized format, optimized for CPU and edge devices. See Deployment.

vLLM A high-throughput inference engine based on PagedAttention, supporting continuous batching, streaming output, and an OpenAI-compatible API. It is the de facto standard for open-source deployment. See Deployment and Serving.

Continuous Batching A technique that dynamically bundles requests at different completion stages into the same batch, allowing new requests to be inserted at any time. This frees GPU utilization from the "wait for all requests in a batch to finish" paradigm. See Deployment.

PagedAttention A memory management scheme for KV cache that borrows the paging concept from operating systems, eliminating fragmentation and supporting near-zero-waste memory sharing. It is the core innovation of vLLM. See Deployment.

TTFT (Time to First Token) The elapsed time from sending a request to receiving the first output token, primarily determined by the speed of the prefill stage (parallel attention over the input). See Deployment.

TPOT (Time Per Output Token) The average time to generate one output token, determined by the speed of the decode stage. TTFT and TPOT together determine the interactive experience. See Deployment.

Speculative Decoding A technique that uses a small draft model plus a large model verifier to advance multiple tokens per forward pass, speeding up inference losslessly while preserving output quality. See Deployment.

VII. Applications ​

RAG (Retrieval-Augmented Generation) A "retrieve then generate" paradigm: retrieve private, real-time, long-tail knowledge into the context before asking the model to answer. This mitigates hallucination and knowledge cutoff issues and supports traceability. See RAG: Retrieval-Augmented Generation.

Agent A system that uses an LLM as its brain and calls tools through a "plan–act–observe" loop to complete multi-step tasks. Tool calling and memory are the two pillars. See LLM-based Agents.

Function Calling / Tool Calling The capability that enables a model to output structured results specifying "which tool to call and with what parameters." The host executes the tool and feeds the result back to the model for further reasoning. This is a foundational capability for Agents. See LLM-based Agents.

MCP (Model Context Protocol) An open protocol proposed by Anthropic that standardizes the connection between "LLM applications ↔ tools/data sources." It is viewed as a "USB-C for tool calling" interface standard. See LLM-based Agents.

Prompt Injection An attack where adversaries embed malicious instructions in external content the model reads (web pages, documents, tool outputs) to induce the model to deviate from its intended task. This is a security threat unique to LLM applications. See Safety and Risks.

Structured Output / JSON Mode Techniques that constrain model output to valid JSON, tables, and other formats (via few-shot examples, schema constraints, or JSON mode). This is an engineering necessity for production deployments. See Prompt Engineering.

VIII. Evaluation and Safety ​

Benchmark An evaluation suite consisting of a fixed dataset and scoring rules, used for cross-model capability comparison. MMLU, GSM8K, and HumanEval are the three dominant benchmarks today. See Evaluation and Benchmarks.

MMLU A comprehensive knowledge benchmark covering 57 disciplines with ~16,000 multiple-choice questions. It has been the primary leaderboard metric for large models for years, though it is also frequently criticized for data contamination. See Evaluation and Benchmarks.

GSM8K ~8,000 elementary/middle school math word problems that test chained reasoning and multi-step arithmetic, used with CoT prompting. See Evaluation and Benchmarks.

HumanEval 164 Python function completion tasks evaluated with pass@k. It is the de facto benchmark for "code models." See Evaluation and Benchmarks.

pass@k A metric measuring the probability of passing tests in at least one out of k samples. It reflects a model's capability ceiling rather than single-run performance and is the standard metric for code and math evaluation. See Evaluation and Benchmarks.

LLM-as-a-Judge Using a strong model (such as GPT-4) to score model outputs in place of human evaluation. It is cheap and scalable but has biases such as position preference, self-preference, and verbosity preference, requiring calibration. See Evaluation in Practice.

Data Contamination The phenomenon where evaluation samples appear in training data, inflating scores. This is the biggest threat to leaderboard credibility, and defenses include deduplication and phased release. See Evaluation and Benchmarks.

MT-Bench A conversational quality benchmark consisting of 80 multi-turn open-ended questions scored by a strong model. It is an important reference for Chatbot Arena and the LLaMA-series conversational model evaluations. See Evaluation and Benchmarks.

Commonly Confused Terms ​

Pretraining / Posttraining / Fine-tuning

Pretraining trains general language capabilities on massive unlabeled corpora (self-supervised, label-free). Posttraining covers all stages after pretraining, and fine-tuning uses labeled data to adapt to specific tasks. For end users, "fine-tuning" often refers specifically to adaptation techniques like SFT and LoRA, distinguishing them from pretraining (changing the model) and posttraining alignment (changing behavior).

SFT / RLHF / DPO

All three are alignment techniques to "make the model more obedient," but at different levels: SFT uses instruction–response pairs for supervised learning (teaching "how to do it"); RLHF optimizes through a reward model + PPO reinforcement learning loop (teaching "what is good"); DPO is a simplified alternative to RLHF, directly fine-tuning on preference pairs without a reward model, at lower cost.

temperature / top-p / top-k

Temperature changes the steepness of the entire probability distribution (global scaling of logits); top-k hard-truncates to the top-k candidates; top-p dynamically retains candidates based on cumulative probability. All three control "randomness" but operate on different dimensions: temperature controls "shape," while top-k/top-p control the "candidate scope." They are typically used together: set temperature first, then pair with top-p.

RAG / Fine-tuning / Prompt Engineering

The "three-piece set" for LLM deployment solves different problems: Prompt engineering modifies input (zero cost, for simple tasks); RAG adds knowledge (for facts, private data, and real-time issues); Fine-tuning modifies behavior (for style, format, and domain-specific behavior). Use RAG for knowledge problems and fine-tuning for behavior problems — a common mistake is "feeding knowledge through fine-tuning."

Further Reading ​

References ​