Appearance
Glossary
This page collects the core terms of the AI hot-concepts field, grouped into six themes: Models & Architecture, Training & Optimization, Application Paradigms, Data & Retrieval, Generative Models, and Evaluation & Safety, plus a quick-reference table of high-frequency interview abbreviations. Each entry pairs a term with a one-sentence definition; pairs of easily confused terms are disentangled in a dedicated section at the end.
How to Use This Page
You don't have to read the terms in order. Start with What Are AI Hot Concepts to build the big picture, then come back here whenever you hit an unfamiliar word. The "Related page" column in each row links to the in-depth discussion of that term on this site; bolded terms are the core concepts of this series.
1. Models & Architecture
This group answers the question of what a large model looks like: from the basic unit to the overall architecture. Hold on to one thread — words are split into tokens, tokens become vectors, vectors "attend" to each other in the Transformer's attention mechanism, and the next token is generated one token at a time.
| Term | Definition | Related page |
|---|---|---|
| Large Language Model (LLM) | A generative model built on the Transformer and pre-trained on massive amounts of text; as scale grows, understanding, reasoning, and conversational abilities emerge | Large Language Models |
| Transformer | A sequence architecture proposed in 2017 that is built entirely on attention and discards recurrence; it is the foundation of GPT, BERT, and virtually every modern LLM | Transformer and Attention |
| Attention | The mechanism that aggregates information by "relevance weights" when processing a sequence; its core formula is softmax(Q·Kᵀ/√d)·V, letting every word "see" its context | Transformer and Attention |
| Token | The basic unit of text a model processes (roughly 0.7 English words or 1 Chinese character); both context windows and API billing are counted in tokens | Large Language Models |
| Embedding | A representation that maps discrete tokens into dense vectors, so semantically similar words sit closer together | Transformer and Attention |
| Mixture of Experts (MoE) | A sparse architecture that splits one large model into multiple "expert" subnetworks and activates only a few of them per token, sustaining far more parameters on the same compute | Large Language Models |
| Decoder-only | An autoregressive architecture that uses only the Transformer decoder with a causal mask (each token can only see what precedes it); adopted by the GPT family and best suited to text generation | Transformer and Attention |
| KV Cache | Caching the K/V vectors of already-generated tokens during inference to avoid recomputing them at every step — the core inference optimization that trades memory for speed | Inference Optimization and Quantization |
| Positional Encoding | The mechanism that supplies token order information to the otherwise order-blind parallel Transformer; the current mainstream is Rotary Position Embedding (RoPE) | Transformer and Attention |
| Context Window | The total number of tokens a model can take in per request, which determines how much it can read in one go; mainstream models now range from 8K to 1M | Large Language Models |
| In-Context Learning (ICL) | The ability to adapt to new tasks from examples and instructions in the prompt alone, without updating weights; it grows stronger with scale and underpins "the prompt is the program" ever since GPT-3 | Large Language Models |
| Emergent Ability | An ability (multi-step reasoning, instruction following, etc.) that appears suddenly once model scale crosses a threshold and is absent in smaller models; it goes hand in hand with scaling laws | Large Language Models |
| Scaling Law | The empirical law that model performance grows as a power law with parameter count, data volume, and compute; the theoretical basis for the "brute force works" route | Large Language Models |
| Native Multimodal | A multimodal model design that handles text, images, and audio in a unified way from pre-training onward, rather than stitching together separate single-modality models after the fact | Multimodal Models |
| Reasoning Model | A model (e.g., o1, DeepSeek-R1) that generates an internal chain of thought before answering and is specifically trained with reinforcement learning — "think it through, then answer" | DeepSeek-R1 and Reasoning Models |
| Reasoning Budget | The amount of thinking (chain-of-thought length / step limit) a reasoning model is allowed to spend; a bigger budget means smarter but slower and costlier answers, and it can be tuned like a sampling parameter | DeepSeek-R1 and Reasoning Models |
| Temperature | The parameter that controls sampling randomness: higher values make output more diverse, lower values more deterministic, and near 0 it approaches greedy decoding | Prompt Engineering |
| Multi-Agent | A system shape in which multiple agents divide the work (orchestration, debate, each carrying its own tools) to complete complex tasks — coordination costs come with the benefits | AI Agents |
| World Model | A paradigm in which a model learns an internal representation of an environment or the world and predicts the next state; a key direction for video generation and embodied AI | Frontier Advances |
| AGI | A hypothetical intelligence that matches or exceeds human level on nearly all cognitive tasks; most hot concepts here are viewed as paths toward it | What Are AI Hot Concepts |
Rule of thumb
Whenever you see a parameter count, first ask "dense or MoE": a MoE model's total parameter count is not its active parameter count — actual compute follows the active parameters. For example, a model with 1.8T total parameters that activates only 37B costs about the same per inference as a 37B model.
2. Training & Optimization
This group answers how a large model is trained: pre-training learns general knowledge → fine-tuning learns skills → alignment adjusts values, plus a two-piece toolkit (LoRA, quantization) for making models smaller and faster.
| Term | Definition | Related page |
|---|---|---|
| Pretraining | The self-supervised learning stage (predicting the next token) on massive unlabeled corpora; it produces the foundation of general knowledge and costs the most | Large Language Models |
| Supervised Fine-Tuning (SFT) | Continuing to train a pre-trained model on instruction–answer or input–output pairs so it learns to follow instructions and produce well-formatted output | Fine-Tuning and PEFT |
| Reinforcement Learning from Human Feedback (RLHF) | First trains a reward model on human preference data, then optimizes the policy with reinforcement learning (PPO); the key alignment technique behind ChatGPT's rise | Alignment: RLHF and DPO |
| Direct Preference Optimization (DPO) | Encodes preference data directly into a loss function and completes alignment in one training pass, eliminating the reward model and PPO's stability headaches; the mainstream choice in the open-source community | Alignment: RLHF and DPO |
| Low-Rank Adaptation (LoRA) | A fine-tuning method that freezes the original weights and trains only low-rank adaptation matrices, cutting trainable parameters by roughly 10,000× so fine-tuning fits on a single GPU | Fine-Tuning and PEFT |
| Quantized LoRA (QLoRA) | Quantizes the base model to 4-bit and then runs LoRA fine-tuning on top, letting consumer GPUs fine-tune 7B–13B models | Fine-Tuning and PEFT |
| Knowledge Distillation | Using a large model's (teacher's) outputs or logits to teach a small model (student), so the student approaches the teacher's performance; a common way to cut cost and boost efficiency | Fine-Tuning and PEFT |
| Quantization | Compressing weights from FP32 down to FP16/INT8/INT4 — halves memory, speeds up inference, and loses little accuracy; the first choice for deployment optimization | Inference Optimization and Quantization |
| Parameter-Efficient Fine-Tuning (PEFT) | An umbrella term for methods that fine-tune by training only a small set of new parameters (LoRA, Adapters, Prefix, etc.); the goal is adapting to downstream tasks at minimal cost | Fine-Tuning and PEFT |
| Instruction Tuning | Fine-tuning on large numbers of instruction–answer examples; essentially a form of SFT that teaches a base model to "take orders" instead of only continuing text | Alignment: RLHF and DPO |
| Post-training | The umbrella term for all training stages after pre-training (SFT, alignment, reasoning enhancement, etc.); it is what turns a "general-knowledge model" into "a genuinely usable product model" | Large Language Models |
| Odds Ratio Preference Optimization (ORPO) | A reward-model-free alignment method proposed in 2024 that merges SFT and preference optimization into a single step, saving both memory and training time | Alignment: RLHF and DPO |
| Identity Preference Optimization (IPO) | An improved variant of DPO that fixes its overfitting regularization with an identity operator, improving the stability and generalization of preference optimization | Alignment: RLHF and DPO |
| Agentic RL | Applying reinforcement learning to agents' planning, tool use, and long-horizon tasks so agents learn by trial and error in an environment; the go-to tool for training reasoning models in 2025 | Frontier Advances |
| Speculative Decoding | An acceleration trick in which a small model "drafts" candidates and the large model verifies several steps at once; inference gets 2–3× faster with identical output | Inference Optimization and Quantization |
Rule of thumb
Use RAG for knowledge questions, fine-tuning for behavior and style — feeding knowledge through fine-tuning is a common mistake: new knowledge is hard to learn and prone to overfitting, while retrieved knowledge never enters the model weights. When picking an alignment method: with little data (a few thousand examples) and a preference for simplicity, go straight to DPO; if you need tight consistency with human preferences and have the budget, step up to RLHF.
3. Application Paradigms
This group answers how to put a large model to work: without training weights, you shape the input, attach external knowledge, and hook up tools, turning the model from a chatbot into a program that gets things done.
| Term | Definition | Related page |
|---|---|---|
| Prompt | The instruction text fed to a model; wording, structure, and examples constrain the output — the cheapest form of "model tuning" there is | Prompt Engineering |
| Few-shot | Giving the prompt a few input–output examples before asking the model to answer, adapting it to new tasks without any weight updates | Prompt Engineering |
| Chain-of-Thought (CoT) | Prompting the model to "reason step by step before answering" and write out its intermediate steps; markedly improves math and logic tasks | Prompt Engineering |
| Retrieval-Augmented Generation (RAG) | A paradigm that retrieves knowledge first and hands it to the model to generate from, addressing three weaknesses at once: stale knowledge, hallucination, and no access to private data | Retrieval-Augmented Generation |
| GraphRAG | A RAG variant that retrieves over a knowledge graph or other graph structure; strong at cross-entity questions that require global aggregation | Knowledge Graphs and Knowledge Injection |
| Agent | A class of LLM applications that can perceive their environment, plan autonomously, call tools, and iterate over multiple steps to finish tasks; the LLM is its "brain" | AI Agents |
| Function Calling | The model emits structured function calls (with arguments) alongside its text; the program executes them and feeds the results back, letting the model "use tools" | AI Agents |
| ReAct (Reasoning + Acting) | An agent design paradigm built on a "think → act → observe → think again" loop, alternating reasoning and action | AI Agents |
| Model Context Protocol (MCP) | Anthropic's open protocol for tool and data access, released in 2024, letting one agent tool ecosystem be reused across models; it became the de facto standard in 2025 | AI Agents |
| Agent Harness | The shell (framework/protocol) that wraps the model–tool–loop runtime, letting an LLM interact safely with the application environment; a key component of agent engineering | AI Agents |
| Agent Skill | Packaging reusable capabilities (prompt + tools + workflow) into standalone modules that agents load on demand — like installing plug-ins for an agent | AI Agents |
| Self-Reflection / Self-Verification / Self-Critique | A family of techniques in which the model reviews and corrects its own output (retracing its reasoning, verifying answers, critiquing drafts); markedly improves reliability on long tasks | AI Agents |
| Self-RAG | A RAG variant in which the model itself decides when to retrieve, what to retrieve, and whether to accept the results — retrieving on demand instead of every time, cheaper and more precise | Retrieval-Augmented Generation |
| Agentic Workflow | Orchestrating multiple LLM steps or multiple agents into a fixed pipeline (retrieve → generate → review → publish); sits between a single call and a fully autonomous agent | AI Agents |
| Graph of Thoughts (GoT) | A prompting paradigm that models reasoning as a graph (with branching, merging, and backtracking); a generalization of Chain-of-Thought (linear) and Tree of Thoughts (tree-shaped) | Prompt Engineering |
| Vibe Coding | A style of programming where you describe intent in natural language, let the AI write the code, and humans only review and fine-tune; a buzzword that took off in 2025 | GitHub Copilot and Code Intelligence |
| Computer Use | The ability of a model to operate the screen, mouse, and keyboard like a human to complete tasks (Claude Computer Use, OpenAI Operator, etc.) | Manus and Agent Applications |
| Copilot | The byword for AI coding assistants (from GitHub Copilot), covering everything from line-level completion and conversational generation to autonomous code changes | GitHub Copilot and Code Intelligence |
Prompt engineering / RAG / fine-tuning: who does what
The three are the LLM practitioner's toolkit, each solving a different problem: prompt engineering changes the input (zero cost, simple tasks); RAG adds knowledge (factual, private, or real-time questions); fine-tuning changes behavior (style, format, domain voice). A real application usually stacks all three: scaffold with prompts, attach RAG when knowledge is missing, fine-tune when behavior is still off.
4. Data & Retrieval
This group answers the engineering details behind RAG and semantic search: how documents get split, how vectors get stored, and how retrieval becomes both fast and accurate.
| Term | Definition | Related page |
|---|---|---|
| Chunking | Splitting long documents into semantically bounded pieces before indexing; chunk size directly affects retrieval quality (too small loses context, too big adds noise) | Vector Databases and Semantic Search |
| Embedding | In the retrieval context, the mapping from text to dense vectors: semantically similar texts end up close together in vector space — the cornerstone of semantic search | Vector Databases and Semantic Search |
| Vector Database | A database purpose-built to store vectors and query them by similarity (Milvus, Qdrant, Chroma, etc.); the storage layer of RAG | Vector Databases and Semantic Search |
| Approximate Nearest Neighbor (ANN) | Retrieval techniques that trade a little accuracy for sub-second queries over millions to billions of vectors; the core algorithm family underneath vector databases | Vector Databases and Semantic Search |
| Hierarchical Navigable Small World (HNSW) | One of the most popular ANN algorithms today; uses a multi-layer graph for "coarse-to-fine" hops — fast, accurate, with tunable memory use | Vector Databases and Semantic Search |
| Rerank | First coarsely recall a few hundred candidates with vectors or keywords, then precisely re-order the top few with a stronger cross-encoder; a two-stage design that lifts retrieval accuracy | Vector Databases and Semantic Search |
| Hybrid Search | Running keyword (BM25) and vector retrieval in parallel and fusing the rankings, balancing exact matching with semantic understanding; more robust in production than pure vector search | Vector Databases and Semantic Search |
| BM25 | The classic scoring function for lexical retrieval, based on term frequency and document length normalization; the "keyword leg" of hybrid RAG retrieval | Vector Databases and Semantic Search |
| Knowledge Graph | A structured knowledge network organized as entity–relation–entity triples, giving RAG a semantic structure it can reason over and aggregate | Knowledge Graphs and Knowledge Injection |
| Cosine Similarity | The cosine of the angle between two vectors; it measures direction rather than magnitude and is the most commonly used similarity metric in vector retrieval | Vector Databases and Semantic Search |
| Neuro-symbolic | A technical route that combines neural networks (for perception and statistics) with symbolic systems (for logic and rules, such as knowledge graphs), taking the best of both | Knowledge Graphs and Knowledge Injection |
| Data Governance | The policies and technical measures governing the quality, compliance, privacy, and copyright of training and usage data; an enterprise necessity in the era of large models | AI Safety and Governance |
Rule of thumb
A retrieval system needs two stages: coarse recall + fine ranking. ANN only narrows candidates from tens of millions down to a few hundred; a reranker then decides the final order — a RAG that relies on raw vector similarity alone usually isn't accurate enough. For smaller datasets (under a million vectors), just use a library like FAISS; no need for a heavyweight vector database.
5. Generative Models
This group answers how images, video, and audio are generated: diffusion models add noise first and then learn to denoise, steered by text conditioning — the de facto standard for text-to-image and text-to-video today.
| Term | Definition | Related page |
|---|---|---|
| Diffusion Model | A family of generative models that gradually adds noise to data until it becomes pure noise, then learns to denoise step by step to recover the data distribution; currently rules text-to-image and text-to-video | Diffusion Models and Generative AI |
| Denoising Diffusion Probabilistic Model (DDPM) | The 2020 foundational work that cast the diffusion process as a trainable denoising network; the common starting point for later diffusion models | Diffusion Models and Generative AI |
| Latent Diffusion Model (LDM) | Runs diffusion in a low-dimensional latent space rather than pixel space, sharply cutting compute; the technical foundation of Stable Diffusion (2022) | Diffusion Models and Generative AI |
| ControlNet | A 2023 work that adds extra conditions (line art, skeletons, depth maps) to diffusion models for precise control over generation structure, making "aim exactly where you point" possible | Diffusion Models and Generative AI |
| Text-to-Image | The generation task of taking a text prompt and producing an image; representative products include Midjourney, Stable Diffusion, and DALL·E | Midjourney and Image Generation |
| Text-to-Video | The generation task of producing coherent video clips from text or images; the landmark product Sora was released in 2024 | Sora and Video Generation |
| Contrastive Language-Image Pre-training (CLIP) | The image–text alignment model proposed by OpenAI in 2021, mapping images and text into a single vector space; the basis of the "text condition" in text-to-image | Diffusion Models and Generative AI |
| Variational Autoencoder (VAE) | A generative model that learns to compress and reconstruct; LDM uses it to squeeze pixel images into latent space. Proposed in 2013 | Diffusion Models and Generative AI |
| Classifier-Free Guidance (CFG) | A sampling trick that amplifies the text constraint using the difference between conditional and unconditional predictions, making generated results stick closer to the prompt | Diffusion Models and Generative AI |
| Multimodal Model | A model that handles text, images, audio, and video together, merging "seeing," "hearing," and "generating" into one model | Multimodal Models |
| Embodied AI | The paradigm of giving AI a "body" (robots) so it can perceive, interact, and learn in the physical world; seen as another path to AGI | Multimodal Models |
| On-device LLM | Models compressed through quantization to run locally on phones, PCs, and edge devices — offline, private, low latency; the deployment form beyond the cloud | Inference Optimization and Quantization |
| AI Scientist | Agent systems that automate research (read the literature → propose hypotheses → run experiments → write papers); a frontier direction of AI for Science | Frontier Advances |
Rule of thumb
Track the evolution of diffusion models along two lines: compute (pixel space → latent space) and control (text → precise ControlNet conditions → video). The gap in generation quality has already narrowed; the competition now centers on controllability, consistency (stable characters across frames), and efficiency.
6. Evaluation & Safety
This group answers whether a model is actually good and actually safe: benchmarks enable apples-to-apples comparison, while red teams and defenses hold the line against attacks and hallucination.
| Term | Definition | Related page |
|---|---|---|
| Benchmark | A standardized test with a fixed question set and scoring rules for measuring model capability; the common yardstick for comparing models side by side | LLM Evaluation and Benchmarks |
| Massive Multitask Language Understanding (MMLU) | A general-knowledge benchmark spanning 57 subjects and about 14,000 multiple-choice questions; after its 2021 release it became a required exam for mainstream models | LLM Evaluation and Benchmarks |
| Hallucination | When a model confidently fabricates facts or logic that don't exist; the root cause is that it learned "plausible sequences," not a "database of facts" | AI Safety and Governance |
| Red Teaming | Adversarial testing in which testers deliberately attack a model (inducing jailbreaks, digging out biases, probing boundaries) to expose its weaknesses | AI Safety and Governance |
| Jailbreak | An attack that bypasses a model's safety alignment through carefully crafted prompts (role-play, fictional scenarios, mixed languages) | AI Safety and Governance |
| Prompt Injection | An attack in which malicious instructions are hidden inside content the model will read (web pages, documents, tool outputs), tricking it away from its intended task | AI Safety and Governance |
| Watermark | A mark embedded in AI-generated content that is hard for human eyes or ears to detect but machine-verifiable, used to trace whether content came from a model | AI Safety and Governance |
| Perplexity | A measure of how surprised a language model is by data (lower means more confident); a common internal metric during pre-training | LLM Evaluation and Benchmarks |
| Grade School Math 8K (GSM8K) | A benchmark of 8.5K grade-school and middle-school math word problems, testing a model's arithmetic and multi-step reasoning | LLM Evaluation and Benchmarks |
| HumanEval | An OpenAI benchmark of 164 programming problems that measures code generation ability with the pass@k metric | LLM Evaluation and Benchmarks |
Three cautions about benchmarks
First, leaderboard gaming: vendors train on test-set-adjacent data, which inflates scores, so check whether the model saw these questions during training. Second, saturation: mainstream models are already near 90% on MMLU, so a single benchmark's discriminating power is fading — look at combined leaderboards. Third, disconnection from reality: a high benchmark score doesn't mean the model works well in your business; ultimately, defer to evaluation on your own task. For the full methodology, see Building an LLM Evaluation Suite.
Appendix: High-Frequency Interview Abbreviations
Abbreviations fly around in interviews and everyday reading. This table collects the most frequent ones — memorize the full name first, then the meaning — because plenty of canned interview questions are really just testing whether you can expand the acronym.
| Abbreviation | Full name | In one sentence |
|---|---|---|
| AGI | Artificial General Intelligence | Hypothetical intelligence that learns and reasons across domains like a human |
| ASI | Artificial Superintelligence | Hypothetical intelligence that far surpasses top humans in nearly every field |
| GPT | Generative Pre-trained Transformer | Generative pre-trained Transformer; the name of OpenAI's model series |
| BERT | Bidirectional Encoder Representations from Transformers | Pre-trained bidirectional encoder representations; a 2018 NLP milestone |
| RAG | Retrieval-Augmented Generation | Retrieve first, then generate — the knowledge-injection paradigm |
| RLHF | Reinforcement Learning from Human Feedback | Reinforcement learning from human preferences; the mainstream alignment technique |
| DPO | Direct Preference Optimization | Reward-model-free simplified alignment method |
| SFT | Supervised Fine-Tuning | Continued training of a model on labeled data |
| PEFT | Parameter-Efficient Fine-Tuning | Adapt by training only a small set of parameters |
| LoRA | Low-Rank Adaptation | Freeze the original weights and train only low-rank matrices |
| QLoRA | Quantized LoRA | 4-bit quantization + LoRA; fine-tunable on consumer GPUs |
| CoT | Chain-of-Thought | Have the model reason step by step before answering |
| MoE | Mixture of Experts | Sparsely activated giant-model architecture |
| MCP | Model Context Protocol | The open standard for agent tool connectivity |
| KV | Key-Value Cache | The caching mechanism behind inference speedups |
| ANN | Approximate Nearest Neighbor | The core technology of large-scale vector retrieval |
| HNSW | Hierarchical Navigable Small World | One of the mainstream ANN algorithms |
| MMLU | Massive Multitask Language Understanding | Multitask language understanding benchmark (57 subjects) |
| CLIP | Contrastive Language-Image Pre-training | Contrastive image–text pre-training model; the basis of text-to-image conditioning |
| RoPE | Rotary Position Embedding | The position encoding scheme used by mainstream LLMs |
| GQA | Grouped Query Attention | Attention variant that reduces inference memory |
| TTS | Text-to-Speech | Making machines speak |
| ASR | Automatic Speech Recognition | Making machines understand speech |
| SOTA | State of the Art | The current best result on a given task |
Easily Confused Terms
Pretraining / fine-tuning / alignment
Pretraining learns "general language knowledge" (massive text, self-supervised, most expensive); fine-tuning (SFT) learns "task skills" (instruction data, cheap, targeted); alignment (RLHF/DPO) adjusts "values and preferences" (making output better match human expectations and be safer). The three form a pipeline: pretraining → SFT → alignment.
LoRA / QLoRA / full fine-tuning
Full fine-tuning updates all parameters (highest ceiling, but expensive); LoRA freezes the original weights and trains only low-rank adaptation matrices (roughly 10,000× fewer trainable parameters); QLoRA additionally quantizes the base model to 4-bit on top of LoRA (memory down another ~3/4). Small teams should start with QLoRA by default and step up to LoRA or full fine-tuning only if results fall short.
RAG / GraphRAG / fine-tuning (three ways to inject knowledge)
RAG retrieves "text snippets" and suits factual Q&A; GraphRAG retrieves "graph structure" and suits cross-entity aggregation questions (e.g., "what are the most relevant themes across the whole dataset"); fine-tuning should not be used to stuff in knowledge at all. Rule of thumb: knowledge you can retrieve should not be fine-tuned in; relationships that require reasoning belong in a graph.
Jailbreak / prompt injection / hallucination
All three mean "the model said what it shouldn't have," but they differ in nature: jailbreak is a user deliberately crafting prompts to bypass safety alignment; prompt injection is malicious instructions passively triggered from data the model reads; hallucination is the model fabricating facts on its own, with no attack involved. The first two are security contests; the last is a capability limitation.
Embedding (retrieval) vs Embedding (inside the model)
The same word, two contexts: inside the model it is a lookup table from vocabulary entries to dense vectors; in a retrieval system it is the product of an encoder that turns text into semantic vectors. The key insight is the same — similar meanings end up close in vector space.
Further Reading
- What Are AI Hot Concepts — the starting point and big picture for every concept on this site
- Concept Boundaries: AI vs ML vs DL vs GenAI — how the umbrella concepts relate to LLM terminology
- A Brief History — the technical timeline behind the terms
- Anatomy of the Architecture — where each term sits in a real LLM system
- Learning Paths: Three Routes — build your knowledge map with this glossary
- Curated Resources — deep-dive material for every term in this table
- Models and Leaderboards at a Glance — a cross-reference of models and benchmark scores
- Paper Reading: Start Here — entry points to the original papers behind the terms
References
- Vaswani et al., Attention Is All You Need, arXiv:1706.03762, 2017 — the original Transformer paper
- Brown et al., Language Models are Few-Shot Learners (GPT-3), arXiv:2005.14165, 2020
- Ouyang et al., Training language models to follow instructions with human feedback (InstructGPT), arXiv:2203.02155, 2022 — the representative RLHF implementation
- Hu et al., LoRA: Low-Rank Adaptation of Large Language Models, arXiv:2106.09685, 2021
- Dettmers et al., QLoRA: Efficient Finetuning of Quantized LLMs, arXiv:2305.14314, 2023
- Rafailov et al., Direct Preference Optimization, arXiv:2305.18290, 2023
- Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, arXiv:2005.11401, 2020
- Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models, arXiv:2210.03629, 2022
- Ho et al., Denoising Diffusion Probabilistic Models, arXiv:2006.11239, 2020
- Rombach et al., High-Resolution Image Synthesis with Latent Diffusion Models (LDM), arXiv:2112.10752, 2022
- Malkov & Yashunin, Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs, arXiv:1603.09320, 2016
- Hendrycks et al., Measuring Massive Multitask Language Understanding (MMLU), arXiv:2009.03300, 2021
- Model Context Protocol official documentation — modelcontextprotocol.io