Skip to content

Glossary

At a glance Core glossary of AI hot concepts: accurate one-sentence definitions of 80+ terms across six categories — model architecture, training & optimization, application paradigms, data & retrieval, generative models, and evaluation & safety — plus a quick-reference table of high-frequency interview abbreviations.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

Glossary ​

This page collects the core terms of the AI hot-concepts field, grouped into six themes: Models & Architecture, Training & Optimization, Application Paradigms, Data & Retrieval, Generative Models, and Evaluation & Safety, plus a quick-reference table of high-frequency interview abbreviations. Each entry pairs a term with a one-sentence definition; pairs of easily confused terms are disentangled in a dedicated section at the end.

How to Use This Page

You don't have to read the terms in order. Start with What Are AI Hot Concepts to build the big picture, then come back here whenever you hit an unfamiliar word. The "Related page" column in each row links to the in-depth discussion of that term on this site; bolded terms are the core concepts of this series.

1. Models & Architecture ​

This group answers the question of what a large model looks like: from the basic unit to the overall architecture. Hold on to one thread — words are split into tokens, tokens become vectors, vectors "attend" to each other in the Transformer's attention mechanism, and the next token is generated one token at a time.

TermDefinitionRelated page
Large Language Model (LLM)A generative model built on the Transformer and pre-trained on massive amounts of text; as scale grows, understanding, reasoning, and conversational abilities emergeLarge Language Models
TransformerA sequence architecture proposed in 2017 that is built entirely on attention and discards recurrence; it is the foundation of GPT, BERT, and virtually every modern LLMTransformer and Attention
AttentionThe mechanism that aggregates information by "relevance weights" when processing a sequence; its core formula is softmax(Q·Kᵀ/√d)·V, letting every word "see" its contextTransformer and Attention
TokenThe basic unit of text a model processes (roughly 0.7 English words or 1 Chinese character); both context windows and API billing are counted in tokensLarge Language Models
EmbeddingA representation that maps discrete tokens into dense vectors, so semantically similar words sit closer togetherTransformer and Attention
Mixture of Experts (MoE)A sparse architecture that splits one large model into multiple "expert" subnetworks and activates only a few of them per token, sustaining far more parameters on the same computeLarge Language Models
Decoder-onlyAn autoregressive architecture that uses only the Transformer decoder with a causal mask (each token can only see what precedes it); adopted by the GPT family and best suited to text generationTransformer and Attention
KV CacheCaching the K/V vectors of already-generated tokens during inference to avoid recomputing them at every step — the core inference optimization that trades memory for speedInference Optimization and Quantization
Positional EncodingThe mechanism that supplies token order information to the otherwise order-blind parallel Transformer; the current mainstream is Rotary Position Embedding (RoPE)Transformer and Attention
Context WindowThe total number of tokens a model can take in per request, which determines how much it can read in one go; mainstream models now range from 8K to 1MLarge Language Models
In-Context Learning (ICL)The ability to adapt to new tasks from examples and instructions in the prompt alone, without updating weights; it grows stronger with scale and underpins "the prompt is the program" ever since GPT-3Large Language Models
Emergent AbilityAn ability (multi-step reasoning, instruction following, etc.) that appears suddenly once model scale crosses a threshold and is absent in smaller models; it goes hand in hand with scaling lawsLarge Language Models
Scaling LawThe empirical law that model performance grows as a power law with parameter count, data volume, and compute; the theoretical basis for the "brute force works" routeLarge Language Models
Native MultimodalA multimodal model design that handles text, images, and audio in a unified way from pre-training onward, rather than stitching together separate single-modality models after the factMultimodal Models
Reasoning ModelA model (e.g., o1, DeepSeek-R1) that generates an internal chain of thought before answering and is specifically trained with reinforcement learning — "think it through, then answer"DeepSeek-R1 and Reasoning Models
Reasoning BudgetThe amount of thinking (chain-of-thought length / step limit) a reasoning model is allowed to spend; a bigger budget means smarter but slower and costlier answers, and it can be tuned like a sampling parameterDeepSeek-R1 and Reasoning Models
TemperatureThe parameter that controls sampling randomness: higher values make output more diverse, lower values more deterministic, and near 0 it approaches greedy decodingPrompt Engineering
Multi-AgentA system shape in which multiple agents divide the work (orchestration, debate, each carrying its own tools) to complete complex tasks — coordination costs come with the benefitsAI Agents
World ModelA paradigm in which a model learns an internal representation of an environment or the world and predicts the next state; a key direction for video generation and embodied AIFrontier Advances
AGIA hypothetical intelligence that matches or exceeds human level on nearly all cognitive tasks; most hot concepts here are viewed as paths toward itWhat Are AI Hot Concepts

Rule of thumb

Whenever you see a parameter count, first ask "dense or MoE": a MoE model's total parameter count is not its active parameter count — actual compute follows the active parameters. For example, a model with 1.8T total parameters that activates only 37B costs about the same per inference as a 37B model.

2. Training & Optimization ​

This group answers how a large model is trained: pre-training learns general knowledge → fine-tuning learns skills → alignment adjusts values, plus a two-piece toolkit (LoRA, quantization) for making models smaller and faster.

TermDefinitionRelated page
PretrainingThe self-supervised learning stage (predicting the next token) on massive unlabeled corpora; it produces the foundation of general knowledge and costs the mostLarge Language Models
Supervised Fine-Tuning (SFT)Continuing to train a pre-trained model on instruction–answer or input–output pairs so it learns to follow instructions and produce well-formatted outputFine-Tuning and PEFT
Reinforcement Learning from Human Feedback (RLHF)First trains a reward model on human preference data, then optimizes the policy with reinforcement learning (PPO); the key alignment technique behind ChatGPT's riseAlignment: RLHF and DPO
Direct Preference Optimization (DPO)Encodes preference data directly into a loss function and completes alignment in one training pass, eliminating the reward model and PPO's stability headaches; the mainstream choice in the open-source communityAlignment: RLHF and DPO
Low-Rank Adaptation (LoRA)A fine-tuning method that freezes the original weights and trains only low-rank adaptation matrices, cutting trainable parameters by roughly 10,000× so fine-tuning fits on a single GPUFine-Tuning and PEFT
Quantized LoRA (QLoRA)Quantizes the base model to 4-bit and then runs LoRA fine-tuning on top, letting consumer GPUs fine-tune 7B–13B modelsFine-Tuning and PEFT
Knowledge DistillationUsing a large model's (teacher's) outputs or logits to teach a small model (student), so the student approaches the teacher's performance; a common way to cut cost and boost efficiencyFine-Tuning and PEFT
QuantizationCompressing weights from FP32 down to FP16/INT8/INT4 — halves memory, speeds up inference, and loses little accuracy; the first choice for deployment optimizationInference Optimization and Quantization
Parameter-Efficient Fine-Tuning (PEFT)An umbrella term for methods that fine-tune by training only a small set of new parameters (LoRA, Adapters, Prefix, etc.); the goal is adapting to downstream tasks at minimal costFine-Tuning and PEFT
Instruction TuningFine-tuning on large numbers of instruction–answer examples; essentially a form of SFT that teaches a base model to "take orders" instead of only continuing textAlignment: RLHF and DPO
Post-trainingThe umbrella term for all training stages after pre-training (SFT, alignment, reasoning enhancement, etc.); it is what turns a "general-knowledge model" into "a genuinely usable product model"Large Language Models
Odds Ratio Preference Optimization (ORPO)A reward-model-free alignment method proposed in 2024 that merges SFT and preference optimization into a single step, saving both memory and training timeAlignment: RLHF and DPO
Identity Preference Optimization (IPO)An improved variant of DPO that fixes its overfitting regularization with an identity operator, improving the stability and generalization of preference optimizationAlignment: RLHF and DPO
Agentic RLApplying reinforcement learning to agents' planning, tool use, and long-horizon tasks so agents learn by trial and error in an environment; the go-to tool for training reasoning models in 2025Frontier Advances
Speculative DecodingAn acceleration trick in which a small model "drafts" candidates and the large model verifies several steps at once; inference gets 2–3× faster with identical outputInference Optimization and Quantization

Rule of thumb

Use RAG for knowledge questions, fine-tuning for behavior and style — feeding knowledge through fine-tuning is a common mistake: new knowledge is hard to learn and prone to overfitting, while retrieved knowledge never enters the model weights. When picking an alignment method: with little data (a few thousand examples) and a preference for simplicity, go straight to DPO; if you need tight consistency with human preferences and have the budget, step up to RLHF.

3. Application Paradigms ​

This group answers how to put a large model to work: without training weights, you shape the input, attach external knowledge, and hook up tools, turning the model from a chatbot into a program that gets things done.

TermDefinitionRelated page
PromptThe instruction text fed to a model; wording, structure, and examples constrain the output — the cheapest form of "model tuning" there isPrompt Engineering
Few-shotGiving the prompt a few input–output examples before asking the model to answer, adapting it to new tasks without any weight updatesPrompt Engineering
Chain-of-Thought (CoT)Prompting the model to "reason step by step before answering" and write out its intermediate steps; markedly improves math and logic tasksPrompt Engineering
Retrieval-Augmented Generation (RAG)A paradigm that retrieves knowledge first and hands it to the model to generate from, addressing three weaknesses at once: stale knowledge, hallucination, and no access to private dataRetrieval-Augmented Generation
GraphRAGA RAG variant that retrieves over a knowledge graph or other graph structure; strong at cross-entity questions that require global aggregationKnowledge Graphs and Knowledge Injection
AgentA class of LLM applications that can perceive their environment, plan autonomously, call tools, and iterate over multiple steps to finish tasks; the LLM is its "brain"AI Agents
Function CallingThe model emits structured function calls (with arguments) alongside its text; the program executes them and feeds the results back, letting the model "use tools"AI Agents
ReAct (Reasoning + Acting)An agent design paradigm built on a "think → act → observe → think again" loop, alternating reasoning and actionAI Agents
Model Context Protocol (MCP)Anthropic's open protocol for tool and data access, released in 2024, letting one agent tool ecosystem be reused across models; it became the de facto standard in 2025AI Agents
Agent HarnessThe shell (framework/protocol) that wraps the model–tool–loop runtime, letting an LLM interact safely with the application environment; a key component of agent engineeringAI Agents
Agent SkillPackaging reusable capabilities (prompt + tools + workflow) into standalone modules that agents load on demand — like installing plug-ins for an agentAI Agents
Self-Reflection / Self-Verification / Self-CritiqueA family of techniques in which the model reviews and corrects its own output (retracing its reasoning, verifying answers, critiquing drafts); markedly improves reliability on long tasksAI Agents
Self-RAGA RAG variant in which the model itself decides when to retrieve, what to retrieve, and whether to accept the results — retrieving on demand instead of every time, cheaper and more preciseRetrieval-Augmented Generation
Agentic WorkflowOrchestrating multiple LLM steps or multiple agents into a fixed pipeline (retrieve → generate → review → publish); sits between a single call and a fully autonomous agentAI Agents
Graph of Thoughts (GoT)A prompting paradigm that models reasoning as a graph (with branching, merging, and backtracking); a generalization of Chain-of-Thought (linear) and Tree of Thoughts (tree-shaped)Prompt Engineering
Vibe CodingA style of programming where you describe intent in natural language, let the AI write the code, and humans only review and fine-tune; a buzzword that took off in 2025GitHub Copilot and Code Intelligence
Computer UseThe ability of a model to operate the screen, mouse, and keyboard like a human to complete tasks (Claude Computer Use, OpenAI Operator, etc.)Manus and Agent Applications
CopilotThe byword for AI coding assistants (from GitHub Copilot), covering everything from line-level completion and conversational generation to autonomous code changesGitHub Copilot and Code Intelligence

Prompt engineering / RAG / fine-tuning: who does what

The three are the LLM practitioner's toolkit, each solving a different problem: prompt engineering changes the input (zero cost, simple tasks); RAG adds knowledge (factual, private, or real-time questions); fine-tuning changes behavior (style, format, domain voice). A real application usually stacks all three: scaffold with prompts, attach RAG when knowledge is missing, fine-tune when behavior is still off.

4. Data & Retrieval ​

This group answers the engineering details behind RAG and semantic search: how documents get split, how vectors get stored, and how retrieval becomes both fast and accurate.

TermDefinitionRelated page
ChunkingSplitting long documents into semantically bounded pieces before indexing; chunk size directly affects retrieval quality (too small loses context, too big adds noise)Vector Databases and Semantic Search
EmbeddingIn the retrieval context, the mapping from text to dense vectors: semantically similar texts end up close together in vector space — the cornerstone of semantic searchVector Databases and Semantic Search
Vector DatabaseA database purpose-built to store vectors and query them by similarity (Milvus, Qdrant, Chroma, etc.); the storage layer of RAGVector Databases and Semantic Search
Approximate Nearest Neighbor (ANN)Retrieval techniques that trade a little accuracy for sub-second queries over millions to billions of vectors; the core algorithm family underneath vector databasesVector Databases and Semantic Search
Hierarchical Navigable Small World (HNSW)One of the most popular ANN algorithms today; uses a multi-layer graph for "coarse-to-fine" hops — fast, accurate, with tunable memory useVector Databases and Semantic Search
RerankFirst coarsely recall a few hundred candidates with vectors or keywords, then precisely re-order the top few with a stronger cross-encoder; a two-stage design that lifts retrieval accuracyVector Databases and Semantic Search
Hybrid SearchRunning keyword (BM25) and vector retrieval in parallel and fusing the rankings, balancing exact matching with semantic understanding; more robust in production than pure vector searchVector Databases and Semantic Search
BM25The classic scoring function for lexical retrieval, based on term frequency and document length normalization; the "keyword leg" of hybrid RAG retrievalVector Databases and Semantic Search
Knowledge GraphA structured knowledge network organized as entity–relation–entity triples, giving RAG a semantic structure it can reason over and aggregateKnowledge Graphs and Knowledge Injection
Cosine SimilarityThe cosine of the angle between two vectors; it measures direction rather than magnitude and is the most commonly used similarity metric in vector retrievalVector Databases and Semantic Search
Neuro-symbolicA technical route that combines neural networks (for perception and statistics) with symbolic systems (for logic and rules, such as knowledge graphs), taking the best of bothKnowledge Graphs and Knowledge Injection
Data GovernanceThe policies and technical measures governing the quality, compliance, privacy, and copyright of training and usage data; an enterprise necessity in the era of large modelsAI Safety and Governance

Rule of thumb

A retrieval system needs two stages: coarse recall + fine ranking. ANN only narrows candidates from tens of millions down to a few hundred; a reranker then decides the final order — a RAG that relies on raw vector similarity alone usually isn't accurate enough. For smaller datasets (under a million vectors), just use a library like FAISS; no need for a heavyweight vector database.

5. Generative Models ​

This group answers how images, video, and audio are generated: diffusion models add noise first and then learn to denoise, steered by text conditioning — the de facto standard for text-to-image and text-to-video today.

TermDefinitionRelated page
Diffusion ModelA family of generative models that gradually adds noise to data until it becomes pure noise, then learns to denoise step by step to recover the data distribution; currently rules text-to-image and text-to-videoDiffusion Models and Generative AI
Denoising Diffusion Probabilistic Model (DDPM)The 2020 foundational work that cast the diffusion process as a trainable denoising network; the common starting point for later diffusion modelsDiffusion Models and Generative AI
Latent Diffusion Model (LDM)Runs diffusion in a low-dimensional latent space rather than pixel space, sharply cutting compute; the technical foundation of Stable Diffusion (2022)Diffusion Models and Generative AI
ControlNetA 2023 work that adds extra conditions (line art, skeletons, depth maps) to diffusion models for precise control over generation structure, making "aim exactly where you point" possibleDiffusion Models and Generative AI
Text-to-ImageThe generation task of taking a text prompt and producing an image; representative products include Midjourney, Stable Diffusion, and DALL·EMidjourney and Image Generation
Text-to-VideoThe generation task of producing coherent video clips from text or images; the landmark product Sora was released in 2024Sora and Video Generation
Contrastive Language-Image Pre-training (CLIP)The image–text alignment model proposed by OpenAI in 2021, mapping images and text into a single vector space; the basis of the "text condition" in text-to-imageDiffusion Models and Generative AI
Variational Autoencoder (VAE)A generative model that learns to compress and reconstruct; LDM uses it to squeeze pixel images into latent space. Proposed in 2013Diffusion Models and Generative AI
Classifier-Free Guidance (CFG)A sampling trick that amplifies the text constraint using the difference between conditional and unconditional predictions, making generated results stick closer to the promptDiffusion Models and Generative AI
Multimodal ModelA model that handles text, images, audio, and video together, merging "seeing," "hearing," and "generating" into one modelMultimodal Models
Embodied AIThe paradigm of giving AI a "body" (robots) so it can perceive, interact, and learn in the physical world; seen as another path to AGIMultimodal Models
On-device LLMModels compressed through quantization to run locally on phones, PCs, and edge devices — offline, private, low latency; the deployment form beyond the cloudInference Optimization and Quantization
AI ScientistAgent systems that automate research (read the literature → propose hypotheses → run experiments → write papers); a frontier direction of AI for ScienceFrontier Advances

Rule of thumb

Track the evolution of diffusion models along two lines: compute (pixel space → latent space) and control (text → precise ControlNet conditions → video). The gap in generation quality has already narrowed; the competition now centers on controllability, consistency (stable characters across frames), and efficiency.

6. Evaluation & Safety ​

This group answers whether a model is actually good and actually safe: benchmarks enable apples-to-apples comparison, while red teams and defenses hold the line against attacks and hallucination.

TermDefinitionRelated page
BenchmarkA standardized test with a fixed question set and scoring rules for measuring model capability; the common yardstick for comparing models side by sideLLM Evaluation and Benchmarks
Massive Multitask Language Understanding (MMLU)A general-knowledge benchmark spanning 57 subjects and about 14,000 multiple-choice questions; after its 2021 release it became a required exam for mainstream modelsLLM Evaluation and Benchmarks
HallucinationWhen a model confidently fabricates facts or logic that don't exist; the root cause is that it learned "plausible sequences," not a "database of facts"AI Safety and Governance
Red TeamingAdversarial testing in which testers deliberately attack a model (inducing jailbreaks, digging out biases, probing boundaries) to expose its weaknessesAI Safety and Governance
JailbreakAn attack that bypasses a model's safety alignment through carefully crafted prompts (role-play, fictional scenarios, mixed languages)AI Safety and Governance
Prompt InjectionAn attack in which malicious instructions are hidden inside content the model will read (web pages, documents, tool outputs), tricking it away from its intended taskAI Safety and Governance
WatermarkA mark embedded in AI-generated content that is hard for human eyes or ears to detect but machine-verifiable, used to trace whether content came from a modelAI Safety and Governance
PerplexityA measure of how surprised a language model is by data (lower means more confident); a common internal metric during pre-trainingLLM Evaluation and Benchmarks
Grade School Math 8K (GSM8K)A benchmark of 8.5K grade-school and middle-school math word problems, testing a model's arithmetic and multi-step reasoningLLM Evaluation and Benchmarks
HumanEvalAn OpenAI benchmark of 164 programming problems that measures code generation ability with the pass@k metricLLM Evaluation and Benchmarks

Three cautions about benchmarks

First, leaderboard gaming: vendors train on test-set-adjacent data, which inflates scores, so check whether the model saw these questions during training. Second, saturation: mainstream models are already near 90% on MMLU, so a single benchmark's discriminating power is fading — look at combined leaderboards. Third, disconnection from reality: a high benchmark score doesn't mean the model works well in your business; ultimately, defer to evaluation on your own task. For the full methodology, see Building an LLM Evaluation Suite.

Appendix: High-Frequency Interview Abbreviations ​

Abbreviations fly around in interviews and everyday reading. This table collects the most frequent ones — memorize the full name first, then the meaning — because plenty of canned interview questions are really just testing whether you can expand the acronym.

AbbreviationFull nameIn one sentence
AGIArtificial General IntelligenceHypothetical intelligence that learns and reasons across domains like a human
ASIArtificial SuperintelligenceHypothetical intelligence that far surpasses top humans in nearly every field
GPTGenerative Pre-trained TransformerGenerative pre-trained Transformer; the name of OpenAI's model series
BERTBidirectional Encoder Representations from TransformersPre-trained bidirectional encoder representations; a 2018 NLP milestone
RAGRetrieval-Augmented GenerationRetrieve first, then generate — the knowledge-injection paradigm
RLHFReinforcement Learning from Human FeedbackReinforcement learning from human preferences; the mainstream alignment technique
DPODirect Preference OptimizationReward-model-free simplified alignment method
SFTSupervised Fine-TuningContinued training of a model on labeled data
PEFTParameter-Efficient Fine-TuningAdapt by training only a small set of parameters
LoRALow-Rank AdaptationFreeze the original weights and train only low-rank matrices
QLoRAQuantized LoRA4-bit quantization + LoRA; fine-tunable on consumer GPUs
CoTChain-of-ThoughtHave the model reason step by step before answering
MoEMixture of ExpertsSparsely activated giant-model architecture
MCPModel Context ProtocolThe open standard for agent tool connectivity
KVKey-Value CacheThe caching mechanism behind inference speedups
ANNApproximate Nearest NeighborThe core technology of large-scale vector retrieval
HNSWHierarchical Navigable Small WorldOne of the mainstream ANN algorithms
MMLUMassive Multitask Language UnderstandingMultitask language understanding benchmark (57 subjects)
CLIPContrastive Language-Image Pre-trainingContrastive image–text pre-training model; the basis of text-to-image conditioning
RoPERotary Position EmbeddingThe position encoding scheme used by mainstream LLMs
GQAGrouped Query AttentionAttention variant that reduces inference memory
TTSText-to-SpeechMaking machines speak
ASRAutomatic Speech RecognitionMaking machines understand speech
SOTAState of the ArtThe current best result on a given task

Easily Confused Terms ​

Pretraining / fine-tuning / alignment

Pretraining learns "general language knowledge" (massive text, self-supervised, most expensive); fine-tuning (SFT) learns "task skills" (instruction data, cheap, targeted); alignment (RLHF/DPO) adjusts "values and preferences" (making output better match human expectations and be safer). The three form a pipeline: pretraining → SFT → alignment.

LoRA / QLoRA / full fine-tuning

Full fine-tuning updates all parameters (highest ceiling, but expensive); LoRA freezes the original weights and trains only low-rank adaptation matrices (roughly 10,000× fewer trainable parameters); QLoRA additionally quantizes the base model to 4-bit on top of LoRA (memory down another ~3/4). Small teams should start with QLoRA by default and step up to LoRA or full fine-tuning only if results fall short.

RAG / GraphRAG / fine-tuning (three ways to inject knowledge)

RAG retrieves "text snippets" and suits factual Q&A; GraphRAG retrieves "graph structure" and suits cross-entity aggregation questions (e.g., "what are the most relevant themes across the whole dataset"); fine-tuning should not be used to stuff in knowledge at all. Rule of thumb: knowledge you can retrieve should not be fine-tuned in; relationships that require reasoning belong in a graph.

Jailbreak / prompt injection / hallucination

All three mean "the model said what it shouldn't have," but they differ in nature: jailbreak is a user deliberately crafting prompts to bypass safety alignment; prompt injection is malicious instructions passively triggered from data the model reads; hallucination is the model fabricating facts on its own, with no attack involved. The first two are security contests; the last is a capability limitation.

Embedding (retrieval) vs Embedding (inside the model)

The same word, two contexts: inside the model it is a lookup table from vocabulary entries to dense vectors; in a retrieval system it is the product of an encoder that turns text into semantic vectors. The key insight is the same — similar meanings end up close in vector space.

Further Reading ​

References ​

  • Vaswani et al., Attention Is All You Need, arXiv:1706.03762, 2017 — the original Transformer paper
  • Brown et al., Language Models are Few-Shot Learners (GPT-3), arXiv:2005.14165, 2020
  • Ouyang et al., Training language models to follow instructions with human feedback (InstructGPT), arXiv:2203.02155, 2022 — the representative RLHF implementation
  • Hu et al., LoRA: Low-Rank Adaptation of Large Language Models, arXiv:2106.09685, 2021
  • Dettmers et al., QLoRA: Efficient Finetuning of Quantized LLMs, arXiv:2305.14314, 2023
  • Rafailov et al., Direct Preference Optimization, arXiv:2305.18290, 2023
  • Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, arXiv:2005.11401, 2020
  • Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models, arXiv:2210.03629, 2022
  • Ho et al., Denoising Diffusion Probabilistic Models, arXiv:2006.11239, 2020
  • Rombach et al., High-Resolution Image Synthesis with Latent Diffusion Models (LDM), arXiv:2112.10752, 2022
  • Malkov & Yashunin, Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs, arXiv:1603.09320, 2016
  • Hendrycks et al., Measuring Massive Multitask Language Understanding (MMLU), arXiv:2009.03300, 2021
  • Model Context Protocol official documentation — modelcontextprotocol.io