Skip to content

Anatomy of the AI Stack

At a glance A top-down dissection of an AI application into five layers — application, capability, model, architecture, and infrastructure — reconnected through the dual lenses of data flow and the training, post-training, and inference lifecycle, making this page the site-wide content map where every layer points to pages that go deeper.

Anatomy of the AI Stack ​

A modern AI application is a five-layer technology stack: infrastructure underneath it, the Transformer architecture as its foundation, three model families supplying the intelligence, a capability layer bolting on knowledge and tools, and an application layer orchestrating the experience. A layperson sees ChatGPT's chat box; an engineer sees layer upon layer of abstraction. This page cuts the iceberg open in full, dissecting it layer by layer from the top down, then looking again through two other lenses: data flow and the timeline. By the end of this page you hold the map to the entire site — the core concepts of every layer, along with the matching case studies, papers, and hands-on guides, all in their proper places on one chart.

This page is written for four kinds of readers: beginners building a global mental model, architects drawing up system designs, candidates preparing for interviews, and developers moving from "able to call an API" to "able to understand the system." Whichever door you came in through, spend ten minutes reading this page first — the road after it gets much smoother.

1. The Big Picture: The AI Application Tech Stack ​

Start with the whole. The diagram below splits an "AI application" into five layers, plus one engineering line that runs across all of them:

text
┌────────────────────────── AI Application Technology Stack ──────────────────────────┐
│                                                                                     │
│  Application Layer    Agents / Chat / AI Search / Image & Video / Coding Assistants │
│                                                                                     │
│  Capability Layer     Prompt Engineering / RAG / Vector Databases / Knowledge Graphs│
│                                                                                     │
│  Model Layer          Large Language Models (LLM) / Multimodal / Diffusion Models   │
│                                                                                     │
│  Architecture Layer   Transformer & the Attention Mechanism                         │
│                                                                                     │
│  Infrastructure Layer GPUs / Inference Engines / Quantization / Deployment /        │
│                       Data Pipelines / Vector Indexing                              │
│                                                                                     │
│  ────────────────────────────────────────────────────────────────────────────────   │
│  Engineering (off the live request path; determines the model's "factory quality"): │
│  Pretraining → Fine-Tuning & PEFT → Alignment → Evaluation → Inference Optimization │
└─────────────────────────────────────────────────────────────────────────────────────┘

The five layers describe the live runtime: requests flow downward, answers come back up. The engineering layer is the manufacturing side — it never sits on a request's path, yet it determines how intelligent, how trustworthy, and how costly the deployed model is. In one sentence: the upper layers consume intelligence, the lower layers produce it, and the engineering layer polishes it.

LayerWhat it is, in one sentenceGo deeper
Application layerThe product shapes and task loops users interact with directlyAgents
Capability layerThe "plug-ins" that let a model retrieve, verify, and be constrainedRAG
Model layerThe three families that supply intelligence: text, multimodal, generationLarge Language Models
Architecture layerThe foundation of every modern model: the attention mechanismTransformer
Infrastructure layerCompute, inference, deployment, data pipelinesInference Optimization & Quantization
Engineering layer (manufacturing side)Training, customization, alignment, evaluation, rolloutFine-Tuning & PEFT

We now dissect the stack layer by layer from the bottom up — starting with the most stable foundation and ending with the user-facing applications.

2. Layer-by-Layer, from the Bottom Up ​

1. The Architecture Layer: Transformer and Attention — the foundation of every modern model ​

This layer is the stack's bedrock. Whether you are using a GPT model or a diffusion model, trace the lineage back three generations and you land on the 2017 paper Attention Is All You Need. Transformer's core contribution was making the attention mechanism the skeleton of the entire network: the model can dynamically decide "where to look and how closely," instead of being forced — as an RNN is — to squeeze information through one sequential pipeline.

Key Transformer componentWhat it doesWhy the model can't live without it
Self-attentionComputes association weights between each token and every other token in the sequenceThe core tool for modeling long-range dependencies
Multi-head attentionRuns several attention heads in parallel, each learning its own relationshipsOne attention head alone is not expressive enough
Positional encodingInjects order information into tokens that are otherwise computed in parallelA Transformer is inherently order-agnostic
Residual connections and LayerNormKeep deep-network training stableWithout them, networks dozens of layers deep simply won't train

The computational core of attention fits in one sentence: take the dot-product similarities between the Query and the Key, softmax them into weights, then take a weighted sum over the Values. For the full derivation and the intuition behind it, see Transformer and the Attention Mechanism.

The one-line verdict

The architecture layer is the most stable of the five — and the last one you should ever modify yourself. Practically all recent progress in mainstream models has come from "keep the Transformer fixed; change the data, the scale, and the training recipe." Understanding this foundation deeply matters far more than chasing each month's model news.

2. The Model Layer: how the three model families divide the work ​

Above the architecture layer live the actual "model species." Mainstream models of the 2020s fall into three families, which differ in output modality, in the tasks they own, and in how they are trained:

FamilyCore input → outputTypical use casesRepresentative examplesGo deeper
Large language models (LLM)Text → textChat, writing, reasoning, codeChatGPT, DeepSeek-R1Large Language Models
Multimodal modelsMixed text/image/audio/video → multimodalImage understanding, video generation, voice interactionSora video generation, Whisper speech AIMultimodal Models
Diffusion modelsText/noise → image/videoImage generation, image editingMidjourneyDiffusion Models & Generative AI

The three families are not substitutes; they are a division of labor: the LLM does the thinking, the diffusion model does the drawing, and the multimodal model does the seeing and hearing. Real products are usually hybrid deployments — a multimodal front end translates the user's image into text, hands it to an LLM for planning, then hands the result to a diffusion model to produce the image.

More than three families

Beyond the big three there are vertical species: AI for Science efforts such as AlphaFold and recommendation systems in the era of large models also hold a place on this family tree. Their underlying architectures are still Transformer variants; only the tasks and training objectives differ.

3. The Capability Layer: bolting on an external brain and external rules ​

The model layer answers "can it"; the capability layer answers "is that enough." Left to parameter memory alone, knowledge goes stale, facts get fabricated, and business rules cannot be enforced — so the industry built a ring of "peripherals" around the model, known collectively as the capability layer:

CapabilityProblem it solvesKey componentsConcept pagePractice page
Prompt engineeringActivates capabilities the model already has, through wordingInstructions, few-shot examples, chain-of-thoughtPrompt EngineeringPrompt Playbook
RAG (retrieval-augmented generation)Injects external facts and curbs hallucinationRetriever + generatorRAGBuild a RAG App from Scratch
Vector databasesThe engine behind semantic retrievalEmbeddings + approximate nearest neighborVector Databases & Semantic SearchBuild a RAG App from Scratch
Knowledge graphsInject structured rules and relationshipsEntities, relations, triplesKnowledge Graphs & Knowledge InjectionCommon Pitfalls & Anti-Patterns

The one-line verdict

The capability layer has been the best value-for-effort engineering lever since 2024: "retrieve first, then generate" (RAG) is two orders of magnitude cheaper than retraining a model, and the payoff is immediate. AI search products like Perplexity are built almost entirely on the "LLM + retrieval + citations" combination (see Perplexity and AI Search).

4. The Application Layer: orchestrating capabilities into closed task loops ​

The capability layer supplies the parts; the application layer assembles them into a product. The star here is the AI agent — it does not just answer a question; it orchestrates perception, planning, tool calls, execution, and self-checking into a closed task loop (see Agents). Today's mainstream application shapes fall into roughly three buckets:

Application shapeCapability mixExampleCase study
Conversational assistantPrompts + conversation memoryChatGPTChatGPT and Conversational AI
AI searchPrompts + RAG + citation renderingPerplexityPerplexity and AI Search
Task agentsPrompts + tool calls + planning loopsGeneral-purpose agents such as ManusManus and Agent Applications

There is also an underrated species at this layer — intelligence embedded into existing products: GitHub Copilot turns an LLM into a resident teammate inside your IDE. Application shapes vary endlessly, but the skeleton never changes: pick a model, attach the capability-layer peripherals, then write a layer of orchestration logic. For the complete walkthrough of building a minimal agent by hand, see Build an Agent from Scratch.

5. The Engineering Layer: making models customizable, trustworthy, and runnable ​

The engineering layer never touches the live request path, but it clears three gates before a model "leaves the factory": customizable (a foundation model may not fit your business), trustworthy (no making things up, no causing harm), and runnable (cost and latency held in check).

Engineering taskWhat it solvesCore methodsConcept pagePractice page
Fine-tuning & PEFTAdapts the model to your domain dataLow-rank adaptation such as LoRA and QLoRAFine-Tuning & PEFTFine-Tune Your Own LLM
AlignmentMakes the model match human preferences and valuesRLHF, DPO, safety trainingAlignment: RLHF and DPO, AI Safety & GovernanceDeploy & Optimize LLM Inference
Inference optimizationCuts latency and costQuantization, distillation, KV cache, speculative decodingInference Optimization & QuantizationDeploy & Optimize LLM Inference
EvaluationTells you whether the model actually worksBenchmarks, human evaluation, automated evalsLLM Evaluation & BenchmarksBuild an LLM Eval Suite

The easiest trap to fall into

The four engineering tasks interlock, and evaluate first, fine-tune second is an iron rule. Many teams skip the cost-effective "prompt → RAG → fine-tuning" progression and burn money on fine-tuning right away — only to end up with the same hallucinations in a different outfit. For the right order and the classic failure scenes, see Common Pitfalls & Anti-Patterns.

3. The Data-Flow View: how one request travels through the layers ​

Layers are the static view; flow is the dynamic one. Zoom in on a single "user asks → answer comes back" exchange and you can watch all five layers cooperate:

text
Input ──▶ Prompt Assembly ──▶ Semantic Retrieval ──▶ Model Inference ──▶ Output Rendering ──▶ Evaluation & Feedback
          (capability)        (capability)           (model + arch)      (application)        (engineering)
StepLayers involvedKey mechanismGo deeper
1. Input normalizationApplicationSession management, instruction wrappingChatGPT and Conversational AI
2. Prompt assemblyCapabilityInstructions + few-shot examples + tool definitionsPrompt Engineering
3. Semantic retrievalCapabilityThe user's question is embedded; relevant documents are recalled from the vector databaseVector Databases & Semantic Search
4. Context injectionCapabilityRetrieved results are stitched into the prompt — the essence of RAGRAG
5. Model inferenceModel + architectureAutoregressive decoding, one token at a timeLarge Language Models
6. Output renderingApplicationStreaming output, citation markers, format validationPerplexity and AI Search
7. Evaluation & feedbackEngineeringOffline evals, online monitoring, error feedback loopsLLM Evaluation & Benchmarks

Steps 3 and 4 are optional but strongly recommended — bare question answering without retrieval shows a markedly higher hallucination rate. The complete end-to-end build is the main body of the Build a RAG App from Scratch practice guide.

4. Training, Post-Training, Inference: the three stages of a model's life ​

The data-flow view answers "how does one request travel"; the timeline view answers "how does a model come to be." Every modern large model is born in three stages:

StageWhat happensKey techniques and scaleConcept page
PretrainingLearns the regularities of language from massive text corporaTransformer + data scale + compute scaleTransformer
Post-trainingFine-tunes for tasks and aligns valuesInstruction tuning, LoRA, RLHF, DPOFine-Tuning & PEFT, Alignment
InferenceOptimizes deployment and serves live trafficQuantization, distillation, inference enginesInference Optimization & Quantization

The three stages have a stark asymmetry: pretraining burns money, post-training burns brains, and inference burns machines. Pretraining sets the ceiling of intelligence (parameter and data scale); post-training sets the floor of behavior (usable? safe?); inference sets the unit cost. What makes DeepSeek-R1 remarkable is precisely that it demonstrated "keep the pretraining skeleton untouched and elicit reasoning ability through reinforcement-learning-style post-training" (see DeepSeek-R1 and Reasoning Models). For the full timeline narrative of the three stages, see A Brief History.

5. The Site-Wide Content Map: seven sections × the layer diagram ​

Finally, pin the whole site onto this layer diagram. The site has seven sections, each with its own job; the mapping looks like this:

SectionWhich layers it coversEntry pages
GuideA global view of all five layers plus the engineering layerLearning Paths, Concept Boundaries
ConceptsThe 14 concepts of the five layers plus the engineering layer, expanded one by oneConcept Overview
Case studiesMostly the application layer, with forays into the model layer10 Classic Case Studies
PapersPrimary literature for the architecture and model layersPaper Map, Reading Paths
PracticeHands-on guides for the capability and engineering layersBuild a RAG App from Scratch, Fine-Tune Your Own LLM
CareersThe roles and skills that correspond to each layerJob Landscape, JD Knowledge Breakdown
ResourcesTerm, dataset, and model quick references for every layerGlossary, Model & Leaderboard Cheat Sheet

How to use this page

Further Reading ​

References ​