Appearance
Anatomy of the AI Stack
A modern AI application is a five-layer technology stack: infrastructure underneath it, the Transformer architecture as its foundation, three model families supplying the intelligence, a capability layer bolting on knowledge and tools, and an application layer orchestrating the experience. A layperson sees ChatGPT's chat box; an engineer sees layer upon layer of abstraction. This page cuts the iceberg open in full, dissecting it layer by layer from the top down, then looking again through two other lenses: data flow and the timeline. By the end of this page you hold the map to the entire site — the core concepts of every layer, along with the matching case studies, papers, and hands-on guides, all in their proper places on one chart.
This page is written for four kinds of readers: beginners building a global mental model, architects drawing up system designs, candidates preparing for interviews, and developers moving from "able to call an API" to "able to understand the system." Whichever door you came in through, spend ten minutes reading this page first — the road after it gets much smoother.
1. The Big Picture: The AI Application Tech Stack
Start with the whole. The diagram below splits an "AI application" into five layers, plus one engineering line that runs across all of them:
text
┌────────────────────────── AI Application Technology Stack ──────────────────────────┐
│ │
│ Application Layer Agents / Chat / AI Search / Image & Video / Coding Assistants │
│ │
│ Capability Layer Prompt Engineering / RAG / Vector Databases / Knowledge Graphs│
│ │
│ Model Layer Large Language Models (LLM) / Multimodal / Diffusion Models │
│ │
│ Architecture Layer Transformer & the Attention Mechanism │
│ │
│ Infrastructure Layer GPUs / Inference Engines / Quantization / Deployment / │
│ Data Pipelines / Vector Indexing │
│ │
│ ──────────────────────────────────────────────────────────────────────────────── │
│ Engineering (off the live request path; determines the model's "factory quality"): │
│ Pretraining → Fine-Tuning & PEFT → Alignment → Evaluation → Inference Optimization │
└─────────────────────────────────────────────────────────────────────────────────────┘The five layers describe the live runtime: requests flow downward, answers come back up. The engineering layer is the manufacturing side — it never sits on a request's path, yet it determines how intelligent, how trustworthy, and how costly the deployed model is. In one sentence: the upper layers consume intelligence, the lower layers produce it, and the engineering layer polishes it.
| Layer | What it is, in one sentence | Go deeper |
|---|---|---|
| Application layer | The product shapes and task loops users interact with directly | Agents |
| Capability layer | The "plug-ins" that let a model retrieve, verify, and be constrained | RAG |
| Model layer | The three families that supply intelligence: text, multimodal, generation | Large Language Models |
| Architecture layer | The foundation of every modern model: the attention mechanism | Transformer |
| Infrastructure layer | Compute, inference, deployment, data pipelines | Inference Optimization & Quantization |
| Engineering layer (manufacturing side) | Training, customization, alignment, evaluation, rollout | Fine-Tuning & PEFT |
We now dissect the stack layer by layer from the bottom up — starting with the most stable foundation and ending with the user-facing applications.
2. Layer-by-Layer, from the Bottom Up
1. The Architecture Layer: Transformer and Attention — the foundation of every modern model
This layer is the stack's bedrock. Whether you are using a GPT model or a diffusion model, trace the lineage back three generations and you land on the 2017 paper Attention Is All You Need. Transformer's core contribution was making the attention mechanism the skeleton of the entire network: the model can dynamically decide "where to look and how closely," instead of being forced — as an RNN is — to squeeze information through one sequential pipeline.
| Key Transformer component | What it does | Why the model can't live without it |
|---|---|---|
| Self-attention | Computes association weights between each token and every other token in the sequence | The core tool for modeling long-range dependencies |
| Multi-head attention | Runs several attention heads in parallel, each learning its own relationships | One attention head alone is not expressive enough |
| Positional encoding | Injects order information into tokens that are otherwise computed in parallel | A Transformer is inherently order-agnostic |
| Residual connections and LayerNorm | Keep deep-network training stable | Without them, networks dozens of layers deep simply won't train |
The computational core of attention fits in one sentence: take the dot-product similarities between the Query and the Key, softmax them into weights, then take a weighted sum over the Values. For the full derivation and the intuition behind it, see Transformer and the Attention Mechanism.
The one-line verdict
The architecture layer is the most stable of the five — and the last one you should ever modify yourself. Practically all recent progress in mainstream models has come from "keep the Transformer fixed; change the data, the scale, and the training recipe." Understanding this foundation deeply matters far more than chasing each month's model news.
2. The Model Layer: how the three model families divide the work
Above the architecture layer live the actual "model species." Mainstream models of the 2020s fall into three families, which differ in output modality, in the tasks they own, and in how they are trained:
| Family | Core input → output | Typical use cases | Representative examples | Go deeper |
|---|---|---|---|---|
| Large language models (LLM) | Text → text | Chat, writing, reasoning, code | ChatGPT, DeepSeek-R1 | Large Language Models |
| Multimodal models | Mixed text/image/audio/video → multimodal | Image understanding, video generation, voice interaction | Sora video generation, Whisper speech AI | Multimodal Models |
| Diffusion models | Text/noise → image/video | Image generation, image editing | Midjourney | Diffusion Models & Generative AI |
The three families are not substitutes; they are a division of labor: the LLM does the thinking, the diffusion model does the drawing, and the multimodal model does the seeing and hearing. Real products are usually hybrid deployments — a multimodal front end translates the user's image into text, hands it to an LLM for planning, then hands the result to a diffusion model to produce the image.
More than three families
Beyond the big three there are vertical species: AI for Science efforts such as AlphaFold and recommendation systems in the era of large models also hold a place on this family tree. Their underlying architectures are still Transformer variants; only the tasks and training objectives differ.
3. The Capability Layer: bolting on an external brain and external rules
The model layer answers "can it"; the capability layer answers "is that enough." Left to parameter memory alone, knowledge goes stale, facts get fabricated, and business rules cannot be enforced — so the industry built a ring of "peripherals" around the model, known collectively as the capability layer:
| Capability | Problem it solves | Key components | Concept page | Practice page |
|---|---|---|---|---|
| Prompt engineering | Activates capabilities the model already has, through wording | Instructions, few-shot examples, chain-of-thought | Prompt Engineering | Prompt Playbook |
| RAG (retrieval-augmented generation) | Injects external facts and curbs hallucination | Retriever + generator | RAG | Build a RAG App from Scratch |
| Vector databases | The engine behind semantic retrieval | Embeddings + approximate nearest neighbor | Vector Databases & Semantic Search | Build a RAG App from Scratch |
| Knowledge graphs | Inject structured rules and relationships | Entities, relations, triples | Knowledge Graphs & Knowledge Injection | Common Pitfalls & Anti-Patterns |
The one-line verdict
The capability layer has been the best value-for-effort engineering lever since 2024: "retrieve first, then generate" (RAG) is two orders of magnitude cheaper than retraining a model, and the payoff is immediate. AI search products like Perplexity are built almost entirely on the "LLM + retrieval + citations" combination (see Perplexity and AI Search).
4. The Application Layer: orchestrating capabilities into closed task loops
The capability layer supplies the parts; the application layer assembles them into a product. The star here is the AI agent — it does not just answer a question; it orchestrates perception, planning, tool calls, execution, and self-checking into a closed task loop (see Agents). Today's mainstream application shapes fall into roughly three buckets:
| Application shape | Capability mix | Example | Case study |
|---|---|---|---|
| Conversational assistant | Prompts + conversation memory | ChatGPT | ChatGPT and Conversational AI |
| AI search | Prompts + RAG + citation rendering | Perplexity | Perplexity and AI Search |
| Task agents | Prompts + tool calls + planning loops | General-purpose agents such as Manus | Manus and Agent Applications |
There is also an underrated species at this layer — intelligence embedded into existing products: GitHub Copilot turns an LLM into a resident teammate inside your IDE. Application shapes vary endlessly, but the skeleton never changes: pick a model, attach the capability-layer peripherals, then write a layer of orchestration logic. For the complete walkthrough of building a minimal agent by hand, see Build an Agent from Scratch.
5. The Engineering Layer: making models customizable, trustworthy, and runnable
The engineering layer never touches the live request path, but it clears three gates before a model "leaves the factory": customizable (a foundation model may not fit your business), trustworthy (no making things up, no causing harm), and runnable (cost and latency held in check).
| Engineering task | What it solves | Core methods | Concept page | Practice page |
|---|---|---|---|---|
| Fine-tuning & PEFT | Adapts the model to your domain data | Low-rank adaptation such as LoRA and QLoRA | Fine-Tuning & PEFT | Fine-Tune Your Own LLM |
| Alignment | Makes the model match human preferences and values | RLHF, DPO, safety training | Alignment: RLHF and DPO, AI Safety & Governance | Deploy & Optimize LLM Inference |
| Inference optimization | Cuts latency and cost | Quantization, distillation, KV cache, speculative decoding | Inference Optimization & Quantization | Deploy & Optimize LLM Inference |
| Evaluation | Tells you whether the model actually works | Benchmarks, human evaluation, automated evals | LLM Evaluation & Benchmarks | Build an LLM Eval Suite |
The easiest trap to fall into
The four engineering tasks interlock, and evaluate first, fine-tune second is an iron rule. Many teams skip the cost-effective "prompt → RAG → fine-tuning" progression and burn money on fine-tuning right away — only to end up with the same hallucinations in a different outfit. For the right order and the classic failure scenes, see Common Pitfalls & Anti-Patterns.
3. The Data-Flow View: how one request travels through the layers
Layers are the static view; flow is the dynamic one. Zoom in on a single "user asks → answer comes back" exchange and you can watch all five layers cooperate:
text
Input ──▶ Prompt Assembly ──▶ Semantic Retrieval ──▶ Model Inference ──▶ Output Rendering ──▶ Evaluation & Feedback
(capability) (capability) (model + arch) (application) (engineering)| Step | Layers involved | Key mechanism | Go deeper |
|---|---|---|---|
| 1. Input normalization | Application | Session management, instruction wrapping | ChatGPT and Conversational AI |
| 2. Prompt assembly | Capability | Instructions + few-shot examples + tool definitions | Prompt Engineering |
| 3. Semantic retrieval | Capability | The user's question is embedded; relevant documents are recalled from the vector database | Vector Databases & Semantic Search |
| 4. Context injection | Capability | Retrieved results are stitched into the prompt — the essence of RAG | RAG |
| 5. Model inference | Model + architecture | Autoregressive decoding, one token at a time | Large Language Models |
| 6. Output rendering | Application | Streaming output, citation markers, format validation | Perplexity and AI Search |
| 7. Evaluation & feedback | Engineering | Offline evals, online monitoring, error feedback loops | LLM Evaluation & Benchmarks |
Steps 3 and 4 are optional but strongly recommended — bare question answering without retrieval shows a markedly higher hallucination rate. The complete end-to-end build is the main body of the Build a RAG App from Scratch practice guide.
4. Training, Post-Training, Inference: the three stages of a model's life
The data-flow view answers "how does one request travel"; the timeline view answers "how does a model come to be." Every modern large model is born in three stages:
| Stage | What happens | Key techniques and scale | Concept page |
|---|---|---|---|
| Pretraining | Learns the regularities of language from massive text corpora | Transformer + data scale + compute scale | Transformer |
| Post-training | Fine-tunes for tasks and aligns values | Instruction tuning, LoRA, RLHF, DPO | Fine-Tuning & PEFT, Alignment |
| Inference | Optimizes deployment and serves live traffic | Quantization, distillation, inference engines | Inference Optimization & Quantization |
The three stages have a stark asymmetry: pretraining burns money, post-training burns brains, and inference burns machines. Pretraining sets the ceiling of intelligence (parameter and data scale); post-training sets the floor of behavior (usable? safe?); inference sets the unit cost. What makes DeepSeek-R1 remarkable is precisely that it demonstrated "keep the pretraining skeleton untouched and elicit reasoning ability through reinforcement-learning-style post-training" (see DeepSeek-R1 and Reasoning Models). For the full timeline narrative of the three stages, see A Brief History.
5. The Site-Wide Content Map: seven sections × the layer diagram
Finally, pin the whole site onto this layer diagram. The site has seven sections, each with its own job; the mapping looks like this:
| Section | Which layers it covers | Entry pages |
|---|---|---|
| Guide | A global view of all five layers plus the engineering layer | Learning Paths, Concept Boundaries |
| Concepts | The 14 concepts of the five layers plus the engineering layer, expanded one by one | Concept Overview |
| Case studies | Mostly the application layer, with forays into the model layer | 10 Classic Case Studies |
| Papers | Primary literature for the architecture and model layers | Paper Map, Reading Paths |
| Practice | Hands-on guides for the capability and engineering layers | Build a RAG App from Scratch, Fine-Tune Your Own LLM |
| Careers | The roles and skills that correspond to each layer | Job Landscape, JD Knowledge Breakdown |
| Resources | Term, dataset, and model quick references for every layer | Glossary, Model & Leaderboard Cheat Sheet |
How to use this page
- Want a quick global picture? Just read this page through.
- Want depth on one layer? Click the layer-name links into the concept pages.
- Want to build? Pick a route on Learning Paths, then get your hands dirty with Build a RAG App from Scratch and Build an Agent from Scratch.
- Prepping for interviews? Self-test against this diagram in the Interview Question Bank — if you can explain what problem each of the five layers solves, you are most of the way there.
Further Reading
- What Is AI: The Hot Concepts — the opening definitions and panoramic overview behind this page
- AI vs ML vs DL vs GenAI vs Agent — translating this page's layer names into concept boundaries
- A Brief History — why this tech stack looks the way it does
- Learning Paths — three routes; pick the one that fits you
- Transformer and the Attention Mechanism — the complete derivation for the architecture layer
- RAG and Agents — the two leads of the capability and application layers
- Common Pitfalls & Anti-Patterns — every layer's classic failure scenes
References
- Attention Is All You Need (Vaswani et al., 2017) — the original Transformer paper, the bedrock literature of the architecture layer
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., 2020) — the founding paper of RAG
- Training Language Models to Follow Instructions with Human Feedback (Ouyang et al., 2022) — InstructGPT/RLHF, the origin of alignment techniques
- LoRA: Low-Rank Adaptation of Large Language Models (Hu et al., 2021) — the original paper of the most mainstream PEFT method
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (2025) — the benchmark technical report on eliciting reasoning through post-training
- Hugging Face Docs — first-party documentation for the fine-tuning, evaluation, and inference toolchains