Theme
Overall Architecture Anatomy
One-sentence positioning: A large model is not "one model," but a complete pipeline from data to product. This article breaks that pipeline into seven stages, annotating for each "what's the input, what's the output, what are the key questions, and which page on this site expands on it." It's the site's master map, and your first stop for building a systems view.
I. Lifecycle Panorama
Visualize the complete lifecycle of a large model as a pipeline:
text
┌─────────────────────────────────────────────┐
│ Model Training Main Thread │
│ │
Data Engineering ──→ Pretraining ──→ Post-training (SFT/alignment) ──→ Evaluation ──→ Deployment & Inference
│ │ │ │ │
│ │ │ │ │
└───────────┴────────────┴───────────────┴── Model weights ─┘
│
┌───────────────────────────────────┘
│ Model Usage Main Thread
▼
Applications (prompting / RAG / Agent)
│
▼
Feedback loop (logs / evaluation / safety monitoring) ───────→ Feeds back into data and alignmentTwo main threads
The entire pipeline can be divided into two main threads:
- Model training thread (data engineering → pretraining → post-training → evaluation → deployment): produces "model weights," with participants mainly being algorithm and training engineers.
- Model usage thread (deployment & inference → applications → feedback loop): consumes "model capability," with participants mainly being application, platform, and product engineers.
The two threads share evaluation (assessing the model on the training side, assessing the system on the application side) and data (creating data on the training side, feeding data back from the application side). Most practitioners only occupy one segment, but seeing the full picture tells you where your work fits and where the value flows.
II. Seven-Stage Dissection
1. Data Engineering
Input: raw crawled internet corpora, books, code, multilingual text; Output: clean, deduplicated, well-mixed pretraining corpora, plus instruction/preference data for post-training.
| Key questions | Description | Related pages |
|---|---|---|
| Where does the data come from, and how is quality controlled? | Cleaning, deduplication (MinHash), toxicity filtering | Pretraining |
| How to mix multilingual / code / math? | "Data is the model" empirical finding | Datasets and Benchmarks Archive |
| Copyright and licensing? | Corpus compliance risks | Safety and Risks |
Data is an underrated bottleneck
Pretraining corpora run into the trillions of tokens, but "quantity" is far from "quality." The industry consensus is: the ceiling of data quality determines the ceiling of model capability. Many companies are stuck not on model training, but on data pipelines. See the "data as model" section in Pretraining.
2. Pretraining
Input: cleaned massive text corpora; Output: a "talks but isn't yet aligned" base model. This is the most expensive stage (thousands of GPUs training for months), and also the source of large model capability — the "predict the next word" training objective compresses knowledge into parameters here.
| Key questions | Description | Related pages |
|---|---|---|
| Training objective and loss | Next-token prediction, cross-entropy | Language Modeling |
| How text becomes tokens | Tokenization and vocabularies | Tokenization and Vocabularies |
| What architecture carries it | Transformer, positional encoding | Transformer Architecture |
| How to determine scale | Optimal mix of parameters, data, compute | Scaling Laws |
| How to save compute | MoE sparse experts | MoE Sparse Expert Models |
| How to schedule training | Learning rate, batch, distributed parallelism | Pretraining |
3. Post-training (SFT and Alignment)
Input: base model + instruction/preference data; Output: an "obedient, natural-speaking, boundary-aware" assistant model (instruct/chat model). Pretraining teaches capability; post-training teaches how to use it.
| Key questions | Description | Related pages |
|---|---|---|
| How to teach the model "to follow instructions"? | Supervised fine-tuning (SFT) | Fine-tuning |
| How to teach the model "to conform to human preferences"? | RLHF / DPO | Alignment |
| How to change only some parameters? | LoRA, QLoRA, and other parameter-efficient fine-tuning methods | Fine-tuning, Fine-tuning in Practice |
| Does capability degrade after fine-tuning? | Catastrophic forgetting | Fine-tuning |
Post-training is the "unsung hero" post-2022
Base models already "can generate," but generation does not equal usefulness. ChatGPT looks more like an "assistant" than a base model precisely because of alignment. Viewing post-training and pretraining as separate is a common beginner mistake — one gives capability, the other gives form. Both are essential.
4. Evaluation
Input: base/post-trained model + evaluation benchmarks; Output: capability profile (what it's strong at, what it's weak at, whether it has degraded). Evaluation spans both threads: the training side asks "did the model do it right?" and the application side asks "is the system good to use?"
| Key questions | Description | Related pages |
|---|---|---|
| Which benchmarks to use? | MMLU / GSM8K / HumanEval, etc. | Evaluation and Benchmarks |
| How to build a custom eval set? | Golden set, LLM-as-a-judge | Evaluation in Practice |
| How to quantify hallucination? | Factuality evaluation | Hallucination |
| What pitfalls do benchmarks have? | Data contamination, leaderboard gaming | Evaluation and Benchmarks |
5. Deployment and Inference
Input: trained model weights + inference requests; Output: low-latency, high-throughput online service. Inference and training are two completely different engineering disciplines: training optimizes for "throughput," while inference optimizes for "latency + GPU memory."
| Key questions | Description | Related pages |
|---|---|---|
| How does the generation process work? | Autoregression, KV Cache, sampling | Inference Fundamentals |
| How to estimate GPU memory? | Weights + KV Cache + activations | Deployment and Servicing |
| How to speed up and cut costs? | Quantization, continuous batching | Deployment and Servicing |
| Which frameworks to use? | vLLM / SGLang / TensorRT-LLM | Framework and Tool Selection |
6. Applications (Prompting / RAG / Agent)
Input: model service + user requests + application orchestration; Output: usable product features. This is the core of the model usage thread, and where most engineers actually work.
| Application form | Mechanism | Related pages |
|---|---|---|
| Direct prompting | Write a good prompt, let the model answer | Prompting, Prompting in Practice |
| RAG | Retrieve external knowledge + generate | RAG, RAG in Practice |
| Agent | Model + tools + planning loop | Agents with LLMs |
| Fine-tuning adaptation | Modify the model for specific tasks | Fine-tuning in Practice |
| Multimodal | Text + image/audio | Multimodal LLMs |
7. Feedback Loop
Input: production logs, user feedback, safety incidents, evaluation failure samples; Output: new data annotation needs, alignment fixes, evaluation set expansion. The closed loop makes the pipeline "get better with use."
| Feedback direction | Purpose |
|---|---|
| Back to data engineering | Accumulate real user questions; produce higher-quality SFT/preference data |
| Back to alignment | Discover harmful outputs; supplement alignment samples and guardrails |
| Back to evaluation | Production failure samples enter regression test sets (golden set) |
| Back to product | Error patterns inform prompt wording and RAG strategy |
The feedback loop is what separates "engineers" from "API callers"
Calling APIs without building a feedback loop means outsourcing all quality improvement to the model vendors. Building your own evaluation and log feedback — even at small scale — is the watershed between "can use" and "can build." See Evaluation in Practice.
III. Stage Quick Reference Table
| Stage | Input | Output | Key pages | Roles |
|---|---|---|---|---|
| ① Data engineering | Raw corpora | Cleaned, mixed corpora / instruction data | Pretraining, Datasets and Benchmarks Archive | Data engineers |
| ② Pretraining | Massive corpora | Base model | Language Modeling, Transformer | Training engineers, researchers |
| ③ Post-training | Base model + instruction/preference data | Aligned assistant model | Fine-tuning, Alignment | Alignment engineers |
| ④ Evaluation | Model + benchmarks | Capability profile | Evaluation and Benchmarks | Evaluation engineers |
| ⑤ Deployment & inference | Model weights + requests | Online service | Inference Fundamentals, Deployment | Inference / platform engineers |
| ⑥ Applications | Service + user + orchestration | Product features | RAG, Agent | Application / algorithm engineers |
| ⑦ Feedback loop | Logs + failure samples | New data and fixes | Evaluation in Practice | Full chain |
1. Typical Failure Modes at Each Stage
Each stage has a recurring "failure point." Knowing what failure looks like in advance is far more efficient than debugging after the fact:
| Stage | Typical failure | Symptoms | Countermeasure |
|---|---|---|---|
| ① Data | Data contamination | Inflated eval scores, poor production performance | Separate train/eval data; deduplicate and audit provenance (see Datasets and Benchmarks Archive) |
| ② Pretraining | Mixed-data imbalance | Good Chinese, poor English; poor code | Review per Scaling Laws and corpus mixing tables |
| ③ Post-training | Alignment tax | General capability degrades after fine-tuning | Mix in general-purpose data; control training steps (see Fine-tuning) |
| ④ Evaluation | Measuring the wrong thing | High leaderboard scores but no business uplift | Add a custom golden set; run regression tests (see Evaluation in Practice) |
| ⑤ Deployment | Memory / latency out of control | OOM or timeout on launch | Estimate first, then select; fallback to quantization and batching (see Deployment and Servicing) |
| ⑥ Applications | Over-promising | Asking the model to do things it's bad at (precise computation, real-time facts) | Capability-tier selection + tool/retrieval supplementation (see Prompting in Practice) |
| ⑦ Feedback loop | Broken loop | Production issues never make it into training data | Feed failure samples into annotation and evaluation pipelines (see Common Pitfalls) |
IV. Key Question Checklist
Each of the seven stages has a "make-or-break" question that comes up in interviews and project management:
text
① Data engineering: Is the corpus clean enough? Is the mix reasonable?
② Pretraining: Does the data-to-params-to-compute ratio follow scaling laws?
③ Post-training: After alignment, is the balance of capability and safety correct?
④ Evaluation: Are we measuring what we actually care about?
⑤ Deployment & inference: Latency, throughput, cost — which two of the triangle did you pick?
⑥ Applications: Prompting, RAG, or Agent — which combo fits the current scenario?
⑦ Feedback loop: Did the failure samples actually make it back into training data?1. From Questions to Decisions: Two Decision Chains
The seven questions can be compressed into two decision chains, the most commonly asked in engineering:
Training-side decision chain (should we train our own model?)
text
Do we have high-quality proprietary data? ── No ──→ Just use existing models (API or open-source)
│ Yes
↓
Can we afford training compute? ── No ──→ Fine-tune (LoRA) instead of pretraining
│ Yes
↓
Does the data/params/compute ratio follow scaling laws? ──→ Refer to [Scaling Laws](/concepts/scaling-laws) to set budget
↓
Do we need alignment after training? ──→ Enter the [Alignment](/concepts/alignment) stageApplication-side decision chain (which combo for the current scenario?)
text
Can the task be solved by writing a good prompt? ── Yes ──→ [Prompting in Practice](/practice/prompting-practice)
│ No
↓
Do we need private/real-time knowledge? ── Yes ──→ [RAG in Practice](/practice/rag-in-practice)
│ No
↓
Do we need multi-step action and tools? ── Yes ──→ [Agents with LLMs](/case-studies/agents-with-llm)
│ No
↓
Do we need fixed format / stable style? ── Yes ──→ Consider [fine-tuning](/concepts/fine-tuning)
↓
└──→ Go back to [Evaluation and Benchmarks](/concepts/evaluation) to re-validate the task definitionBoth decision chains end at evaluation — without evaluation, any "selection" is a guess. This is why evaluation sits at the pivot point of the lifecycle.
V. Talent Capability Map
Viewing roles from the lifecycle perspective, each stage is a class of role with a corresponding set of skills (see Careers & JD's JD breakdown for details):
| Stage | Role | Core skills | Related concept pages |
|---|---|---|---|
| ① Data engineering | Data engineer / corpus engineer | Crawling, cleaning, deduplication, mixing | Pretraining, Datasets and Benchmarks Archive |
| ② Pretraining | Training engineer / researcher | Distributed training, scaling laws, hyperparameter tuning | Pretraining, Scaling Laws |
| ③ Post-training | Alignment engineer / fine-tuning engineer | SFT, RLHF/DPO, LoRA | Fine-tuning, Alignment, Fine-tuning in Practice |
| ④ Evaluation | Evaluation engineer | Benchmarks, evaluation set design, LLM-as-a-judge | Evaluation and Benchmarks, Evaluation in Practice |
| ⑤ Deployment & inference | Inference optimization / platform engineer | KV Cache, quantization, batching | Inference Fundamentals, Deployment |
| ⑥ Applications | Application algorithm / Agent engineer | Prompting, RAG, tool calling, productization | RAG, Agent, Prompting in Practice |
| ⑦ Full chain | Tech lead / head of algorithm | System architecture, evaluation framework, cost management | Common Pitfalls |
One role often spans multiple stages
A frontline "large model algorithm engineer" typically spans ③④⑥ (fine-tuning + evaluation + applications), while a "training engineer" focuses on ②. In interviews, first ask which stage the candidate's role maps to, then prepare accordingly — this "role map" thinking is repeatedly emphasized in Careers & JD.
1. Self-Assessment Checklist
Go through the seven stages, score yourself in three tiers (proficient / familiar / blank), and identify your next step:
| Stage | Self-assessment question | Proficient | Familiar | Blank |
|---|---|---|---|---|
| ① Data | Can you explain deduplication (MinHash) and data mixing? | ☐ | ☐ | ☐ |
| ② Pretraining | Can you draw the next-token training loop? | ☐ | ☐ | ☐ |
| ③ Post-training | Can you explain the three steps of RLHF and LoRA's principle? | ☐ | ☐ | ☐ |
| ④ Evaluation | Can you design a 20-item golden set? | ☐ | ☐ | ☐ |
| ⑤ Deployment | Can you estimate the GPU memory for a 7B INT8 model? | ☐ | ☐ | ☐ |
| ⑥ Applications | Can you explain the selection boundaries of prompting / RAG / Agent? | ☐ | ☐ | ☐ |
| ⑦ Feedback loop | Have production failure samples made it into the eval set? | ☐ | ☐ | ☐ |
How to use: For "blank", go read the corresponding concept page and practice page; for "familiar", write a one-pager to explain it to yourself — if you can't explain it clearly, reread; for "proficient", try explaining it to someone else or write a blog post. This checklist is also a self-assessment version of the "JD knowledge point breakdown" in Careers & JD.
VI. Where to Start
With the panorama, how do you make it concrete? Three options:
- Read the seven stages in order: Follow the "Systematic Deep Dive" path in Learning Paths, reading the corresponding pages for each of the seven stages one by one.
- Start by building one stage: Begin with Building a Large Model from Scratch, get the minimum viable version of ② running, then fill in the other stages.
- Put yourself on the map first: Against the role landscape in Careers & JD, find your position, and use the Interview Question Bank to check your weak spots.
1. The 48-Hour Minimum Closed Loop
If you want to "run through" all seven stages in the least time possible, you can complete a minimal closed loop in 48 hours (using an open-source small model with local deployment or an API):
| Time slot | Action | Stages covered |
|---|---|---|
| Hour 1 | Pick a narrow task (e.g., classify user questions into 5 intents); hand-write 20 golden-set examples | ④ Evaluation design |
| Hours 2–6 | Pick a model, write the first version of the prompt, run baseline scores | ⑥ Applications |
| Hours 7–24 | Iterate on prompts and examples following Evaluation in Practice methods, logging every change and score | ④ Evaluation + ⑥ Applications |
| Hours 25–40 | Wrap the service with an API gateway or local vLLM; attach request logs | ⑤ Deployment |
| Hours 41–48 | Collect failure samples; go back to the golden set to add examples; write a one-page retrospective | ⑦ Feedback loop |
This loop deliberately avoids ①②③ (data, pretraining, post-training) — that's the "model side." This minimal loop demonstrates the full "application side" chain. After running it, you'll have personally experienced a complete model lifecycle. Remaining questions (should I fine-tune? should I use RAG?) will naturally surface, and the corresponding pages will have context.
One-sentence summary
Large model engineering = data × model × alignment × evaluation × deployment × applications — six essential elements, seven interlocking stages. The map is drawn; pick a stage and go — check the Glossary for terms, Mainstream Model Profiles for model info.
Further Reading
- Learning Paths: Three Routes — Translating the seven stages on this page into an 8-week reading plan
- What Is a Large Language Model — The "product" of the lifecycle: a three-layer definition of an LLM
- A Brief History — How the seven stages were filled in step by step by history
- Language Modeling — Stage ②'s foundation: the next-word prediction paradigm
- Alignment — Stage ③ fully expanded: RLHF and DPO
- Deployment and Servicing — Stage ⑤'s engineering details: memory estimation and quantization
References
- Andrej Karpathy · State of GPT (lecture video, 2023) — Karpathy details the full GPT pipeline from data to inference; the most intuitive popularization of the "lifecycle panorama" concept
- Hugging Face · Transformers Official Documentation — Authoritative docs for architecture implementation and training/inference toolchains
- OpenAI · Language Models are Few-Shot Learners (2020) — The original paper on pretraining paradigm and scaling effects
- Google · Attention Is All You Need (2017) — The original paper on the stage ② architectural cornerstone
- OpenAI · Training Language Models to Follow Instructions (InstructGPT, 2022) — The original paper on stage ③'s alignment paradigm
- Anthropic · Constitutional AI (blog) — Industrial practice reference for alignment engineering and feedback loops