Skip to content

Overall Architecture Anatomy

At a glance A complete lifecycle map of large model systems: data engineering → pretraining → post-training → evaluation → deployment & inference → applications → feedback loop, dissecting inputs, outputs, and key questions at each stage, and mapping to pages across the site.

Overall Architecture Anatomy ​

One-sentence positioning: A large model is not "one model," but a complete pipeline from data to product. This article breaks that pipeline into seven stages, annotating for each "what's the input, what's the output, what are the key questions, and which page on this site expands on it." It's the site's master map, and your first stop for building a systems view.

I. Lifecycle Panorama ​

Visualize the complete lifecycle of a large model as a pipeline:

text
                    ┌─────────────────────────────────────────────┐
                    │              Model Training Main Thread      │
                    │                                             │
  Data Engineering ──→ Pretraining ──→ Post-training (SFT/alignment) ──→ Evaluation ──→ Deployment & Inference
      │           │            │               │          │
      │           │            │               │          │
      └───────────┴────────────┴───────────────┴── Model weights ─┘
                                                        │
                    ┌───────────────────────────────────┘
                    │              Model Usage Main Thread
                    ▼
              Applications (prompting / RAG / Agent)
                    │
                    ▼
              Feedback loop (logs / evaluation / safety monitoring) ───────→ Feeds back into data and alignment

Two main threads

The entire pipeline can be divided into two main threads:

  • Model training thread (data engineering → pretraining → post-training → evaluation → deployment): produces "model weights," with participants mainly being algorithm and training engineers.
  • Model usage thread (deployment & inference → applications → feedback loop): consumes "model capability," with participants mainly being application, platform, and product engineers.

The two threads share evaluation (assessing the model on the training side, assessing the system on the application side) and data (creating data on the training side, feeding data back from the application side). Most practitioners only occupy one segment, but seeing the full picture tells you where your work fits and where the value flows.

II. Seven-Stage Dissection ​

1. Data Engineering ​

Input: raw crawled internet corpora, books, code, multilingual text; Output: clean, deduplicated, well-mixed pretraining corpora, plus instruction/preference data for post-training.

Key questionsDescriptionRelated pages
Where does the data come from, and how is quality controlled?Cleaning, deduplication (MinHash), toxicity filteringPretraining
How to mix multilingual / code / math?"Data is the model" empirical findingDatasets and Benchmarks Archive
Copyright and licensing?Corpus compliance risksSafety and Risks

Data is an underrated bottleneck

Pretraining corpora run into the trillions of tokens, but "quantity" is far from "quality." The industry consensus is: the ceiling of data quality determines the ceiling of model capability. Many companies are stuck not on model training, but on data pipelines. See the "data as model" section in Pretraining.

2. Pretraining ​

Input: cleaned massive text corpora; Output: a "talks but isn't yet aligned" base model. This is the most expensive stage (thousands of GPUs training for months), and also the source of large model capability — the "predict the next word" training objective compresses knowledge into parameters here.

Key questionsDescriptionRelated pages
Training objective and lossNext-token prediction, cross-entropyLanguage Modeling
How text becomes tokensTokenization and vocabulariesTokenization and Vocabularies
What architecture carries itTransformer, positional encodingTransformer Architecture
How to determine scaleOptimal mix of parameters, data, computeScaling Laws
How to save computeMoE sparse expertsMoE Sparse Expert Models
How to schedule trainingLearning rate, batch, distributed parallelismPretraining

3. Post-training (SFT and Alignment) ​

Input: base model + instruction/preference data; Output: an "obedient, natural-speaking, boundary-aware" assistant model (instruct/chat model). Pretraining teaches capability; post-training teaches how to use it.

Key questionsDescriptionRelated pages
How to teach the model "to follow instructions"?Supervised fine-tuning (SFT)Fine-tuning
How to teach the model "to conform to human preferences"?RLHF / DPOAlignment
How to change only some parameters?LoRA, QLoRA, and other parameter-efficient fine-tuning methodsFine-tuning, Fine-tuning in Practice
Does capability degrade after fine-tuning?Catastrophic forgettingFine-tuning

Post-training is the "unsung hero" post-2022

Base models already "can generate," but generation does not equal usefulness. ChatGPT looks more like an "assistant" than a base model precisely because of alignment. Viewing post-training and pretraining as separate is a common beginner mistake — one gives capability, the other gives form. Both are essential.

4. Evaluation ​

Input: base/post-trained model + evaluation benchmarks; Output: capability profile (what it's strong at, what it's weak at, whether it has degraded). Evaluation spans both threads: the training side asks "did the model do it right?" and the application side asks "is the system good to use?"

Key questionsDescriptionRelated pages
Which benchmarks to use?MMLU / GSM8K / HumanEval, etc.Evaluation and Benchmarks
How to build a custom eval set?Golden set, LLM-as-a-judgeEvaluation in Practice
How to quantify hallucination?Factuality evaluationHallucination
What pitfalls do benchmarks have?Data contamination, leaderboard gamingEvaluation and Benchmarks

5. Deployment and Inference ​

Input: trained model weights + inference requests; Output: low-latency, high-throughput online service. Inference and training are two completely different engineering disciplines: training optimizes for "throughput," while inference optimizes for "latency + GPU memory."

Key questionsDescriptionRelated pages
How does the generation process work?Autoregression, KV Cache, samplingInference Fundamentals
How to estimate GPU memory?Weights + KV Cache + activationsDeployment and Servicing
How to speed up and cut costs?Quantization, continuous batchingDeployment and Servicing
Which frameworks to use?vLLM / SGLang / TensorRT-LLMFramework and Tool Selection

6. Applications (Prompting / RAG / Agent) ​

Input: model service + user requests + application orchestration; Output: usable product features. This is the core of the model usage thread, and where most engineers actually work.

Application formMechanismRelated pages
Direct promptingWrite a good prompt, let the model answerPrompting, Prompting in Practice
RAGRetrieve external knowledge + generateRAG, RAG in Practice
AgentModel + tools + planning loopAgents with LLMs
Fine-tuning adaptationModify the model for specific tasksFine-tuning in Practice
MultimodalText + image/audioMultimodal LLMs

7. Feedback Loop ​

Input: production logs, user feedback, safety incidents, evaluation failure samples; Output: new data annotation needs, alignment fixes, evaluation set expansion. The closed loop makes the pipeline "get better with use."

Feedback directionPurpose
Back to data engineeringAccumulate real user questions; produce higher-quality SFT/preference data
Back to alignmentDiscover harmful outputs; supplement alignment samples and guardrails
Back to evaluationProduction failure samples enter regression test sets (golden set)
Back to productError patterns inform prompt wording and RAG strategy

The feedback loop is what separates "engineers" from "API callers"

Calling APIs without building a feedback loop means outsourcing all quality improvement to the model vendors. Building your own evaluation and log feedback — even at small scale — is the watershed between "can use" and "can build." See Evaluation in Practice.

III. Stage Quick Reference Table ​

StageInputOutputKey pagesRoles
① Data engineeringRaw corporaCleaned, mixed corpora / instruction dataPretraining, Datasets and Benchmarks ArchiveData engineers
② PretrainingMassive corporaBase modelLanguage Modeling, TransformerTraining engineers, researchers
③ Post-trainingBase model + instruction/preference dataAligned assistant modelFine-tuning, AlignmentAlignment engineers
④ EvaluationModel + benchmarksCapability profileEvaluation and BenchmarksEvaluation engineers
⑤ Deployment & inferenceModel weights + requestsOnline serviceInference Fundamentals, DeploymentInference / platform engineers
⑥ ApplicationsService + user + orchestrationProduct featuresRAG, AgentApplication / algorithm engineers
⑦ Feedback loopLogs + failure samplesNew data and fixesEvaluation in PracticeFull chain

1. Typical Failure Modes at Each Stage ​

Each stage has a recurring "failure point." Knowing what failure looks like in advance is far more efficient than debugging after the fact:

StageTypical failureSymptomsCountermeasure
① DataData contaminationInflated eval scores, poor production performanceSeparate train/eval data; deduplicate and audit provenance (see Datasets and Benchmarks Archive)
② PretrainingMixed-data imbalanceGood Chinese, poor English; poor codeReview per Scaling Laws and corpus mixing tables
③ Post-trainingAlignment taxGeneral capability degrades after fine-tuningMix in general-purpose data; control training steps (see Fine-tuning)
④ EvaluationMeasuring the wrong thingHigh leaderboard scores but no business upliftAdd a custom golden set; run regression tests (see Evaluation in Practice)
⑤ DeploymentMemory / latency out of controlOOM or timeout on launchEstimate first, then select; fallback to quantization and batching (see Deployment and Servicing)
⑥ ApplicationsOver-promisingAsking the model to do things it's bad at (precise computation, real-time facts)Capability-tier selection + tool/retrieval supplementation (see Prompting in Practice)
⑦ Feedback loopBroken loopProduction issues never make it into training dataFeed failure samples into annotation and evaluation pipelines (see Common Pitfalls)

IV. Key Question Checklist ​

Each of the seven stages has a "make-or-break" question that comes up in interviews and project management:

text
① Data engineering:      Is the corpus clean enough? Is the mix reasonable?
② Pretraining:           Does the data-to-params-to-compute ratio follow scaling laws?
③ Post-training:         After alignment, is the balance of capability and safety correct?
④ Evaluation:            Are we measuring what we actually care about?
⑤ Deployment & inference: Latency, throughput, cost — which two of the triangle did you pick?
⑥ Applications:          Prompting, RAG, or Agent — which combo fits the current scenario?
⑦ Feedback loop:         Did the failure samples actually make it back into training data?

1. From Questions to Decisions: Two Decision Chains ​

The seven questions can be compressed into two decision chains, the most commonly asked in engineering:

Training-side decision chain (should we train our own model?)

text
Do we have high-quality proprietary data? ── No ──→ Just use existing models (API or open-source)
        │ Yes
        ↓
Can we afford training compute? ── No ──→ Fine-tune (LoRA) instead of pretraining
        │ Yes
        ↓
Does the data/params/compute ratio follow scaling laws? ──→ Refer to [Scaling Laws](/concepts/scaling-laws) to set budget
        ↓
Do we need alignment after training? ──→ Enter the [Alignment](/concepts/alignment) stage

Application-side decision chain (which combo for the current scenario?)

text
Can the task be solved by writing a good prompt? ── Yes ──→ [Prompting in Practice](/practice/prompting-practice)
        │ No
        ↓
Do we need private/real-time knowledge? ── Yes ──→ [RAG in Practice](/practice/rag-in-practice)
        │ No
        ↓
Do we need multi-step action and tools? ── Yes ──→ [Agents with LLMs](/case-studies/agents-with-llm)
        │ No
        ↓
Do we need fixed format / stable style? ── Yes ──→ Consider [fine-tuning](/concepts/fine-tuning)
        ↓
        └──→ Go back to [Evaluation and Benchmarks](/concepts/evaluation) to re-validate the task definition

Both decision chains end at evaluation — without evaluation, any "selection" is a guess. This is why evaluation sits at the pivot point of the lifecycle.

V. Talent Capability Map ​

Viewing roles from the lifecycle perspective, each stage is a class of role with a corresponding set of skills (see Careers & JD's JD breakdown for details):

StageRoleCore skillsRelated concept pages
① Data engineeringData engineer / corpus engineerCrawling, cleaning, deduplication, mixingPretraining, Datasets and Benchmarks Archive
② PretrainingTraining engineer / researcherDistributed training, scaling laws, hyperparameter tuningPretraining, Scaling Laws
③ Post-trainingAlignment engineer / fine-tuning engineerSFT, RLHF/DPO, LoRAFine-tuning, Alignment, Fine-tuning in Practice
④ EvaluationEvaluation engineerBenchmarks, evaluation set design, LLM-as-a-judgeEvaluation and Benchmarks, Evaluation in Practice
⑤ Deployment & inferenceInference optimization / platform engineerKV Cache, quantization, batchingInference Fundamentals, Deployment
⑥ ApplicationsApplication algorithm / Agent engineerPrompting, RAG, tool calling, productizationRAG, Agent, Prompting in Practice
⑦ Full chainTech lead / head of algorithmSystem architecture, evaluation framework, cost managementCommon Pitfalls

One role often spans multiple stages

A frontline "large model algorithm engineer" typically spans ③④⑥ (fine-tuning + evaluation + applications), while a "training engineer" focuses on ②. In interviews, first ask which stage the candidate's role maps to, then prepare accordingly — this "role map" thinking is repeatedly emphasized in Careers & JD.

1. Self-Assessment Checklist ​

Go through the seven stages, score yourself in three tiers (proficient / familiar / blank), and identify your next step:

StageSelf-assessment questionProficientFamiliarBlank
① DataCan you explain deduplication (MinHash) and data mixing?☐☐☐
② PretrainingCan you draw the next-token training loop?☐☐☐
③ Post-trainingCan you explain the three steps of RLHF and LoRA's principle?☐☐☐
④ EvaluationCan you design a 20-item golden set?☐☐☐
⑤ DeploymentCan you estimate the GPU memory for a 7B INT8 model?☐☐☐
⑥ ApplicationsCan you explain the selection boundaries of prompting / RAG / Agent?☐☐☐
⑦ Feedback loopHave production failure samples made it into the eval set?☐☐☐

How to use: For "blank", go read the corresponding concept page and practice page; for "familiar", write a one-pager to explain it to yourself — if you can't explain it clearly, reread; for "proficient", try explaining it to someone else or write a blog post. This checklist is also a self-assessment version of the "JD knowledge point breakdown" in Careers & JD.

VI. Where to Start ​

With the panorama, how do you make it concrete? Three options:

  1. Read the seven stages in order: Follow the "Systematic Deep Dive" path in Learning Paths, reading the corresponding pages for each of the seven stages one by one.
  2. Start by building one stage: Begin with Building a Large Model from Scratch, get the minimum viable version of ② running, then fill in the other stages.
  3. Put yourself on the map first: Against the role landscape in Careers & JD, find your position, and use the Interview Question Bank to check your weak spots.

1. The 48-Hour Minimum Closed Loop ​

If you want to "run through" all seven stages in the least time possible, you can complete a minimal closed loop in 48 hours (using an open-source small model with local deployment or an API):

Time slotActionStages covered
Hour 1Pick a narrow task (e.g., classify user questions into 5 intents); hand-write 20 golden-set examples④ Evaluation design
Hours 2–6Pick a model, write the first version of the prompt, run baseline scores⑥ Applications
Hours 7–24Iterate on prompts and examples following Evaluation in Practice methods, logging every change and score④ Evaluation + ⑥ Applications
Hours 25–40Wrap the service with an API gateway or local vLLM; attach request logs⑤ Deployment
Hours 41–48Collect failure samples; go back to the golden set to add examples; write a one-page retrospective⑦ Feedback loop

This loop deliberately avoids ①②③ (data, pretraining, post-training) — that's the "model side." This minimal loop demonstrates the full "application side" chain. After running it, you'll have personally experienced a complete model lifecycle. Remaining questions (should I fine-tune? should I use RAG?) will naturally surface, and the corresponding pages will have context.

One-sentence summary

Large model engineering = data × model × alignment × evaluation × deployment × applications — six essential elements, seven interlocking stages. The map is drawn; pick a stage and go — check the Glossary for terms, Mainstream Model Profiles for model info.

Further Reading ​

References ​