Theme
A Brief History
One-sentence positioning: Large language models did not appear out of nowhere in the 2020s. They are the inevitable endpoint of a nearly eighty-year thread of "using statistics and neural networks to teach machines to understand language." This article provides the complete timeline, three paradigm shifts, and a technology roadmap, helping you replace the intuitive feeling of "GPT is amazing" with a judgment of "why it had to appear this way, and where it goes next."
I. Timeline Overview
Key milestones across five phases: "ideation → statistical era → neural era → pretraining era → scale and alignment era."
| Era | Event | One-line significance |
|---|---|---|
| 1948 | Shannon publishes A Mathematical Theory of Communication, proposing information entropy | Provides the mathematical language for "describing language with probability"; the starting point of statistical language models |
| 1951 | Shannon proposes the N-gram language model | The first computable "predict the next word" model |
| 1966 | ELIZA chatbot (Weizenbaum) | Early "conversation program," pure rule matching, unrelated to learning |
| 1980s | Backpropagation algorithm popularized (Rumelhart et al., 1986) | Neural networks became trainable; the engine of deep learning |
| 1997 | Hochreiter & Schmidhuber propose LSTM | The classic architecture for long-sequence memory; sequence modeling's dominant tool until 2017 |
| 2003 | Bengio et al. publish A Neural Probabilistic Language Model | The pioneering work of neural language models: replacing n-gram counting with neural networks |
| 2013 | Word2Vec (Mikolov et al.) | Word embeddings made "semantic similarity" computable; the precursor of representation learning |
| 2014 | Seq2Seq (Sutskever et al.) and attention introduced to machine translation (Bahdanau et al.) | Sequence modeling + attention, paving the way for the Transformer |
| 2017 | Vaswani et al. publish Attention Is All You Need | The Transformer is born: pure attention, no recurrence, fully parallelizable |
| 2018.6 | GPT-1 (OpenAI, 117M parameters) | Generative pretraining + discriminative fine-tuning; the decoder-only route is established |
| 2018.10 | BERT (Google, 340M parameters) | Masked language modeling + encoder; swept 11 NLP benchmarks |
| 2019.2 | GPT-2 (OpenAI, 1.5B parameters) | Larger scale + zero-shot; demonstrates "pretraining itself is capability" |
| 2019.10 | T5 (Google, up to 11B parameters) | Encoder-decoder + text-to-text unified framework |
| 2020.5 | GPT-3 (OpenAI, 175B parameters) | Empirical validation of scaling laws: few-shot via prompting, not fine-tuning |
| 2020.1 / 2022.3 | Scaling Laws (Kaplan et al.) / Chinchilla (DeepMind) | Loss declines as a power law with scale; the optimal data-to-parameter ratio of ~1:20 |
| 2021–2022 | GShard / Switch Transformer / GLaM / LaMDA / PaLM | MoE sparsification and ultra-large scale (PaLM: 540B parameters) |
| 2022.3 | InstructGPT (OpenAI) | RLHF alignment milestone: making models "follow instructions" |
| 2022.11.30 | ChatGPT released | Alignment + conversational product form; 1 million users in 5 days, igniting the world |
| 2023.2–7 | Llama 1 / Llama 2 (Meta) open-sourced | The open-source ecosystem starting point; local deployment becomes possible |
| 2023.3 | GPT-4 (OpenAI) | Multimodal input + top 10% human performance on multiple exams; pushing capability ceilings |
| 2023.12 | Gemini (Google) / Mixtral (Mistral AI) | Closed-source multimodal benchmarking + open-source MoE for efficient inference |
| 2024 | GPT-4o (native multimodal real-time), Llama 3, DeepSeek-V3, Sora | Real-time multimodal, open-source catching up to closed-source, MoE at scale, video generation |
| 2024.9–2025 | o1 series (inference-time scaling), DeepSeek-R1 open-source reproduction | Adding "thinking time" to the capability formula; a new extension axis |
| 2024–2025 | RAG engineering, Agent (tool calling) productization | From "stronger models" to "stronger systems" |
Parameters and timeline are based on public releases. GPT-4's parameter count was never officially disclosed, so no specific number is given here.
1. The Chinese LLM Development Thread (2019–2025)
China's LLM progress started slightly later but kept pace tightly, running largely parallel to the global main thread (see Llama and the Open-Source Ecosystem and MoE and Ultra-Large Models for details):
| Era | Event | Significance |
|---|---|---|
| 2019 | Google proposes ALBERT (lightweight BERT, parameter sharing to reduce memory) | Efficient pretraining exploration |
| 2021 | Baidu's Wenxin (ERNIE) series iterates | Industry models get started |
| 2022 | Zhipu open-sources GLM-130B; Alibaba's Qwen project launched | Open-source + industry advance on two fronts |
| 2023 | Wenxin Yiyan, Qwen, ChatGLM, Baichuan, Zhipu, etc. released in rapid succession | "A hundred models battle"; Chinese LLM ecosystem takes shape |
| 2024 | Qwen open-source series, DeepSeek-V2/V3 released | Open-source models catching up to closed-source; MoE efficiency leading |
| 2025 | DeepSeek-R1 open-sourced with reproduction of inference-time scaling | Pushing global attention to "thinking models" with an open-source posture |
Global vs. China: a perspective on reading the timeline
China's catch-up priorities shifted in stages: model usability (2022–2023, solving "do we have one") → open-source and efficiency (2024, solving "can we afford it") → frontier methods (2025, participating in "how to build"). This rhythm corroborates the judgment in Frontier Progress that "open-source is catching up to closed-source."
II. Three Paradigm Shifts
Nearly eighty years of history can be compressed into three shifts. Understanding these shifts is understanding "why now."
1. First shift: From rules to statistics (1940s–1980s)
Before statistical methods, the mainstream of language processing was manual rules: linguists and programmers hand-wrote grammar rules, dictionaries, and transformation rules (e.g., ELIZA, early machine translation). The problem with this approach was obvious: you can never write enough rules; language always has new exceptions.
Shannon's information theory and N-gram brought the first turn: replacing manual rules with statistical frequencies from corpora. The core idea was that language can be described probabilistically, and the best guess for "the next word" comes from statistics of similar contexts in the corpus. This idea was directly inherited by all modern language models — the P(wₙ|w₁,…,wₙ₋₁) objective, from 1951 to 2025, unchanged. See Language Modeling.
2. Second shift: From statistics to neural networks (2003–2017)
Statistical models' fatal flaw was data sparsity: as the n-gram window grows, the combination count explodes, and most word sequences never appear in the corpus. Bengio's 2003 solution was: use neural networks to map words to continuous vectors, letting "similar semantics" share statistical strength — so unseen sentences still get reasonable probabilities.
The next decade saw neural methods spread everywhere: Word2Vec (word embeddings), LSTM (long sequences), Seq2Seq (sequence-to-sequence), attention mechanisms (soft alignment). By 2014–2016, neural machine translation had fully replaced statistical machine translation. But models in this phase still needed to be trained separately for each task — the seed of "one model for all tasks" (pretraining) had not yet sprouted.
3. Third shift: Pretraining + scale + alignment (2017–present)
The third shift includes three mutually reinforcing forces:
| Force | Representative | Content |
|---|---|---|
| Architecture | Transformer (2017) | Pure attention, fully parallelizable, highly scalable — making "training hundred-billion-parameter models" possible |
| Scale | GPT-3 / Scaling Laws (2020) | Loss declines as a power law with parameters, data, and compute; emergent capabilities appear with scale |
| Alignment | InstructGPT / ChatGPT (2022) | Pretrained models "can talk but aren't useful"; RLHF makes them "follow instructions and speak naturally" |
The completion of the shift is marked by ChatGPT in late 2022: it fills all three forces simultaneously — Transformer architecture + large-scale pretraining + deep alignment — and reached the public in conversational product form. Previous shifts changed "how models are built"; this shift changed "how humans use intelligence."
Why are shifts incremental rather than a continuous slope?
Because each shift is a "dimension switch": from writing rules to counting statistics is changing the answer source; from statistics to neural networks is changing the model family; from neural to pretraining+scale is changing the "capability generation mechanism" (learned representations + scale amplification + alignment shaping). Progress within the same dimension is climbing a slope (GPT-2 → GPT-3), while cross-dimension is a shift (statistical → neural → scale + alignment).
III. Key Technology Roadmap
Around "model forms in the Transformer era," four routes evolve in parallel. Understanding this branching lets you read the family relationships in case study pages like GPT Series and BERT and the Encoder Family.
1. Comparing the four routes
| Route | Architecture | Training objective | Representative models | Strengths |
|---|---|---|---|---|
| Decoder-only | Unidirectional autoregressive Transformer | Predict the next word | GPT-1/2/3/4, Llama, Qwen, DeepSeek | Generation; mainstream in the LLM era |
| Encoder-only | Bidirectional Transformer | Masked language modeling (MLM) | BERT, RoBERTa, ELECTRA | Understanding (classification, retrieval, embeddings) |
| Encoder-Decoder | Encoder reads in + decoder generates | Text-to-text | T5, BART, FLAN-T5 | Balance of understanding + generation (translation, summarization) |
| Hybrid / Next-gen | Sparse experts (MoE), state-space models, etc. | Various | Mixtral, DeepSeek-V3, Mamba | Efficiency and long sequences (see MoE) |
2. Why the decoder won
Before 2019, the encoder (BERT) ruled — leading on understanding-task benchmarks. After 2020, the balance reversed, and the decoder became the de facto standard. Three reasons:
- Generation subsumes understanding: A model that can generate can naturally do understanding (just make it "answer"), while understanding models struggle to cover generation.
- Unified pretraining objective: Next-token prediction needs no labels and has infinite corpora, making it more suitable for scaling laws (see Scaling Laws) than MLM.
- Alignment and conversation: RLHF, multi-turn dialogue, and tool calling all naturally build on the decoder's "autoregressive generation + instructions" form.
text
GPT route = decoder architecture + massive pretraining + scale amplification + alignment (RLHF)3. The encoder's legacy
The decoder won today's market, but the encoder route left three legacies that are still actively used:
| Legacy | Use case | Representative |
|---|---|---|
| Text vectorization (embeddings) | Input features for retrieval, clustering, semantic similarity computation | BERT, E5, bge (see the retrieval step in RAG) |
| Bidirectional understanding benchmarks | Classification, extraction, information retrieval — "understanding-first" tasks still commonly use them | RoBERTa, ELECTRA |
| Training methodology | Masked modeling, contrastive learning, and other pretraining techniques inherited by multimodal fields | Vision-language model alignment heads |
Reading BERT and the Encoder Family separately reveals: it hasn't "died" — it's just gone from "lead actor" to "infrastructure."
IV. Three Patterns from the Timeline
1. Simple objective × massive scale = complex capability
The three shifts never changed that "naively simple" training objective — predict the next word. What changed was scale (data and parameters) and architectural efficiency (Transformer's parallelism). History repeatedly proves: in this field, "smarter objectives" are far less effective than "more data + better scaling" — this is the quantitative conclusion that Scaling Laws repeatedly emphasizes.
2. Capabilities emerge indirectly from "surrogate tasks"
From Word2Vec (predicting neighbor words but learning semantic structure) to GPT (predicting the next word but learning to reason), a model's "real capabilities" were never directly required by the training task. They are structures the model was forced to learn when compressing data. History reminds us: what an LLM can or cannot do is often only known after scale is reached, and cannot be fully predicted by design alone. This is why the "capability layer" chapter in What Is a Large Language Model exists.
3. Every "breakthrough" relied on an "application form," not just raw capability
The neural language model existed in 2003, but the public perception only started with ChatGPT. Similarly, BERT's capabilities existed in 2018, but it took years for developers to use it and for products to consume it. Model capability → product form → public acceptance has a long transmission chain. Looking at history today, don't forget: behind every "sudden breakout" is a longer, less visible accumulation line.
4. Two unresolved debates on the timeline
History isn't settled doctrine. Two debates remain open; keep a critical mind when reading the timeline:
| Debate | Two sides | Why it matters |
|---|---|---|
| Is emergence a real capability or a measurement artifact? | One side argues reasoning and other capabilities "jump" as scale grows (phase transition); the other side argues that with smoother metrics, the jumps disappear, being artifacts of "non-continuous measurement." | Determines "whether capability is an inevitable result of scale" and whether we should expect bigger models to automatically be stronger. See Scaling Laws. |
| Is scale growth sustainable? | Optimists believe "more data + bigger models" will keep working; skeptics point out that high-quality text data is peaking (the data wall), and synthetic data plus inference-time scaling may change the rules. | Determines where compute and data investment goes. Inference-time scaling, MoE, and synthetic data are the three current alternative paths. See Frontier Progress. |
Neither debate has a standard answer, but asking about them is itself the ability to read history well — history provides not conclusions, but a genealogy of questions.
V. How to Keep Tracking This Timeline
History doesn't end in 2025. There are three entry points for tracking evolution:
| Entry point | Content | Site link |
|---|---|---|
| Paper map | Key papers laid out as a map by timeline and theme | Paper Map |
| Case study pages | Deep-dive breakdowns of each model family | GPT Series, ChatGPT and Conversational Models, Llama and the Open-Source Ecosystem, MoE and Ultra-Large Models, Multimodal LLMs |
| Frontier tracking | Trend summaries and representative papers from 2023–2025 | Frontier Progress |
Suggested reading order
First read What Is a Large Language Model to build concepts, then come back to this page to string the timeline together, then dive deeper along whichever branch interests you (architecture → Transformer Architecture; scale → Scaling Laws; alignment → Alignment). The value of history is not memorizing dates; it's so that when you read any new model, you can answer: which route is it on, and which step did it push forward?
1. From Timeline to First-Hand Literature
The timeline is just an index. What's truly worth reading are the papers the index points to. Here's a suggestion for "which paper to read first" by interest area:
| Interest area | First paper | Why read it first |
|---|---|---|
| Architecture | Attention Is All You Need | The common foundation of all modern architectures; the starting point of the paper map |
| Scale | Scaling Laws / Chinchilla | The quantitative main thread, explaining "why bigger is better" and the "1:20 ratio" |
| Alignment | InstructGPT | The origin of the RLHF paradigm; ChatGPT's direct predecessor |
| Generation paradigm | GPT-3 | Establishing the few-shot paradigm |
| Open-source ecosystem | The Llama paper and announcement | The starting point of "open-source catching up to closed-source" |
| Inference-time scaling | DeepSeek-R1 technical report | The newest dimension post-2024; first-hand info on papers and new models |
A foundation for reading papers
Reading papers directly can be steep. We recommend first building conceptual foundations in Language Modeling and Transformer, then deep-reading according to the "background → method → experiment → limitations" four-part template in Classic Paper Deep Dives.
Further Reading
- What Is a Large Language Model — The "endpoint" of this page's timeline: three-layer definition and capability inventory
- Transformer Architecture — A full mechanism breakdown of the 2017 shift
- Scaling Laws — The quantitative main thread post-2020: why bigger models are better
- Alignment — The decisive step of 2022: RLHF and DPO
- Paper Map — A full map translating this timeline from a "paper perspective"
- GPT Series: From GPT-1 to GPT-4o — The family history of the decoder route
References
- Shannon. A Mathematical Theory of Communication (1948) — The foundational work on information entropy and information theory, the source of statistical language models
- Bengio et al. A Neural Probabilistic Language Model (2003) — The pioneering paper of neural language models
- Vaswani et al. Attention Is All You Need (2017) — The original Transformer paper, the architectural cornerstone of the third shift
- Devlin et al. BERT: Pre-training of Deep Bidirectional Transformers (2018) — The encoder route's representative, sweeping NLP benchmarks
- Brown et al. Language Models are Few-Shot Learners (2020) — GPT-3, the opening of the scale era
- Ouyang et al. Training Language Models to Follow Instructions with Human Feedback (InstructGPT, 2022) — The original RLHF alignment paper
- OpenAI · Introducing ChatGPT (2022) — ChatGPT release notes; the public moment of the third shift
- Google DeepMind · Introducing Gemini (2023) — A representative release in the multimodal large model era