Skip to content

A Brief History

At a glance From information theory in 1948 to multimodal and agents in 2025, nearly eighty years of large language model evolution: N-gram, neural language models, Transformers, GPT/BERT, scaling laws, RLHF, and ChatGPT's three paradigm shifts.

A Brief History ​

One-sentence positioning: Large language models did not appear out of nowhere in the 2020s. They are the inevitable endpoint of a nearly eighty-year thread of "using statistics and neural networks to teach machines to understand language." This article provides the complete timeline, three paradigm shifts, and a technology roadmap, helping you replace the intuitive feeling of "GPT is amazing" with a judgment of "why it had to appear this way, and where it goes next."

I. Timeline Overview ​

Key milestones across five phases: "ideation → statistical era → neural era → pretraining era → scale and alignment era."

EraEventOne-line significance
1948Shannon publishes A Mathematical Theory of Communication, proposing information entropyProvides the mathematical language for "describing language with probability"; the starting point of statistical language models
1951Shannon proposes the N-gram language modelThe first computable "predict the next word" model
1966ELIZA chatbot (Weizenbaum)Early "conversation program," pure rule matching, unrelated to learning
1980sBackpropagation algorithm popularized (Rumelhart et al., 1986)Neural networks became trainable; the engine of deep learning
1997Hochreiter & Schmidhuber propose LSTMThe classic architecture for long-sequence memory; sequence modeling's dominant tool until 2017
2003Bengio et al. publish A Neural Probabilistic Language ModelThe pioneering work of neural language models: replacing n-gram counting with neural networks
2013Word2Vec (Mikolov et al.)Word embeddings made "semantic similarity" computable; the precursor of representation learning
2014Seq2Seq (Sutskever et al.) and attention introduced to machine translation (Bahdanau et al.)Sequence modeling + attention, paving the way for the Transformer
2017Vaswani et al. publish Attention Is All You NeedThe Transformer is born: pure attention, no recurrence, fully parallelizable
2018.6GPT-1 (OpenAI, 117M parameters)Generative pretraining + discriminative fine-tuning; the decoder-only route is established
2018.10BERT (Google, 340M parameters)Masked language modeling + encoder; swept 11 NLP benchmarks
2019.2GPT-2 (OpenAI, 1.5B parameters)Larger scale + zero-shot; demonstrates "pretraining itself is capability"
2019.10T5 (Google, up to 11B parameters)Encoder-decoder + text-to-text unified framework
2020.5GPT-3 (OpenAI, 175B parameters)Empirical validation of scaling laws: few-shot via prompting, not fine-tuning
2020.1 / 2022.3Scaling Laws (Kaplan et al.) / Chinchilla (DeepMind)Loss declines as a power law with scale; the optimal data-to-parameter ratio of ~1:20
2021–2022GShard / Switch Transformer / GLaM / LaMDA / PaLMMoE sparsification and ultra-large scale (PaLM: 540B parameters)
2022.3InstructGPT (OpenAI)RLHF alignment milestone: making models "follow instructions"
2022.11.30ChatGPT releasedAlignment + conversational product form; 1 million users in 5 days, igniting the world
2023.2–7Llama 1 / Llama 2 (Meta) open-sourcedThe open-source ecosystem starting point; local deployment becomes possible
2023.3GPT-4 (OpenAI)Multimodal input + top 10% human performance on multiple exams; pushing capability ceilings
2023.12Gemini (Google) / Mixtral (Mistral AI)Closed-source multimodal benchmarking + open-source MoE for efficient inference
2024GPT-4o (native multimodal real-time), Llama 3, DeepSeek-V3, SoraReal-time multimodal, open-source catching up to closed-source, MoE at scale, video generation
2024.9–2025o1 series (inference-time scaling), DeepSeek-R1 open-source reproductionAdding "thinking time" to the capability formula; a new extension axis
2024–2025RAG engineering, Agent (tool calling) productizationFrom "stronger models" to "stronger systems"

Parameters and timeline are based on public releases. GPT-4's parameter count was never officially disclosed, so no specific number is given here.

1. The Chinese LLM Development Thread (2019–2025) ​

China's LLM progress started slightly later but kept pace tightly, running largely parallel to the global main thread (see Llama and the Open-Source Ecosystem and MoE and Ultra-Large Models for details):

EraEventSignificance
2019Google proposes ALBERT (lightweight BERT, parameter sharing to reduce memory)Efficient pretraining exploration
2021Baidu's Wenxin (ERNIE) series iteratesIndustry models get started
2022Zhipu open-sources GLM-130B; Alibaba's Qwen project launchedOpen-source + industry advance on two fronts
2023Wenxin Yiyan, Qwen, ChatGLM, Baichuan, Zhipu, etc. released in rapid succession"A hundred models battle"; Chinese LLM ecosystem takes shape
2024Qwen open-source series, DeepSeek-V2/V3 releasedOpen-source models catching up to closed-source; MoE efficiency leading
2025DeepSeek-R1 open-sourced with reproduction of inference-time scalingPushing global attention to "thinking models" with an open-source posture

Global vs. China: a perspective on reading the timeline

China's catch-up priorities shifted in stages: model usability (2022–2023, solving "do we have one") → open-source and efficiency (2024, solving "can we afford it") → frontier methods (2025, participating in "how to build"). This rhythm corroborates the judgment in Frontier Progress that "open-source is catching up to closed-source."

II. Three Paradigm Shifts ​

Nearly eighty years of history can be compressed into three shifts. Understanding these shifts is understanding "why now."

1. First shift: From rules to statistics (1940s–1980s) ​

Before statistical methods, the mainstream of language processing was manual rules: linguists and programmers hand-wrote grammar rules, dictionaries, and transformation rules (e.g., ELIZA, early machine translation). The problem with this approach was obvious: you can never write enough rules; language always has new exceptions.

Shannon's information theory and N-gram brought the first turn: replacing manual rules with statistical frequencies from corpora. The core idea was that language can be described probabilistically, and the best guess for "the next word" comes from statistics of similar contexts in the corpus. This idea was directly inherited by all modern language models — the P(wₙ|w₁,…,wₙ₋₁) objective, from 1951 to 2025, unchanged. See Language Modeling.

2. Second shift: From statistics to neural networks (2003–2017) ​

Statistical models' fatal flaw was data sparsity: as the n-gram window grows, the combination count explodes, and most word sequences never appear in the corpus. Bengio's 2003 solution was: use neural networks to map words to continuous vectors, letting "similar semantics" share statistical strength — so unseen sentences still get reasonable probabilities.

The next decade saw neural methods spread everywhere: Word2Vec (word embeddings), LSTM (long sequences), Seq2Seq (sequence-to-sequence), attention mechanisms (soft alignment). By 2014–2016, neural machine translation had fully replaced statistical machine translation. But models in this phase still needed to be trained separately for each task — the seed of "one model for all tasks" (pretraining) had not yet sprouted.

3. Third shift: Pretraining + scale + alignment (2017–present) ​

The third shift includes three mutually reinforcing forces:

ForceRepresentativeContent
ArchitectureTransformer (2017)Pure attention, fully parallelizable, highly scalable — making "training hundred-billion-parameter models" possible
ScaleGPT-3 / Scaling Laws (2020)Loss declines as a power law with parameters, data, and compute; emergent capabilities appear with scale
AlignmentInstructGPT / ChatGPT (2022)Pretrained models "can talk but aren't useful"; RLHF makes them "follow instructions and speak naturally"

The completion of the shift is marked by ChatGPT in late 2022: it fills all three forces simultaneously — Transformer architecture + large-scale pretraining + deep alignment — and reached the public in conversational product form. Previous shifts changed "how models are built"; this shift changed "how humans use intelligence."

Why are shifts incremental rather than a continuous slope?

Because each shift is a "dimension switch": from writing rules to counting statistics is changing the answer source; from statistics to neural networks is changing the model family; from neural to pretraining+scale is changing the "capability generation mechanism" (learned representations + scale amplification + alignment shaping). Progress within the same dimension is climbing a slope (GPT-2 → GPT-3), while cross-dimension is a shift (statistical → neural → scale + alignment).

III. Key Technology Roadmap ​

Around "model forms in the Transformer era," four routes evolve in parallel. Understanding this branching lets you read the family relationships in case study pages like GPT Series and BERT and the Encoder Family.

1. Comparing the four routes ​

RouteArchitectureTraining objectiveRepresentative modelsStrengths
Decoder-onlyUnidirectional autoregressive TransformerPredict the next wordGPT-1/2/3/4, Llama, Qwen, DeepSeekGeneration; mainstream in the LLM era
Encoder-onlyBidirectional TransformerMasked language modeling (MLM)BERT, RoBERTa, ELECTRAUnderstanding (classification, retrieval, embeddings)
Encoder-DecoderEncoder reads in + decoder generatesText-to-textT5, BART, FLAN-T5Balance of understanding + generation (translation, summarization)
Hybrid / Next-genSparse experts (MoE), state-space models, etc.VariousMixtral, DeepSeek-V3, MambaEfficiency and long sequences (see MoE)

2. Why the decoder won ​

Before 2019, the encoder (BERT) ruled — leading on understanding-task benchmarks. After 2020, the balance reversed, and the decoder became the de facto standard. Three reasons:

  1. Generation subsumes understanding: A model that can generate can naturally do understanding (just make it "answer"), while understanding models struggle to cover generation.
  2. Unified pretraining objective: Next-token prediction needs no labels and has infinite corpora, making it more suitable for scaling laws (see Scaling Laws) than MLM.
  3. Alignment and conversation: RLHF, multi-turn dialogue, and tool calling all naturally build on the decoder's "autoregressive generation + instructions" form.
text
GPT route = decoder architecture + massive pretraining + scale amplification + alignment (RLHF)

3. The encoder's legacy ​

The decoder won today's market, but the encoder route left three legacies that are still actively used:

LegacyUse caseRepresentative
Text vectorization (embeddings)Input features for retrieval, clustering, semantic similarity computationBERT, E5, bge (see the retrieval step in RAG)
Bidirectional understanding benchmarksClassification, extraction, information retrieval — "understanding-first" tasks still commonly use themRoBERTa, ELECTRA
Training methodologyMasked modeling, contrastive learning, and other pretraining techniques inherited by multimodal fieldsVision-language model alignment heads

Reading BERT and the Encoder Family separately reveals: it hasn't "died" — it's just gone from "lead actor" to "infrastructure."

IV. Three Patterns from the Timeline ​

1. Simple objective × massive scale = complex capability ​

The three shifts never changed that "naively simple" training objective — predict the next word. What changed was scale (data and parameters) and architectural efficiency (Transformer's parallelism). History repeatedly proves: in this field, "smarter objectives" are far less effective than "more data + better scaling" — this is the quantitative conclusion that Scaling Laws repeatedly emphasizes.

2. Capabilities emerge indirectly from "surrogate tasks" ​

From Word2Vec (predicting neighbor words but learning semantic structure) to GPT (predicting the next word but learning to reason), a model's "real capabilities" were never directly required by the training task. They are structures the model was forced to learn when compressing data. History reminds us: what an LLM can or cannot do is often only known after scale is reached, and cannot be fully predicted by design alone. This is why the "capability layer" chapter in What Is a Large Language Model exists.

3. Every "breakthrough" relied on an "application form," not just raw capability ​

The neural language model existed in 2003, but the public perception only started with ChatGPT. Similarly, BERT's capabilities existed in 2018, but it took years for developers to use it and for products to consume it. Model capability → product form → public acceptance has a long transmission chain. Looking at history today, don't forget: behind every "sudden breakout" is a longer, less visible accumulation line.

4. Two unresolved debates on the timeline ​

History isn't settled doctrine. Two debates remain open; keep a critical mind when reading the timeline:

DebateTwo sidesWhy it matters
Is emergence a real capability or a measurement artifact?One side argues reasoning and other capabilities "jump" as scale grows (phase transition); the other side argues that with smoother metrics, the jumps disappear, being artifacts of "non-continuous measurement."Determines "whether capability is an inevitable result of scale" and whether we should expect bigger models to automatically be stronger. See Scaling Laws.
Is scale growth sustainable?Optimists believe "more data + bigger models" will keep working; skeptics point out that high-quality text data is peaking (the data wall), and synthetic data plus inference-time scaling may change the rules.Determines where compute and data investment goes. Inference-time scaling, MoE, and synthetic data are the three current alternative paths. See Frontier Progress.

Neither debate has a standard answer, but asking about them is itself the ability to read history well — history provides not conclusions, but a genealogy of questions.

V. How to Keep Tracking This Timeline ​

History doesn't end in 2025. There are three entry points for tracking evolution:

Entry pointContentSite link
Paper mapKey papers laid out as a map by timeline and themePaper Map
Case study pagesDeep-dive breakdowns of each model familyGPT Series, ChatGPT and Conversational Models, Llama and the Open-Source Ecosystem, MoE and Ultra-Large Models, Multimodal LLMs
Frontier trackingTrend summaries and representative papers from 2023–2025Frontier Progress

Suggested reading order

First read What Is a Large Language Model to build concepts, then come back to this page to string the timeline together, then dive deeper along whichever branch interests you (architecture → Transformer Architecture; scale → Scaling Laws; alignment → Alignment). The value of history is not memorizing dates; it's so that when you read any new model, you can answer: which route is it on, and which step did it push forward?

1. From Timeline to First-Hand Literature ​

The timeline is just an index. What's truly worth reading are the papers the index points to. Here's a suggestion for "which paper to read first" by interest area:

Interest areaFirst paperWhy read it first
ArchitectureAttention Is All You NeedThe common foundation of all modern architectures; the starting point of the paper map
ScaleScaling Laws / ChinchillaThe quantitative main thread, explaining "why bigger is better" and the "1:20 ratio"
AlignmentInstructGPTThe origin of the RLHF paradigm; ChatGPT's direct predecessor
Generation paradigmGPT-3Establishing the few-shot paradigm
Open-source ecosystemThe Llama paper and announcementThe starting point of "open-source catching up to closed-source"
Inference-time scalingDeepSeek-R1 technical reportThe newest dimension post-2024; first-hand info on papers and new models

A foundation for reading papers

Reading papers directly can be steep. We recommend first building conceptual foundations in Language Modeling and Transformer, then deep-reading according to the "background → method → experiment → limitations" four-part template in Classic Paper Deep Dives.

Further Reading ​

References ​