Skip to content

Evolutionary History

Quick overview From the 1956 Dartmouth Conference to the large-model era of the 2020s, machine learning has experienced three winters and two paradigm shifts. This article uses a timeline to clarify the rise-and-fall logic of AI/ML/DL over seventy years, and the technical reasons behind each "cooling" and "heating."

Evolutionary History ​

Machine learning didn't explode overnight. Over the past seventy years, it has experienced three waves and two winters—each "cooling" was not bad luck, but a mismatch between the model capacity available at the time and compute/data. Understanding this rise-and-fall arc is far more important than memorizing a list of dates—it helps you see where today's large-model hype fits in the historical picture.

I. A three-act overview ​

text
1950s─1980s   Symbolic AI era: humans write rules, machines reason   → First winter (expert system decline)
1980s─2000s   Statistical learning era: machines learn patterns from data → Second winter (deep learning lacked compute)
2012─present  Deep learning era: compute × data × scaling laws     → Large-model wave (explosion from 2022)

The essence of each shift is "the same problems solved with a new paradigm": rules → feature engineering → representation learning → scaling laws.

II. The pre-deep learning era (1956–2011) ​

1956: Dartmouth Conference, AI is born ​

The term "artificial intelligence" was born at a conference at Dartmouth College in the summer of 1956. Attendees included McCarthy, Minsky, Shannon, and others who would later be called "fathers of AI." The conference proposal stated an ambitious goal that still seems astonishing today: "machines that use language, form abstractions, and solve problems that only humans can solve." This goal has not been fully achieved to this day—keeping that in mind helps you stay calm about extreme narratives of "AI is dead/AGI is here."

1957–1969: The Perceptron and the first wave ​

In 1957, Frank Rosenblatt invented the Perceptron—the first neural network that could "learn from samples," demonstrating learning ability on simulators, and the media began hailing "thinking machines." The wave was cooled by Minsky & Papert's 1969 book Perceptrons, which rigorously proved that a single-layer perceptron couldn't even represent XOR (exclusive OR). It wasn't neural networks that failed—it was "single-layer" that failed—but the conclusion was misread as "neural networks are a dead end," and the first AI winter began. The theory of multi-layer networks (backpropagation) would not be systematically reinvented until 1986 by Rumelhart, Hinton, and Williams.

1960s–1980s: Expert systems and the second wave ​

During the winter period, the mainstream was symbolic AI: encoding expert knowledge into if-then rules to build expert systems. They were commercially hugely successful in the 1980s (e.g., DEC's XCON configuration system saved tens of millions of dollars annually), heralding the second AI wave. But the fatal flaw of expert systems was clear: rules are written by humans and cannot cover the open-ended problems of the real world—the knowledge acquisition bottleneck caused maintenance costs to explode. The expert system bubble burst in the late 1980s, and the second AI winter began.

1980s–2000s: The statistical learning revival ​

Against the backdrop of setbacks for both neural networks and expert systems, statistical learning emerged as the new orthodoxy, laying the foundation for today's "machine learning" discipline:

  • 1986: Backpropagation was systematically formulated, making multi-layer perceptrons trainable;
  • 1995: Vapnik proposed the Support Vector Machine (SVM), which became the king of classic methods with its rigorous statistical learning theory (VC dimension, structural risk minimization);
  • 1997: Freund & Schapire proposed AdaBoost, birthing the boosting idea that would later evolve into XGBoost, LightGBM, and other tabular data powerhouses;
  • 2001: Breiman proposed Random Forest, a combination of bagging + decision trees that proved extremely robust in practice;
  • 2006: Hinton published the deep belief network (DBN) paper, coining the term "deep learning"—but compute limitations prevented it from taking off.

The consensus of this era was "feature engineering sets the ceiling": whoever hand-designed the best features had the best model. This consensus was completely shattered in 2012.

III. The deep learning revolution (2012–2017) ​

2012: AlexNet—the ignition point ​

At the 2012 ImageNet competition (ILSVRC), Alex Krizhevsky's AlexNet (an 8-layer CNN) brought the top-5 error rate down from 26.2% to 15.3%—nearly 10 percentage points below second place. Its three pillars of success: GPU parallel computation (using two GTX 580s), the ReLU activation function, and Dropout regularization + data augmentation. Note: CNNs themselves were not new (LeNet existed since 1998); the 2012 breakthrough was that compute finally caught up with model capacity—the first complete realization of "technological breakthrough = model × compute × data."

2014–2016: Three golden years of deep learning ​

  • 2014: GANs (Goodfellow) proposed, a new paradigm for generative modeling; VGG/GoogLeNet deepened and widened CNNs;
  • 2015: ResNet used residual connections to build networks up to 152 layers, achieving a top-5 error rate of 3.57%—surpassing human-level performance (about 5.1%) for the first time—the "residual" trick solved the problem of training deep networks and became the standard for all subsequent large networks. Also in 2015, the AlphaGo paper was published; in March 2016, AlphaGo defeated Lee Sedol 4:1, igniting "machine learning" in the public for the first time;
  • 2014–2017: Sequence modeling moved from RNN/LSTM to attention mechanisms. In 2014, Bahdanau proposed attention; in 2017, Vaswani et al. published 《Attention Is All You Need》, birthing the Transformer and setting off the NLP revolution from 2018 onward.

The theme of this period can be summarized as: CNNs reigned in images, RNNs struggled in sequences, and attention began to appear.

IV. The pre-training and large-model era (2018–present) ​

2018–2019: The pre-training paradigm is established ​

  • 2018: BERT (340M parameters, Transformer encoder, bidirectional pre-training + fine-tuning) swept 11 NLP benchmarks. BERT established the "pretrain-finetune" paradigm: first pre-train on unlabeled corpora (learn language), then fine-tune on small datasets (learn tasks). GPT-1 (110M parameters, unidirectional Transformer decoder) proposed the "generative pre-training" route the same year.
  • 2019: GPT-2 (1.5B parameters) was delayed due to being "too dangerous," and the media began discussing large models' generative capabilities for the first time; the BERT family (RoBERTa, ALBERT) became the de facto standard for NLP.

2020: Scaling laws are established ​

GPT-3 (175B parameters) published the paper Language Models are Few-Shot Learners, demonstrating a conclusion previously thought to be a hallucination: once models reach a certain size, they can accomplish unseen tasks from prompts alone (few-shot). Kaplan et al.'s concurrent paper Scaling Laws for Neural Language Models empirically established the pattern that "loss decreases as a power law with parameters/data/compute"—"just make it bigger" finally got a mathematical formulation. This principle was later refined by Chinchilla (2022) to say "data and parameters should scale together."

2022–2023: ChatGPT and the conversational AI explosion ​

  • November 2022: OpenAI released ChatGPT (GPT-3.5 + RLHF alignment), which hit one million users in two weeks, becoming the fastest-growing consumer application in history. RLHF (Reinforcement Learning from Human Feedback) enabled models to "learn to converse," and was the technical key to this wave of product explosions;
  • March 2023: GPT-4 was released, reaching the top 10% level on multiple professional exams; since then, various players (Anthropic Claude, Google Gemini, Meta Llama, etc.) have joined the race, and the open-source community (Llama, Mistral, Qwen) has rapidly caught up.

2023–2026: Multimodal and Agent era ​

  • 2023–2024: Multimodal large models (GPT-4V, Gemini, Claude 3) unified text/images/audio; the open-source models caught up to one generation behind closed-source models;
  • 2024–2025: AI Agent wave—large models moved from "chatting" to "doing work": agent products like Claude Code, Cursor, OpenHands exploded; models wrapped in a harness (tools, loops, permissions) became the protagonists; RAG, fine-tuning, and distillation became the three essentials for enterprise adoption;
  • 2025–2026: Reasoning models (o1/o3, DeepSeek-R1, Claude, etc.) pushed reasoning ability to new levels using "chain of thought + reinforcement learning"; the industry shifted from "competing on parameters" to "competing on reasoning efficiency and engineering."

V. Two threads that run through it all ​

Looking back at seventy years, all the ups and downs can be traced to two threads:

Thread 1: Where does knowledge come from? Rules (written by humans) → features (designed by humans) → representations (learned by models) → scaling (stacking data + compute). Every shift pushed "the human" one step further. Today's LLM "in-context learning" is already at the extreme end of this line—where "even task definitions are provided by data."

Thread 2: Where is the bottleneck? Every winter is an imbalance in the same formula:

Model capability = algorithm × compute × data

The perceptron died from insufficient compute and layers, expert systems from incomplete data coverage, and 2000s-era deep learning from insufficient compute. Large models of the 2020s are the result of all three becoming abundant simultaneously. So to judge whether a technology "will take off," look at the shortest board in the trio, not the longest.

Three reminders from history

  1. Hype ≠ maturity: The 1960s and 1980s also had "AI can do everything" media narratives; after the bubbles burst, only real capability remained;
  2. Paradigms don't simply repeat: Today's large models share the same "data/knowledge bottleneck" anxiety as the expert systems of old, but scaling laws and RLHF provide paths that are unprecedented in history;
  3. Technology history is gradual: Attention in Transformers, residuals in ResNet, backpropagation—every "miracle" today has its parts in history. Reading papers with this timeline in mind will make things much clearer.

Further reading ​

References ​