Appearance
A Brief History of the AI Boom
One-sentence definition: The history of artificial intelligence is the history of grand capability narratives being repeatedly audited against resource realities — every boom began with the abilities of some technical approach being overestimated, and every winter arrived because a mismatch between model capacity and data, compute, and cost went ignored. When Turing asked in 1950 whether machines can think, no one could have foreseen that by 2025 models would read images, write code, do mathematical reasoning, and even call tools on their own; but no one could have foreseen, either, that over those seventy-five years AI would twice fall into "winters" so deep that academia and capital abandoned it almost simultaneously.
Understanding this evolutionary line matters far more than memorizing a string of dates. It lets you answer three questions: where in history does today's LLM boom sit? Which of today's "disruptions" have in fact happened before? And — when everyone is saying "AI can do anything," which shortcoming should you be watching? This article is the timeline companion to the site's panoramic overview; for the boundaries between concepts see AI vs ML vs DL vs GenAI vs Agent, and for the complete skeleton of a modern AI system see Anatomy of a Modern AI System.
1. The Opening: A Framework of Three Booms and Two Winters
The mainstream historical narrative summarizes this journey as three booms and two winters:
text
First boom 1950s–1960s Symbolism · Perceptron: hand-written rules / neuron simulation → 1970s, first winter
Second boom 1980s Expert systems · Knowledge engineering: if-then rules commercialized → late 1980s, second winter
Third boom 2012–present Deep learning → LLMs → Agents: data × compute × scaling laws (no global "winter" yet)| Stage | Representative technology | Core assumption | Why it burst | What it left behind |
|---|---|---|---|---|
| 1950s–1960s symbolic wave | Perceptron, logical reasoning, early NLP | "Thinking = symbol manipulation"; write the right rules and you get intelligence | Rules can never be complete; single-layer perceptrons have hard expressive limits | The mathematical foundations of search, logic, and knowledge representation |
| 1970s, first winter | — | — | The Lighthill Report and DARPA funding cuts | The lesson that "AI must deliver on its engineering promises" |
| 1980s expert-systems wave | XCON, MYCIN, knowledge engineering | "Knowledge = a rule base"; stack enough rules to cover a domain | The knowledge-acquisition bottleneck; maintenance costs exploding | The intellectual origin of knowledge graphs / ontologies |
| 2012–2015 deep learning revolution | AlexNet, ResNet, GAN | Representations can be learned automatically | (Did not burst — carried into the next stage) | CNNs, representation learning, generative models |
| 2018–2021 pretrained LLM era | BERT, the GPT series | Large-scale pretraining + scaling laws | (Did not burst — accelerating evolution) | The foundation-model paradigm |
| 2022–2025 generative AI wave | ChatGPT, GPT-4, diffusion models, agents | Scale + alignment = a general-purpose assistant | Still unsettled: cost, safety, evaluation | Mass-market AI products and infrastructure |
The key judgment behind this framework: every "winter" was never the total death of the technology, but the falsification of one particular approach's assumptions; and every "spring" arrived when the same old ideas gained new support in compute, data, or algorithms. Looked at through this lens, you will stay skeptical of both extreme narratives at once — "AGI is just around the corner" and "AI is dead."
2. The Prelude and the First Boom (1950–1969)
1950: The Turing Test — the problem gets its first precise formulation
In 1950, Alan Turing published Computing Machinery and Intelligence in the journal Mind, proposing the famous Turing Test: if a machine can converse in a way that leaves a human unable to tell whether it is a machine or a person, it should be considered "able to think." Turing's contribution was not that he defined AI, but that he turned a philosophical question into a workable experiment — and this "judge intelligence by behavior" path remains the spiritual ancestor of LLM evaluation today (modern benchmarks are essentially engineered variants of the Turing Test; see LLM Evaluation and Benchmarks).
1956: The Dartmouth Conference — AI is officially born
In the summer of 1956, John McCarthy and colleagues organized a two-month workshop at Dartmouth College, where the term "Artificial Intelligence" was formally coined. The goal stated in the proposal still looks radical today: "to make machines use language, form abstractions and concepts, solve kinds of problems now reserved for humans." That goal has still not been fully achieved — remember this; it is your anchor for staying clear-headed in the face of every "AI can do anything" marketing pitch.
1957: The Perceptron — the first machine that "learns"
In 1957, Frank Rosenblatt proposed the Perceptron, the first primitive neural network that could learn weights from examples, and demonstrated image classification in a simulator. The media promptly hyped a "thinking machine." In 1969, Minsky and Papert rigorously proved in Perceptrons that a single-layer perceptron cannot even represent XOR. Note: this was not "neural networks don't work" but "single-layer doesn't work" — yet the field misread the result as neural networks being a dead end, directly fueling the first winter.
1966: ELIZA — the earliest case of "chatbot illusion"
In 1966, Joseph Weizenbaum at MIT wrote ELIZA, a program that played the role of a psychotherapist through simple pattern matching. The striking part: many users, fully aware it was a program, still confided in it — history's first demonstration of "emotional projection onto a program," and the first exposure of "people over-interpreting machine intelligence." Today's conversational AI is engineered completely differently, but the psychological mechanism of "users treating fluent text as genuine intelligence" has not changed in seventy years.
Early milestones
| Year | Event | Significance in one line |
|---|---|---|
| 1950 | Turing publishes Computing Machinery and Intelligence | Proposes the Turing Test; the AI question is stated precisely for the first time |
| 1956 | The Dartmouth Conference | The discipline of "artificial intelligence" is born |
| 1957 | Rosenblatt proposes the Perceptron | The first learnable neural-network prototype |
| 1958 | Rosenblatt publishes the perceptron paper | Triggers a media frenzy over "machines that think" |
| 1966 | Weizenbaum releases ELIZA | The first conversational program, and the first case of "over-interpretation" |
| 1969 | Minsky & Papert publish Perceptrons | The single-layer limit is proven; the fuse of the first winter |
3. Two Winters and Two Thaws (1970s–1990s)
The 1970s: the first winter — the "AI hype" gets settled
In the early 1970s, the early optimistic predictions collapsed across the board: machine translation performed far below its promises (the 1966 ALPAC report had already poured cold water), and the perceptron line had been theoretically discredited. In 1973, the British mathematician James Lighthill submitted a report to the UK government stating bluntly that most of AI's promises were "unfulfilled and unlikely to be fulfilled"; from 1974 onward, DARPA made deep cuts to AI research funding. The essence of a winter is a collapse of trust: progress had not stopped — it simply could not match the publicity. The lesson, in one sentence: "hype is no substitute for maturity."
The 1980s: expert systems and knowledge engineering — the second boom
During the winter, the mainstream pivoted to the engineering of symbolism: encoding domain-expert knowledge as if-then rules to build expert systems. In the 1980s, DEC's XCON configuration system, the medical-diagnosis MYCIN, and the geological-analysis PROSPECTOR succeeded one after another and became major commercial wins — the second boom had arrived. Alongside it came the flourishing of knowledge engineering: semantic networks, frames, and ontologies became AI's core toolbox — precisely the academic ancestor of today's knowledge graphs. One clarification: the name "Knowledge Graph" was not coined by Google until 2012, but the approach of "structure human knowledge and hand it to machines" had been established since the knowledge engineering of the 1980s.
1986: Backpropagation revives neural networks — rising from the dead
Just as expert systems were at their peak, in 1986 Rumelhart, Hinton, and Williams published a paper in Nature that systematically presented the backpropagation algorithm, making the training of multi-layer neural networks possible and briefly reviving connectionism. But from the late 1980s into the early 1990s, the second winter arrived in turn: the expert-system knowledge-acquisition bottleneck (rules written by hand, unable to cover open-ended problems, maintenance costs exploding) burst the bubble, and the neural networks of the day, starved of compute, could not deliver on the hype either. Both currents ebbed at once, and by the early 1990s AI had nearly retreated back into a small academic circle.
1997: Deep Blue — a human champion beaten on a narrow task
In May 1997, IBM's Deep Blue defeated world chess champion Garry Kasparov 3.5:2.5, becoming the first machine to beat a human world champion in an intellectual game. Deep Blue had no "deep learning" and no large model; it relied on brute-force search + evaluation functions + specialized hardware — proving that in closed tasks with clear rules and enumerable compute, machines can crush humans. It also set the stage for the later AlphaGo narrative (2016, machine learning + Monte Carlo tree search). For a full breakdown of Deep Blue and AlphaGo, see the site's case study and the "AI for Science" spirit it embodies — applying AI to clearly bounded, measurable tasks.
Winters and thaws: milestones
| Year | Event | Significance in one line |
|---|---|---|
| 1966 | The ALPAC report | Machine-translation promises are settled; early cooling begins |
| 1973 | The Lighthill Report | The UK government cuts funding; the emblem of the first winter |
| 1974 | DARPA cuts AI funding | The US follows; the term "AI winter" is officially born |
| 1970s–1980s | Expert systems rise (XCON, MYCIN, etc.) | The high-water mark of commercialized symbolism |
| 1986 | The backpropagation paper is published | Multi-layer networks become trainable; connectionism revives |
| Late 1980s | The expert-system bubble bursts | The knowledge-acquisition bottleneck triggers the second winter |
| 1997 | Deep Blue defeats Kasparov | The signature victory of narrow-task AI |
4. The Deep Learning and LLM Era (2012–Present)
2012: AlexNet — the detonation point and the "error-rate cliff"
At the 2012 ImageNet Large Scale Visual Recognition Challenge (ILSVRC), Alex Krizhevsky's AlexNet (an 8-layer convolutional network trained on two GTX 580 GPUs) pushed the top-5 error rate down to 15.3% in one stroke, nearly 11 percentage points below the runner-up — previous winning results had hovered around 25%–26%. The "error-rate cliff" became the public evidence of the deep learning revolution: the core consensus in vision flipped completely from "feature engineering sets the ceiling" to "representations can be learned automatically." Note that CNNs themselves were nothing new (LeNet already existed in 1998); the 2012 breakthrough was essentially compute (GPUs) finally catching up with the model — the first complete cash-in of the formula "technical breakthrough = algorithm × compute × data." For a close reading of the original paper, see Classic Papers Deep Dive.
2014: GAN — a new paradigm for generative models
In 2014, Ian Goodfellow proposed Generative Adversarial Networks (GAN): pit a generator against a discriminator until the generator can fool the discriminator. GAN opened the wave of "getting machines to sample new examples from a data distribution" and set the tone for the generative AI that followed. Together with 2015's ResNet (152 layers; a top-5 error rate of 3.57%, surpassing the human baseline of 5.1% for the first time), it formed deep learning's golden two years of 2014–2015.
2017: The Transformer — Attention Is All You Need
In 2017, Vaswani et al. published Attention Is All You Need, introducing the Transformer architecture: it completely abandoned recurrence and convolution, modeling sequences purely with self-attention. Its three great advantages — parallel computation (GPU-friendly), long-range dependency modeling, and scalability — made it the foundation of nearly every large model within a few years. Virtually all of today's "large language models" and "multimodal models" are Transformer variants. For a mechanical breakdown of the architecture, see The Transformer and Attention; for the original paper, see Classic Papers Deep Dive.
2018: BERT and GPT-1 — the pretraining paradigm is established
2018 was NLP's "paradigm year":
- GPT-1 (OpenAI, June, ~117M parameters): validated the "generative pretraining + fine-tuning" approach;
- BERT (Google, October, 340M parameters): swept 11 NLP benchmarks with a bidirectional Transformer encoder and established the "pretrain + fine-tune" paradigm — learn language from unlabeled corpora first, then adapt to tasks on small datasets.
This paradigm later evolved into "pretraining + prompting/alignment": today's models are no longer fine-tuned per task; instead, behavior is shaped on demand through prompting, alignment (RLHF/DPO), and fine-tuning (LoRA). See Large Language Models for details.
2020: GPT-3 and scaling laws — "brute-force scale works miracles" gets a mathematical form
In May 2020, OpenAI released GPT-3 (175B parameters), and the paper's title — Language Models are Few-Shot Learners — called out directly what had previously seemed like a near-fantasy: once a model is big enough, it can complete tasks it was never specifically trained on, using nothing but prompts (few-shot / in-context learning). In the same period, Kaplan et al.'s Scaling Laws for Neural Language Models gave an empirical law that "loss falls as a power law with parameters, data, and compute" — scaling laws turned "brute-force scale works miracles" into a predictable engineering curve for the first time. In 2022, Chinchilla further refined this to "parameters and data must grow in lockstep." This thread remains the theoretical bedrock of the industry's arms race, and the reference frame for judging any large model's "potential and cost" (see LLM Evaluation and Benchmarks).
2021: DALL·E and diffusion models — generative AI's home field shifts
In 2021, OpenAI's DALL·E brought "text-to-image" into the mainstream; in the same period, diffusion models — starting from the 2020 DDPM paper and applied to conditional generation in 2021 through models like GLIDE — rose rapidly, and in 2022 they ignited image generation via DALL·E 2 and the open-source Stable Diffusion. The route battle between diffusion models and GANs thus began (see Diffusion Models and Generative AI and Midjourney and Image Generation). Their significance goes beyond images: the video-generation models of 2024 are also built on diffusion/autoregressive hybrid routes.
November 2022: ChatGPT — the detonation point from technology to the masses
On November 30, 2022, OpenAI released ChatGPT (GPT-3.5 plus RLHF alignment); it passed one million users within two weeks, becoming one of the fastest-growing consumer applications in history. The technical key was not model scale but training the model into a "conversational assistant": pretraining provides capability, alignment provides "obedience." From that moment, AI stepped out of papers and developer communities into everyday life, and with it came systematic demand for prompt engineering. For the complete post-mortem of this case, see ChatGPT and Conversational AI.
2023: GPT-4, multimodality, and Year One of agents
- GPT-4 (March 2023): reached the top 10% of human test-takers on a range of professional exams, pushing "general capability" to a new height;
- Multimodality: GPT-4V, Gemini, and others unified text, images, and audio into a single model — see Multimodal Models;
- Year One of agents: applications like AutoGPT pushed LLMs from "chatting" to "getting work done" — model + tool calls + looping decisions; see AI Agents and Manus and Agent Applications.
2024–2025: open-source catch-up, reasoning models, and video generation
| Time | Event | Significance in one line |
|---|---|---|
| 2024.02 | OpenAI releases Sora, a video-generation model | Text-to-video enters the "physically plausible" stage; see Video Generation |
| 2024.04 | Meta releases Llama 3 | Open models close in on closed-source flagships |
| 2024.09 | OpenAI releases the o1 reasoning model | "Chain-of-thought + reinforcement learning" opens the reasoning paradigm |
| 2024–2025 | The Qwen and DeepSeek series iterate rapidly | China's open-source community becomes a major global force |
| 2025.01 | DeepSeek-R1 goes open source | Reproduces o1-level reasoning at an extremely low cost, stunning the industry; see DeepSeek-R1 and Reasoning Models |
| 2025 | GPT-4.5/GPT-5, Claude 4, Gemini 2.5, and more | The industry shifts from "competing on parameters" to "competing on reasoning, efficiency, and agent capability" |
The key shift in the 2024–2025 industry landscape is closed-source and open-source running on parallel tracks: closed-source giants (OpenAI, Anthropic, Google) keep a capability lead, while the open-source community (Llama, Qwen, DeepSeek) closes the gap to "one generation behind" at low cost. For a quick reference to the latest models and leaderboards, see Models and Leaderboards at a Glance; for the cost and optimization questions of self-hosted deployment, see Deployment and Inference Optimization in Practice.
LLM-era milestones (quick reference)
| Year | Event | Significance |
|---|---|---|
| 2012 | AlexNet ignites deep learning | The error-rate cliff; algorithm × compute × data finally cash in |
| 2014 | GAN proposed | A new paradigm for generative models |
| 2017 | The Transformer paper is published | The architectural starting point of modern LLMs |
| 2018 | BERT and GPT-1 | The "pretrain + fine-tune" paradigm is established |
| 2020 | GPT-3 and scaling laws | "Brute-force scale works miracles" gets a mathematical form |
| 2021 | DALL·E and the rise of diffusion models | Generative AI pivots to images/video |
| 2022.11 | ChatGPT released | AI enters everyday life |
| 2023 | GPT-4, multimodality, Year One of agents | General capability + autonomous action |
| 2024–2025 | Open-source parity, o1/R1 reasoning, Sora video | The industry shifts from competing on parameters to efficiency and engineering |
5. Three Threads That Run Through Everything
Seen across seventy-five years, every rise and fall reduces to three threads. They matter more than any single date.
Thread 1: Capability leaps — perception → language → reasoning → multimodal autonomy
| Stage | Landmark | Problem solved |
|---|---|---|
| Perceptual intelligence | 2012 AlexNet image recognition, 2016 AlphaGo | Seeing: recognition and pattern matching |
| Linguistic intelligence | 2018 BERT, 2020 GPT-3 | Speaking: understanding and generating natural language |
| Reasoning intelligence | 2024–2025 o1, DeepSeek-R1 | Thinking: multi-step logic and mathematical reasoning |
| Multimodal + autonomy | 2023–2025 GPT-4V, Sora, agents | Doing: cross-modal perception + tool calls + completing tasks |
Each leap adds to what came before rather than replacing it: today's agents need perception (reading images and text), language (conversation), and reasoning (planning steps) all online at once.
Thread 2: Architecture evolution — symbolic → statistical → Transformer → diffusion
Symbolism (rules written by people) → Statistical learning (rules learned from data) → Transformer (attention + scaling laws) → Diffusion models (sampling new examples)- Symbolic → statistical: from "people write rules" to "machines learn rules from data" — the first fundamental paradigm shift;
- Statistical → Transformer: from "task-specific models + feature engineering" to "universal pretraining + on-demand adaptation" — the second;
- Transformer → diffusion: generative models branch off on their own, sampling from a distribution via "noise-up, noise-down," and become the mainstream for image/video/audio generation.
Thread 3: The industry landscape — closed-source giants vs. the open-source community
| Dimension | Closed-source camp | Open-source camp |
|---|---|---|
| Representatives | OpenAI, Anthropic, Google DeepMind | Meta Llama, Mistral, Qwen, DeepSeek |
| Strengths | A generation ahead in capability, mature ecosystems, heavy safety investment | Low cost, self-hostable, customizable, transparent |
| Risks | API dependence, data leaving your jurisdiction, vendor lock-in | Safety misuse, maintenance burden, lack of support |
| Historical parallel | The 1980s "vendor knowledge bases" | The free-software movement since the 1990s |
The fact of 2025: the two camps chase each other, with the gap narrowed to "within one generation." For practitioners, choosing a route is not a matter of faith but of finding the optimum across cost, compliance, and capability needs — for a judgment framework, see Learning Paths and Common Pitfalls and Anti-Patterns.
6. Technological Optimism and the Lessons of the Winters
The explosion of capability in the LLM era makes it easy to forget: every "can do anything" claim in history was eventually settled by reality. Today, three weak spots are just as hard to ignore:
| Weak spot | Concrete manifestation | Where to go on this site |
|---|---|---|
| Cost | The compute, electricity, and water consumption of training and inference are enormous; a single frontier-model training run costs tens of millions of dollars | Inference Optimization and Quantization, Deployment and Inference Optimization in Practice |
| Safety | Hallucination, jailbreaking, bias, deepfakes, agents acting beyond their permissions | AI Safety and Governance |
| Evaluation | Benchmark saturation and "leaderboard gaming"; whether capabilities are real and reproducible | LLM Evaluation and Benchmarks |
Three reminders history gives us today
- Hype is not maturity: both the 1960s and the 1980s had "AI can do anything" media narratives; what remained after the bubbles burst was the real substance. To judge whether a technology is production-ready, look at evaluation methods and business metrics, not demo videos.
- Paradigms do not simply repeat: today's LLMs share the "knowledge/data bottleneck" anxiety of the expert-system era, but the paths offered by scaling laws and RLHF are historically unprecedented — so "this time is different" is not entirely wrong either. Listen to both.
- Technical history is incremental: the Transformer's attention, ResNet's residuals, backpropagation, statistical learning — every "miracle" today has parts you can find in history. Read papers and case studies with this timeline in hand, and everything gets clearer.
7. The Timeline: Which Years You Should Actually Remember
You do not need to memorize every line above. There are only 6 milestones worth burning into memory — each represents one "paradigm change":
| Year | Milestone | Why you must remember it |
|---|---|---|
| 1956 | The Dartmouth Conference | The AI discipline is born; its stated goals remain unachieved to this day |
| 2012 | AlexNet ignites deep learning | The first complete cash-in of the formula algorithm × compute × data |
| 2017 | The Transformer paper | The architectural starting point of every LLM today |
| 2020 | GPT-3 + scaling laws | "Brute-force scale works miracles" becomes a predictable engineering law |
| 2022 | ChatGPT released | AI moves from the tech circle into everyday life |
| 2025 | Reasoning models (o1/R1) and open-source parity | The industry shifts from competing on parameters to reasoning efficiency and engineering |
Remember these 6 years and you have grasped the "trunk" of the whole field; everything else can be filled in as needed. For how to turn this historical line into a competitive edge in interviews and on your resume, see Interview Questions and Glossary.
Further Reading
- What Is AI: A Panoramic Overview — the conceptual foundation of this article's conclusions
- AI vs ML vs DL vs GenAI vs Agent — where one concept ends and the next begins
- Anatomy of a Modern AI System — the complete skeleton of a modern AI system
- Learning Paths — three learning routes organized around the historical threads
- The Transformer and Attention — the mainline technology from 2017 to today
- Diffusion Models and Generative AI — the technical roots of image/video generation
- AI Agents — the "getting work done" era since 2023
- ChatGPT and Conversational AI — the complete post-mortem of the mass-market detonation point
- DeepSeek-R1 and Reasoning Models — the flagship case of the 2025 reasoning paradigm
- Classic Papers Deep Dive — the original papers at history's key nodes
- Frontier Progress — the future, viewed from history
- Models and Leaderboards at a Glance — the 2025 model lineup of every camp
References
- Turing. Computing Machinery and Intelligence (Mind, 1950) — the original Turing Test paper
- McCarthy et al. A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence (1955) — the founding document of AI
- Rosenblatt. The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain (1958) — the original perceptron paper
- Weizenbaum. ELIZA—A Computer Program for the Study of Natural Language Communication (1966) — the first conversational program
- Rumelhart, Hinton, Williams. Learning representations by back-propagating errors (Nature, 1986) — backpropagation systematized
- IBM. Deep Blue (official history archive) — the official record of the 1997 victory over Kasparov
- Krizhevsky, Sutskever, Hinton. ImageNet Classification with Deep Convolutional Neural Networks (NeurIPS 2012) — AlexNet, the detonation point of deep learning
- Goodfellow et al. Generative Adversarial Networks (NeurIPS 2014) — the original GAN paper
- Vaswani et al. Attention Is All You Need (NeurIPS 2017) — the Transformer
- Devlin et al. BERT: Pre-training of Deep Bidirectional Transformers (NAACL 2019) — the pretraining paradigm established
- Brown et al. Language Models are Few-Shot Learners (NeurIPS 2020) — GPT-3, 175B parameters
- Kaplan et al. Scaling Laws for Neural Language Models (2020) — scaling laws
- Ho et al. Denoising Diffusion Probabilistic Models (2020) — DDPM, the starting point of diffusion models
- Ramesh et al. Zero-Shot Text-to-Image Generation (DALL·E, 2021) — text-to-image
- OpenAI. Introducing ChatGPT (2022) — the official launch blog post
- OpenAI. GPT-4 Technical Report (2023) — GPT-4
- OpenAI. Introducing OpenAI o1 (2024) — the reasoning-model launch blog post
- DeepSeek-AI. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (2025) — the DeepSeek-R1 paper
- OpenAI. Sora: Creating video from text (2024) — the official video-generation blog post