Theme
The GPT Series: From GPT-1 to GPT-4o
GPT (Generative Pre-trained Transformer) is a family of decoder-only Transformer-based generative language models iterated by OpenAI since 2018. Through the three-step approach of "pretraining + fine-tuning/alignment + scaling," GPT has turned "next-token prediction" from an academic technique into the foundation of general AI. This page breaks down the full evolution from GPT-1 through GPT-4o by generation; for the underlying knowledge of architecture and training mechanisms, see Transformer Architecture Explained and Language Modeling: The Next-Token Prediction Paradigm; for the broader historical context, see Evolution Overview.
I. The GPT Trajectory: A One-Sentence Summary
Each iteration of GPT follows the same core thread — decoder architecture + larger scale + better alignment with human intent:
Trajectory: Transformer Decoder
+ Massive text autoregressive pretraining (next-token prediction)
+ Scaling (parameter, data, compute growth)
+ Alignment with human preferences (RLHF / instruction tuning)
+ Modality expansion (image / audio / video)This trajectory diverged from the path represented by BERT — "encoder + bidirectional + understanding" (see BERT and the Encoder Family for comparison) — and became the dominant paradigm for large models in the 2020s. Understanding the logic of the GPT series is, in many ways, understanding half the history of large models.
| Dimension | GPT Path Choice | Contrasting Path |
|---|---|---|
| Architecture | Decoder-only, autoregressive | Encoder (BERT, bidirectional), encoder-decoder (T5) |
| Training objective | Next-token prediction | Masked language modeling (MLM) |
| Source of capability | Scale + data + alignment | Deep bidirectional context |
| Output form | Generative (continuation, dialogue, code) | Understanding (classification, extraction, retrieval embeddings) |
| Key milestones | GPT-3 / ChatGPT / GPT-4o | BERT / T5 / retrieval models |
Now let's break them down by timeline.
II. First Generation: GPT-1 (2018) — The Paradigm of "Generative Pretraining + Fine-tuning"
1. Background and Paper
In 2018, the NLP mainstream was "train a separate model for each task." Context-aware word embeddings (like ELMo, Word2Vec) had just started gaining traction. OpenAI's Radford et al. published Improving Language Understanding by Generative Pre-Training, proposing using language modeling as a pretraining task, then lightly fine-tuning for downstream tasks — this is the origin of the name "Generative Pre-training" (GPT). The paper was not posted to arXiv officially; it was released as an OpenAI technical report on their website.
2. Specs and Key Innovations
| Item | Specification |
|---|---|
| Release date | June 2018 |
| Parameters | ~117 million (12-layer decoder, 768-dim) |
| Training data | BookCorpus (~7,000 unpublished books, ~5 GB) |
| Pretraining objective | Autoregressive "next-token prediction" |
| Fine-tuning approach | Task-adaptive input transformation + supervised fine-tuning |
| Vocabulary | ~40K (BPE) |
Three key innovations:
- Two-stage paradigm: Unsupervised pretraining (language modeling on books corpus) → supervised fine-tuning (classification / reasoning / QA on downstream tasks). Pretraining learns general language knowledge; fine-tuning only learns task format.
- Task input transformation (Traversal-Style): Unifies classification, textual entailment, similarity, and QA into "text sequences with separators," all handled by the same decoder output, avoiding architecture changes per task.
- Transformer decoder: Ditched RNNs, used self-attention + causal masking, enabling highly parallel training.
3. Results and Significance
GPT-1 outperformed the SOTA using ELMo + manual feature engineering on 9 of 12 tasks (including an average improvement of about 4+ points on GLUE). Its significance was establishing three things: decoder-only is viable, language modeling is a universal pretraining task, and fine-tuning is a cheap task adaptation method — all three of which were later amplified to the extreme by the GPT series.
Why decoder-only?
At the time, the mainstream academic view was that "bidirectional encoders (like BERT) are needed for understanding, while decoders can only generate." GPT-1 proved that with enough pretraining data, a unidirectional decoder holds its own on understanding tasks — and is naturally suited for generation. This laid the groundwork for the later "one model solves everything" vision.
III. Second Generation: GPT-2 (2019) — Zero-Shot and "Dangerous Speech"
1. Background and Paper
In February 2019, OpenAI released Language Models Are Unsupervised Multitask Learners with an ambitious claim: language models are themselves multitask learners (unsupervised multitask learners) — as long as pretraining is good enough, you don't need task-specific fine-tuning; just provide a natural-language prompt and the model can do translation, summarization, QA.
2. Specs and Key Innovations
| Item | Specification |
|---|---|
| Release date | February 2019 (small models released first due to "safety" concerns; full weights released about 8 months later) |
| Parameters | 1.5 billion (48 layers, 1600-dim, vocab size 50,257) |
| Training data | WebText: ~8 million highly upvoted Reddit link pages (~40 GB of text) |
| Pretraining objective | Same as GPT-1, but purely autoregressive, no fine-tuning (zero-shot) |
| Tokenization | Byte-level BPE (byte-level splitting to handle OOV words and multilingual text), vocab fixed at 50,257 |
Key innovations:
- Zero-shot multitask: No fine-tuning, no few-shot — just "prompt + separator" describing the task, and the model executes directly (e.g., for translation:
"English: ... French: ..."). - Byte-level BPE: Dropped tokenization granularity to the byte level, allowing any text (emoji, code, low-resource languages) to be split, with a fixed vocab of 50,257.
- Data quality over everything: WebText filtered for high-quality web pages via Reddit upvoted links, elevating "data quality" to the same importance as model size — a principle elaborated in Pretraining: Data and Objectives.
3. Phased Release and "AI Safety" Controversy
OpenAI initially released only small versions, citing that the model was "too powerful and could be misused (e.g., generating fake news)," releasing full weights only 8 months later. This sparked enormous debate: "Is capability itself dangerous?" vs. "The consequence of not releasing is the community cannot study alignment." Looking back today, GPT-2's "danger" was more marketing posture and self-restraint, but it was the first time the potential risks of large models and release strategies were put on the table — a theme that evolved into the entire Safety and Risk chapter.
4. Significance
GPT-2 proved that without fine-tuning, relying only on scale and data, models can still handle multitask work; the idea that "the prompt itself is part of the input" took root. It performed comparably to specialized models on benchmarks like LAMBADA, WikiText, and reading comprehension, paving the way for GPT-3's few-shot capabilities.
IV. Third Generation: GPT-3 (2020) — Scale Is Capability
1. Background and Paper
In May 2020, Brown et al. published Language Models Are Few-Shot Learners, jumping parameters from 1.5B to 175B and formally introducing In-Context Learning: without updating any parameters, by simply providing a few examples (few-shot), the model can complete new tasks.
2. Specifications
| Item | Specification |
|---|---|
| Release date | May 2020 (paper); API gradually opened the same year |
| Parameters | 175 billion (96 layers, 12,288-dim), with 8 versions of different sizes released simultaneously |
| Training data | Common Crawl (filtered), WebText2, Books1/2, Wikipedia — ~300 billion tokens (~45 TB of text) |
| Training compute | ~3,640 PFLOPS-days |
| Vocabulary | 50,257 (byte-level BPE) |
3. Key Innovations
- Few-shot in-context learning: No fine-tuning needed for a task — just put examples in the input:
Input: "Translate to Chinese: Apple → 苹果; Orange → 橙子; Banana →"
Output: "香蕉"- Empirical scaling laws: GPT-3 used the same architecture, same data, only scaled up — capability improved continuously with parameters, perfectly matching the power-law curves in the Scaling Laws page, and also validating the existence of "emergent abilities."
- Capability inflection point: The 175B model approached the fine-tuned SOTA of the time on arithmetic, code, reading comprehension, and translation; it scored 71.8% on SUPERGLUE (not yet human-level), but the zero-shot/few-shot "generality" was proven to the public for the first time: "large models can work without training."
4. Limitations and Lessons
GPT-3 also exposed important problems: confidently fabricating facts (hallucinations), toxic biases, unstable mathematical reasoning, and output misaligned with human intent (ask it to summarize, and it might continue the story instead). The answer to all of these points in one direction: with pretraining alone, the model has "capability" but not "intent alignment" — which is exactly what the next phase, InstructGPT, would address.
"Capability ≠ Alignment"
GPT-3's lesson: the bigger the model, the more confidently it's wrong. Scale only solves "can it," not "should it"; alignment (alignment) is the critical step that directs capability toward human intent.
V. The Alignment Era: InstructGPT and ChatGPT (2022)
1. InstructGPT: The RLHF Three-Step Pipeline
In March 2022, Ouyang et al. published Training language models to follow instructions with human feedback, proposing the RLHF (Reinforcement Learning from Human Feedback) three-step pipeline (see Alignment: RLHF and DPO for details):
① Collect human-written instruction-response pairs → SFT fine-tuning for initial assistant
② Have humans rank multiple responses → Train a reward model (RM) to learn "which is better"
③ Use PPO reinforcement learning to maximize the reward, with KL penalty to prevent divergenceThe key finding: InstructGPT at just 1.3B parameters outperformed GPT-3 at 175B on human evaluation — alignment made a "small model" more likable than a "big but wild" one. This "alignment lever" was later proven just as critical as the "scale lever."
2. ChatGPT: The Productization Inflection Point
On November 30, 2022, OpenAI released ChatGPT (based on the GPT-3.5 series, applying RLHF further in dialogue scenarios, with conversational data and safety guardrails). It hit 1 million users in 5 days and 100 million monthly active users in two months — the fastest-growing consumer app in history. ChatGPT's story, tech stack, and product impact get their own page: ChatGPT and Conversational Models.
3. Significance of the Alignment Era
InstructGPT/ChatGPT proved: with the same foundation, with or without alignment determines whether it's a "completion machine" or a "general assistant." After this, "pretraining + SFT + RLHF/DPO" became the standard post-training pipeline for all major large models (Llama, Qwen, Claude, Gemini, etc.). See Fine-tuning: SFT and Parameter-Efficient Fine-tuning.
One commonly discussed cost to clarify here: the "alignment tax" — models may score slightly lower on standard benchmarks due to safety and compliance constraints. This was pronounced in the InstructGPT era, and later research (DPO, Safe RLHF, etc.) worked to reduce it (see Alignment: RLHF and DPO). Understanding the tradeoff — "accept a small cost for usability" — explains why all commercial models are willing to pay this tax: capability without alignment cannot be delivered to real users.
VI. Fourth Generation: GPT-4 (2023) — Multimodal and the "Exam Monster"
1. Release and Technical Report
On March 14, 2023, OpenAI released GPT-4, accompanied by the GPT-4 Technical Report (on arXiv). The report deliberately did not disclose parameters, training data, or compute (industry speculation pointed to a MoE architecture with ~1.8T total parameters, unconfirmed by OpenAI; refer to official releases). Instead, it used extensive capability benchmarks and safety evaluations to demonstrate model quality.
2. Key Capabilities
| Dimension | GPT-4 Performance |
|---|---|
| Multimodal input | Accepts image + text (reasoning over charts, screenshots, handwritten content) |
| Exam level | Top 10% on Uniform Bar Exam, 700+ on SAT Math, perfect on multiple AP exams (approaching human level) |
| Standard benchmarks | ~86% on MMLU 5-shot (compared to ~70% for GPT-3.5) |
| Context window | 8K at release, later extended to 32K and 128K |
| Reasoning stability | Significant improvements over GPT-3.5 in factuality, instruction following, and jailbreak resistance |
| Parameter scale | Not disclosed (speculated MoE, ~1.8T total parameters) |
3. Lessons
- Multimodal is the next input modality: GPT-4V (the vision version) made "reading diagrams" real. See Multimodal LLMs.
- Benchmarks shift toward "human tasks": Using real exams like the bar exam, SAT, and MMLU replaced traditional NLP benchmarks — spawning the "human exam-style benchmark" wave discussed in Evaluation and Benchmarks.
- MoE speculation: If GPT-4 indeed uses a sparse expert model, it means "activated parameters << total parameters" has become the internal choice for flagship models. See MoE and Ultra-Large-Scale Models.
- Safety evaluation front-loaded: The technical report devoted significant effort to red-teaming and safety guardrails, moving alignment from "post-hoc patch" to "built-in training."
VII. Fifth Generation: GPT-4o (2024) — Native Multimodal and Real-Time Interaction
1. Release
On May 13, 2024, OpenAI released GPT-4o ("o" = omni, fully multimodal), accompanied by a System Card. It unified text, image, audio, and video into the same model (native multimodal) and made it available to all users for free.
2. Key Innovations
| Item | Specification |
|---|---|
| Release date | May 2024 |
| Input modalities | Text / image / audio / video (unified token space) |
| Output modalities | Text / image (DALL·E integration) / audio (TTS) |
| Real-time capability | Average voice response ~320ms, approaching human conversation rhythm |
| Interaction paradigm | Voice conversations, interruptions, emotional expression, real-time visual understanding |
| Cost | Significantly lower API pricing than GPT-4, plus free access for all users |
3. Significance
GPT-4o moved large models from a "typing interface" to a "conversation interface": real-time voice, visual understanding, and emotion recognition all happen within the same model, pushing latency from seconds to sub-second. It also validated two paths: native multimodal (no longer needing CLIP stitching) and "capability democratization" (top-tier models freely available). This also further evolved the product form of ChatGPT and Conversational Models.
4. The Continuation (Post-2024 H2)
After GPT-4o, OpenAI's lineage branched into two threads: the o series (o1/o3, test-time scaling from September 2024, using RL to teach the model "slow thinking," see Frontier Progress) and the GPT-5 series (released August 2025, unifying reasoning and multimodal). Details for these threads are subject to official releases; this page focuses through GPT-4o.
VIII. Technical Commonalities: Decoders, Context, and Engineering Details
The previous sections covered "what happened" by generation. This section fills in the engineering commonalities across generations — the underlying reasons GPT has been able to iterate so reliably.
1. Four Engineering Pillars of the Decoder Architecture
| Pillar | Role | Relation to GPT Series |
|---|---|---|
| Causal mask | Ensures autoregressivity (each position only sees the left side) | Pretraining and inference behavior are consistent; naturally suited for generation |
| KV Cache | Caches historical attention key-values, avoiding recomputation during generation | Determines generation speed and concurrency costs (see Inference Fundamentals) |
| Position encoding | Encodes token order | GPT-2/3 used absolute position embeddings; later versions generally adopted RoPE-like extrapolation approaches |
| Layer normalization + residual | Stabilizes deep network training | Shifted from Post-Norm to Pre-Norm, supporting ultra-deep network scaling |
2. Evolution of the Context Window
| Model | Context Window |
|---|---|
| GPT-2 | 1,024 |
| GPT-3 | 2,048 |
| GPT-3.5 / text-davinci-003 | 4,096 |
| GPT-4 | 8K → 32K → 128K |
| GPT-4o | 128K |
| o1 / o3 | 128K–200K+ (subject to official releases) |
The driving force behind the context window expansion each generation is not an architectural revolution, but the cumulative advances in position encoding extrapolation and attention engineering (FlashAttention, sparse attention). See Context and Long Context for details.
Attention's computational cost grows quadratically with sequence length: expanding from 2K to 128K increases the attention computation for a single forward pass by roughly 4,096× — which is precisely why long-context models must rely on FlashAttention-level memory optimization, sparse attention, and KV cache management. Every GPT version has re-negotiated the tradeoff between "context length" and "inference cost": GPT-4 started at 8K and gradually opened 32K/128K — first ensuring inference economics with shorter windows, then relaxing with engineering improvements. For users, long context is not free: feeding hundreds of pages of documents into the context is far more expensive than using RAG to bring only relevant snippets. This is precisely why RAG: Retrieval-Augmented Generation continues to thrive even in the long-context era.
3. API and Model Codenames: How the Product Layer Evolved
| Time | Model Codename | Notes |
|---|---|---|
| 2020–2021 | davinci / text-davinci-001 | GPT-3 base API |
| 2022 | text-davinci-002 / 003 | Aligned instruction models (InstructGPT productized) |
| 2023.3 | gpt-3.5-turbo, gpt-4 | Dialogue-optimized price revolution and flagship capability |
| 2023.11 | gpt-4-turbo (128K) | Long context + massive price reduction |
| 2024.5 | gpt-4o | Fully multimodal + free access |
4. Three-Stage Post-Training Pipeline Specs
The GPT series' post-training pipeline has standardized into three stages: pretraining → SFT → RLHF. The input, output, and major cost for each stage:
| Stage | Input | Output | Major Cost |
|---|---|---|---|
| Pretraining | Trillions of tokens of web/books/code | Base model (can continue text) | Hundreds of millions to billions of dollars in GPU compute |
| SFT | Tens of thousands to hundreds of thousands of instruction-response pairs | Instruction-following model | Single machine to small cluster |
| RLHF | Human preference rankings + reward model | Aligned assistant | Moderate compute + large-scale human annotation |
5. Three Constants and Three Changes
Three constants:
- Autoregressive objective unchanged: From GPT-1 to GPT-4o, the core goal has always been "predict the next token." Every architectural adjustment serves to make this goal train better and work better.
- Scale lever unchanged: Parameters, data, and compute remain the primary drivers of capability (see Scaling Laws).
- Two-stage paradigm unchanged: The skeleton of "pretraining + post-training (fine-tuning/alignment)" has never changed.
Three changes:
- Alignment went from "optional" to "mandatory": After InstructGPT, a base model without alignment is virtually impossible to ship as a product.
- Modalities went from "text only" to "fully multimodal": GPT-4V/4o made input no longer restricted to text.
- Capability went from "single-turn" to "multi-step": The o-series' test-time reasoning expansion moved models from "fast answers" to "slow thinking," making Agents and tool use the new battlefield (see LLM-Based Agents).
6. Data-to-Parameter Ratio Lessons
During GPT-3 training, the intuition of "more parameters is better" led to using ~300 billion tokens to support 175 billion parameters. The 2022 Chinchilla paper showed that for a fixed compute budget, parameters and data should roughly follow a 1:20 ratio — meaning 175 billion parameters would need about 3.5 trillion tokens to reach their full potential, and GPT-3 was "data-undertrained." This meant it hadn't pushed its parameters to the limit. Subsequent models generally improved from hundreds of billions to trillions of training tokens, correcting this lesson. The direct implication for practitioners: before adding parameters, first confirm whether your data scale matches them. The same logic applies to small teams doing fine-tuning — rather than blindly increasing model size, first ensure data quality and coverage. For more scaling details, see Scaling Laws.
This lesson also explains the shift in post-training focus across the GPT series: as the marginal return on pretraining data diminished, alignment and instruction tuning became a cheaper "capability lever" — InstructGPT's 1.3B model beating GPT-3's 175B was essentially shifting limited compute from "stacking parameters" to "shaping behavior." How parameters, data, and alignment combine is the cost equation every large model project must answer.
IX. Safety, Evaluation, and Ecosystem Impact
1. GPT Series and Evaluation Benchmarks
Every GPT generation reshapes the evaluation landscape: GPT-1/2 era used traditional benchmarks like GLUE; GPT-3 spawned the "few-shot evaluation" concept; after InstructGPT, human evaluation and "preference rates" became core metrics; GPT-4 pushed real exams like the bar exam, SAT, and MMLU to the forefront, triggering an "exam-style benchmark" wave (methodology in Evaluation and Benchmarks). The cat-and-mouse game between evaluation and models (benchmark contamination, data leakage) also became a public issue. See Datasets and Benchmarks Archive for the full evaluation landscape.
2. Safety and Controversy Timeline
| Event | Time | Controversy Focus |
|---|---|---|
| GPT-2 phased release | 2019 | "Capability is dangerous" vs. "Non-disclosure hinders alignment research" |
| GPT-3 bias and hallucination reports | 2020 | Social bias and incorrect facts in large models |
| ChatGPT jailbreak wave | 2022–2023 | Prompt injection, DAN-style jailbreaks, content misuse |
| GPT-4 red-teaming | 2023 | Whether front-loaded safety evaluation is sufficient |
| Training data copyright lawsuits | 2023 onward | Large-scale scraping and copyright compliance (subject to judicial outcomes) |
For systematic treatment of this topic, see Safety and Risk and Hallucination: Causes and Mitigation.
3. Paradigm Shifts for the Industry
At most times, GPT models are not isolated champions but "industry capability benchmarks": each generation triggers a chain reaction of benchmark upgrades, training paradigms, product forms, and open-source ecosystems — the 2023 multimodal wave, 2024 MoE and reasoning models, 2025 Agent-ification, almost all of which have echoes in GPT version iterations. For the full coordinate system, see Model Compendium and Paper Map.
The one image that summarizes the GPT series
Pretraining gives "knowledge," SFT gives "format," RLHF gives "willingness" — every GPT generation tweaks the balance of these three, plus a "scale knob" (parameters and data) and a "modality knob" (input/output range).
X. Practical Perspective: How to Choose and Use GPT-Style Models
After covering history and mechanisms, let's get practical: what are your paths for using GPT-style models today, and what are the costs and pitfalls?
1. Usage Method Decision Table
| Method | Approach | Cost | Suitable For |
|---|---|---|---|
| Direct API call | Prompt engineering + streaming output | Pay-per-token | Rapid validation, general assistant |
| RAG | External knowledge retrieval + generation | Adds retrieval/storage costs | Knowledge bases, private data |
| Fine-tuning (LoRA) | Customize style/format/domain behavior | One-time training | Behavior-level problems (see Fine-tuning: SFT and PEFT) |
| Distillation | Train small model on GPT answers | One-time | High concurrency, low cost |
| Open-source alternative | Switch to Llama / Qwen / DeepSeek | Compute investment | Data-sensitive, long-term cost reduction |
2. A Token Cost Ledger for One Conversation
Cost per QA ≈ (input tokens + output tokens) × unit price
Input: system prompt + conversation history + user message + retrieved snippets
Output: actual tokens generated within max_tokens limitThree cost-control tactics: ① Trim redundant history (see Context and Long Context); ② Use small models for simple tasks (gpt-4o-mini-type); ③ Set max_tokens and budget circuit breakers.
3. When to Choose Open-Source vs. GPT
| Decision Dimension | Choose GPT Family | Choose Open-Source |
|---|---|---|
| Capability ceiling | Strongest reasoning/multimodal | Slightly behind but sufficient |
| Data security | Rely on vendor | Fully privatized |
| Cost curve | Pay-as-you-go | One-time investment |
| Ecosystem/tools | Official API ecosystem | Hugging Face community |
| Compliance | Requires data cross-border review | Requires self-check of weight compliance |
4. Common Misconceptions and Anti-Patterns
| Misconception | Correct Approach |
|---|---|
| Using GPT as a database for facts | Use RAG/search + ask for citations |
| Using fine-tuning to teach new knowledge | That's RAG's job |
| Obsessing over latest codenames | Evaluate by task, not name |
| Thinking temperature tuning fixes bad output | Fix prompts and examples first |
| Blindly stuffing the context window | Trim, summarize, route |
There is no "best" model, only the "most suitable" one
The GPT series is a capability benchmark, not a universal answer. The right path for model selection is: define the task → build an evaluation set → compare multiple models → score by cost/safety/latency. Evaluation methods in Evaluations in Practice.
XI. Overview Table and Route Summary
1. One-Per-Generation Summary Table
| Codename | Time | Parameters | Key Data | Key Innovation | One-Sentence Significance |
|---|---|---|---|---|---|
| GPT-1 | 2018.6 | ~117M | BookCorpus (~5 GB) | Generative pretraining + fine-tuning | Established decoder-only and two-stage paradigm |
| GPT-2 | 2019.2 | 1.5B | WebText (~40 GB) | Zero-shot multitask, byte-level BPE | Proved "language models are multitask learners" |
| GPT-3 | 2020.5 | 175B | ~300B tokens | Few-shot in-context learning, scale is capability | Large model concept established, emergent abilities went mainstream |
| InstructGPT | 2022.3 (paper) | 1.3B (main experiment) | 13K instructions + 33K preference pairs | RLHF three-step (SFT→RM→PPO) | Alignment paradigm established, small beats big |
| ChatGPT | 2022.11 | GPT-3.5 series | Dialogue data + RLHF | Dialogue productization + safety guardrails | LLMs went from technology to mass product |
| GPT-4 | 2023.3 | Not disclosed (speculated MoE ~1.8T) | Not disclosed | Multimodal input, exam-level capability | Capability leap + multimodal + safety front-loaded |
| GPT-4o | 2024.5 | Not disclosed | Not disclosed | Native multimodal, real-time voice | Interaction paradigm shifted from typing to conversation |
2. Five Milestone Lessons
- Same architecture, scale for capability (GPT-1 → GPT-3): Same decoder + next-token prediction, parameters ×1,500 → emergent translation, code, and reasoning capabilities — scaling laws are the most predictable lever.
- Alignment is the second lever (GPT-3 → InstructGPT): A 13B aligned model beats an unaligned 175B model, proving "willingness" matters as much as "capability."
- Product is the third lever (InstructGPT → ChatGPT): Technology arrives first, then productization (conversational interface + free + guardrails) ignites the mass market.
- Modality expansion is the fourth step (GPT-4 → GPT-4o): From pure text to "native fully multimodal," input/output freedom keeps expanding.
- Every upgrade amplifies the same recipe: Decoder + scale + data quality + alignment + modality — these five pieces together define the "GPT path."
One sentence to remember the GPT series
GPT path = decoder-only architecture + scale and data + RLHF alignment + modality expansion; each step does one thing well, and through "stacking" achieves paradigm-level leaps.
XII. FAQ Quick Answers
| Question | Quick Answer |
|---|---|
| How many parameters did GPT-1 have? | ~117M, 12-layer decoder |
| Why is GPT-2's vocab 50,257? | Fixed vocab from byte-level BPE, covering arbitrary text |
| How large was GPT-3's training data? | ~300B tokens (Common Crawl, Books, WebText2, Wikipedia) |
| Difference between ChatGPT and InstructGPT? | Same RLHF technology; ChatGPT applied SFT/preference training in dialogue scenarios and productized it |
| Why doesn't GPT-4 disclose parameter count? | OpenAI keeps it secret for competitive and safety reasons (subject to official release) |
| What does the "o" in GPT-4o mean? | Omni (fully multimodal) — one model handling text/image/audio/video |
| How do o1 and GPT-4o complement each other? | o1 emphasizes reasoning (slow thinking), 4o emphasizes real-time multimodal (fast interaction) |
| How large is the context window? | GPT-4 starts at 8K/32K/128K, 4o at 128K, o series longer (subject to official) |
| Is ChatGPT free? | Basic version is free; Plus/Team unlock stronger models and features for a fee |
| Will GPT replace BERT? | For generation tasks, yes; for retrieval, classification, and distillation, encoders are still widely used |
| What preparation is needed for using GPT? | Prompt engineering + evaluation set + cost budget, then consider RAG/fine-tuning |
| Will it get dumber the more you use it? | The model itself doesn't change; the product layer can improve experience via RAG/feedback loops |
Note: OpenAI's naming (GPT-3.5, GPT-4, GPT-4o, o1, GPT-5) does not always increment strictly by generation — there are also many sub-versions within the same generation (turbo, mini, omni, etc.). Always refer to the official API documentation for model codenames and specs. This naming style also reminds us: evaluate models through actual testing, not by codenames — the same codename may point to different weights at different times, and historical records and evaluation baselines must be updated with versions.
What to do when a new GPT version appears
First, check the official release notes and System Card, then do three things: ① Run your own golden set comparison against the previous version; ② Confirm changes to context/pricing/rate limits; ③ Assess safety and compliance differences. During version transitions, the evaluation set is the only trustworthy ruler.
XIII. Further Reading
- Evolution Overview — placing GPT in the full timeline from 1940s–2020s
- Transformer Architecture Explained — the self-attention mechanism GPT is built on
- Alignment: RLHF and DPO — the alignment technology behind InstructGPT onward
- ChatGPT and Conversational Models — the full productization story of GPT products
- Multimodal LLMs — the multimodal path of GPT-4V/4o
- MoE and Ultra-Large-Scale Models — the sparse expert technology behind GPT-4's speculated architecture
- Core Paper Deep Dives — paper-by-paper deep dives on GPT-1/2/3 and InstructGPT
References
- Radford et al. Improving Language Understanding by Generative Pre-Training (GPT-1, 2018) — GPT-1 original paper (OpenAI official PDF)
- Radford et al. Language Models Are Unsupervised Multitask Learners (GPT-2, 2019) — GPT-2 paper and release notes
- Brown et al. Language Models Are Few-Shot Learners (GPT-3, 2020) — GPT-3 paper (arXiv)
- Ouyang et al. Training language models to follow instructions with human feedback (InstructGPT, 2022) — RLHF original paper (arXiv)
- OpenAI. Introducing ChatGPT (2022.11.30) — ChatGPT official release blog
- OpenAI. GPT-4 Technical Report (2023) — GPT-4 technical report (arXiv)
- OpenAI. Hello GPT-4o (2024.5.13) — GPT-4o official release blog