Theme
ChatGPT and Conversational Models
ChatGPT is OpenAI's generative AI assistant with dialogue as the interaction interface, released on November 30, 2022. It was the first complete public deployment of "alignment technology (RLHF) + dialogue productization." In 5 days it hit 1 million users; in two months, monthly active users exceeded 100 million. ChatGPT turned large models from an academic term into a phenomenon-level product, igniting the entire industry. Its technical foundation (the GPT series) is covered in The GPT Series: From GPT-1 to GPT-4o; its historical context in Evolution Overview.
I. What Is ChatGPT: A One-Sentence Positioning
ChatGPT = InstructGPT-style alignment technology + SFT/RLHF for dialogue scenarios + a carefully crafted product shell. Its model isn't mysterious — based on the GPT-3.5 series (later gradually upgraded to GPT-4 / GPT-4o / o series) — the core was further shaping the "instruction following" capability into a "multi-turn dialogue" capability, paired with safety guardrails, streaming output, conversation memory, and other product features.
User: Explain what a black hole is in one sentence.
ChatGPT: A black hole is a region of spacetime where gravity is so strong that nothing, not even light, can escape. Its boundary is called the event horizon.
User: Would a black hole evaporate?
ChatGPT: Yes. Hawking radiation... (multi-turn context maintained)II. Background: Why Late 2022?
1. Technological Conditions Were Ready
- 2020 GPT-3 proved scale is capability, but "could only continue text, wouldn't listen" (see GPT Series);
- March 2022 InstructGPT paper proved the RLHF three-step pipeline (SFT → reward model → PPO) could make a 1.3B model beat a 175B model, with alignment mechanisms maturing (see Alignment: RLHF and DPO);
- OpenAI internally extended RLHF from "instruction following" to "dialogue," using multi-turn dialogue data for SFT and preference training.
2. Productization Decision
OpenAI released ChatGPT for free as a "research preview" on November 30, 2022, with ~100K human feedback samples for safety refinement. Free access + web dialogue interface + streaming output let it bypass API barriers and reach everyday users directly — a completely different release strategy from the earlier "paper + API" model.
III. The Dialogue System Tech Stack: Four Layers
| Layer | Content | Notes |
|---|---|---|
| Base | GPT-3.5 series (decoder-only pretrained model) | Provides language capability and commonsense |
| Alignment layer | SFT (dialogue data fine-tuning) → reward model → PPO/RLHF | Learns to be "helpful, honest, harmless" (see Alignment: RLHF and DPO) |
| Guardrail layer | Content policy + refusal + system prompts | Handles sensitive content, jailbreak attacks (see Safety and Risk) |
| Product layer | Conversation management, streaming output, multi-turn memory, speed optimization | Determines "feel" and retention |
1. SFT Dialogue Data
The key differentiator for dialogue models is data: using "instruction-response pairs" (InstructGPT used 13K), further covering "multi-turn dialogue trees" — the same question can have multiple reasonable response paths. Annotators structure dialogues as trees, and the model learns "when to probe, when to summarize, when to admit not knowing."
2. Adapting RLHF for Dialogue
Preference data expanded from "which single-turn response is better" to "the overall experience of a multi-turn dialogue": the reward model learns to judge the quality of an entire conversation, and PPO optimization maintains KL distance from the SFT model to prevent output drift.
3. Safety Guardrails
ChatGPT's guardrails are a combination of tactics: during training, safety-related instructions are mixed into SFT and RLHF (the model learns to refuse); during inference, system prompt constraints are layered (e.g., "you are an AI assistant and must not output illegal content"); after release, red-teaming and user feedback continuously patch vulnerabilities. The arms race of guardrails (jailbreaks, prompt injection) is the core topic of Safety and Risk.
Guardrail design also involves a "safety vs. experience" tension: guardrails that are too strict cause frequent refusals, making users feel the model "got dumber"; too loose, and jailbreak/harmful output risk rises. ChatGPT's guardrail practice is "layered" — system prompts provide top-level rules, training uses safety preference pairs to calibrate behavior, and runtime classifiers intercept high-risk requests. The balance isn't about "absolute safety" but "predictable safety": letting users know what they can ask, what will be refused, and giving reasons for refusals.
Why are guardrails important?
Dialogue models are "generalized instruction followers" — they treat every input as an instruction. Without guardrails, models would fluently provide bomb recipes, discriminatory views, or even self-declared fake identities. Guardrails aren't optional — they're the entry ticket for dialogue products.
IV. Why It Broke Through: Capability × Interaction × Timing
| Factor | Specific Manifestation |
|---|---|
| Enough capability | Multi-turn dialogue, writing, coding, translation, summarization work out of the box, and "chats like a person" |
| Interaction paradigm | Natural-language dialogue replaces "search box + click links" — zero learning cost |
| Free access | Bypasses paid APIs; anyone can use it with a browser |
| Product polish | Streaming output (characters appear in seconds), speed, stability, minimal UI |
| Timing window | Late 2022 was on the eve of the AI boom; social media spread ignited it (1M users in 5 days, 100M MAU in 2 months) |
The essence of breaking through: previously, the public's contact with AI was "search engines returning links" or "voice assistants answering fixed questions"; ChatGPT was the first time the public experienced an "agent that converses with you" — an experience perceptible to anyone without any technical background. It turned LLM capability from "tech community consensus" into "social consensus."
V. Competitor Timeline: A Global Race
After ChatGPT's release, the world quickly followed. Below are the major competitors and product events (times subject to official releases):
| Time | Product / Event | Vendor / Key Point |
|---|---|---|
| 2022.11.30 | ChatGPT released | OpenAI, ignited globally |
| 2023.02 | Microsoft New Bing (Bing Chat) | Integrated ChatGPT tech into search engine |
| 2023.03 | Baidu Ernie Bot released | China's first batch of LLM dialogue products |
| 2023.03 | Anthropic released Claude 1 | Focused on safety and constitutional AI |
| 2023.03.21 | Google opened Bard (later renamed Gemini) | Google's official entry |
| 2023.04 | Alibaba Qwen released | Domestic application-layer follow-up |
| 2023.05 | iFlytek Spark, ChatGPT iOS App | Multi-platform acceleration |
| 2023.07 | Anthropic Claude 2; Meta open-sourced Llama 2 | Open-source camp expanded (see Llama and the Open-Source Ecosystem) |
| 2023.08 | ByteDance Doubao released | Domestic traffic play |
| 2023.10 | Moonshot Kimi released | Focused on ultra-long context |
| 2023.12 | Google released Gemini 1.0 | Native multimodal path |
| 2024.05 | OpenAI released GPT-4o | Free, real-time voice multimodal |
| Post-2025 | o1/o3, GPT-5, Gemini 2.x, Claude 4/5 | Reasoning models and multimodal free-for-all |
1. Differentiation of Chinese Vendors
Domestic dialogue products (Ernie Bot, Qwen, Doubao, Kimi, DeepSeek dialogue version, etc.) generally compete along three paths: free traffic entry (Doubao, Ernie), ultra-long context (Kimi), and open-source weights + extreme cost-effectiveness (DeepSeek; see Llama and the Open-Source Ecosystem). They form offsetting competition with OpenAI through Chinese corpus, compliance review, and mobile distribution.
2. Evaluation Dimensions for Dialogue Products
As competitors multiply, "usable" is no longer the threshold — the competition is about: speed (TTFT/throughput), multimodal, context length, tool-calling capability, cost, and safety compliance. These dimensions are systematically compared in Model Compendium.
VI. Evaluating Dialogue Models: Helpfulness / Safety / Multi-Turn Consistency
Dialogue is open-ended, and the traditional "standard-answer accuracy" breaks down. The industry has developed three categories of evaluation:
Dialogue model evaluation shouldn't only happen before launch. Daily operations need a "daily/weekly" rhythm: sample and manually review a few conversations daily (label good/bad), run a golden set regression weekly, and compare with Arena and internal leaderboards monthly. Every prompt change, model upgrade, and knowledge base update must pass through this evaluation pipeline before going live. Making evaluation "daily" is the only way to prevent dialogue products from "slowly degrading" — both models and data change, and yesterday's good results may not hold today. Evaluation system setup is in Evaluations in Practice.
| Evaluation Dimension | What It Tests | Representative Methods |
|---|---|---|
| Helpfulness | Does the answer solve the user's problem? Is it accurate? | MT-Bench (GPT-4 scoring), LMSYS Chatbot Arena (human blind-test Elo) |
| Safety | Does it refuse harmful requests? Is it neutral and harmless? | Red-teaming, jailbreak eval sets, HarmBench, etc. |
| Multi-turn consistency | Stays on track in long conversations, no self-contradiction, remembers context | Multi-turn dialogue test sets, character consistency eval |
1. Human Blind Testing: Chatbot Arena
LMSYS Chatbot Arena lets users vote between two anonymous model responses, generating a "crowdsourced leaderboard" via Elo rankings. Its value lies in real users + real need distribution, avoiding benchmark contamination; its flaws are sample bias and speed/style preferences. It is the most credible dialogue capability leaderboard post-2023.
2. Model Judges: MT-Bench
Using GPT-4 to score two models' responses (1–10) across 80 multi-turn questions covering writing, reasoning, math, code, and roleplay. The advantage is low-cost reproducibility; the disadvantage is the judge model's own biases (position bias, self-preference), requiring cross-validation — see Evaluations in Practice.
3. Don't Forget the Guardrails Themselves
Before a dialogue product launches, it must pass safety acceptance: jailbreak success rate, harmful content refusal rate, and data leakage rate. Many models that "rank first on capability leaderboards" fail in safety tests — the balance between capability and safety is the core engineering challenge for dialogue products. See Evaluation and Benchmarks.
Leaderboards ≠ Your Scenario
The #1 model on Arena may not be the best choice for your customer service, your legal team, or your education scenario. Dialogue model evaluation needs all three: leaderboards + custom golden set + online metrics (retention, positive rate, complaint rate).
VII. ChatGPT's Product Evolution: From Dialogue to Ecosystem
| Stage | Time | Key Events |
|---|---|---|
| 1.0 Dialogue assistant | 2022.11–2023.3 | ChatGPT launch, GPT-4 on web, paid subscription ChatGPT Plus |
| 2.0 Multi-platform & tools | 2023.3–2024.5 | iOS/Android apps, browsing with Bing, code interpreter, image understanding (GPT-4V), plugins/function calling |
| 3.0 Fully multimodal free | 2024.5–2024.9 | GPT-4o free access, real-time voice, memory features, Sora video integration |
| 4.0 Reasoning & search | 2024.9–2025 | o1 reasoning model, ChatGPT Search (deep search), Operator (Agent browser operation), GPT-5 unified series |
The main thread of product evolution is from "chatting" to "working": adding tools (browsing, code, images), adding memory, adding Agent capability — exactly the topic of LLM-Based Agents. ChatGPT evolved from a "dialogue product" to a "personal AI assistant platform."
VIII. Engineering Details of Dialogue Systems
Dialogue products look like "just a chat box," but behind the scenes is a complete engineering stack. This section adds product-layer and engineering-layer details.
1. Model Lineage: From text-davinci to gpt-3.5-turbo
ChatGPT's foundation isn't a single model but a continuously cost-optimizing model line:
| Codename | Time | Characteristics |
|---|---|---|
| text-davinci-002/003 | 2022 | InstructGPT productized, expensive and slow per call |
| gpt-3.5-turbo | 2023.3 | Dialogue-optimized + price dropped to ~1/10, supports multi-turn message arrays |
| gpt-4 / gpt-4-turbo | 2023 | Stronger reasoning, 128K context |
| gpt-4o | 2024.5 | Fully multimodal, speed and price drop further |
gpt-3.5-turbo's "message array" interface (system/user/assistant alternating) became the industry standard — it encodes "conversation history" directly into the API, simplifying multi-turn management.
2. How Dialogue Data Is Built
A dialogue model's "usability" depends on high-quality multi-turn data:
- Dialogue trees: the same request is annotated with multiple valid response paths, covering "probing, clarifying, refusing, admitting not knowing";
- Role distribution: covers both user roleplay (probing, interrupting, challenging) and assistant roleplay (guiding, summarizing, correcting);
- Quality over quantity: better to have fewer, high-quality samples than many low-quality ones — low-quality data directly pollutes RLHF preferences (data methodology in Pretraining: Data and Objectives).
3. Product Engineering: Streaming, Concurrency, and Cost
| Engineering Point | Approach | Effect |
|---|---|---|
| Streaming output (SSE) | Tokens sent character by character | First-character latency drops to milliseconds, greatly improving experience |
| Continuous batching | Dynamically merging multiple requests for inference | GPU utilization jumps from single digits to 90%+ |
| KV Cache | Caching historical attention key-values | Generation speed improves several-fold (see Inference Fundamentals) |
| Cost control | Routing (simple tasks go to small models), rate limiting | Per-conversation cost is controllable |
The complete deployment checklist is in Deployment and Servicing.
4. Conversation Management and Multi-Turn Memory
- History truncation: very long conversations are trimmed by token budget, keeping only the most recent N turns;
- Summary compression: early conversation is summarized and fed back into the system prompt (see Context and Long Context);
- Context injection: current time, user preferences, and enterprise knowledge are injected via system prompt.
5. Dialogue Model Evaluation Practice
Before launch, run at least three categories of metrics: capability leaderboard (Arena Elo / MT-Bench), safety metrics (jailbreak success rate, refusal rate, data leakage rate), and business metrics (retention, positive rate, complaint rate, task completion rate). Evaluation system setup is in Evaluations in Practice and Evaluation and Benchmarks.
The three layers of dialogue product polish
Model determines the ceiling, data determines personality, engineering determines experience — the same GPT-4 foundation can produce wildly different dialogue products from different teams. Getting the data loop running (user feedback flowing back for retraining) is more valuable than chasing new models.
IX. Underlying Lessons from Dialogue Models
- Productization is a technology multiplier: InstructGPT's alignment technology lay in the paper for 8 months; it was the three product decisions of "free + dialogue interface + guardrails" that leveraged the whole world.
- Dialogue is the "native UI" for LLMs: dialogue can express any task, enable multi-turn error correction, and carry tool calling — it became the standard interaction layer for all large model applications.
- Guardrails determine product survival: no matter how strong the technology, if safety isn't solid, large-scale commercialization is impossible; the moat of dialogue models = capability × safety × cost × experience.
- The race has no finish line: from late 2022 to 2025, dialogue model capability has gone through several generations, but the race for "context, memory, Agent, multimodal, low cost" continues.
X. Product Methodology from ChatGPT
ChatGPT's success belongs not just to the model but to product. This section breaks it down into reusable methodology.
1. Four Elements of Dialogue Products
| Element | Meaning | Negative Example |
|---|---|---|
| Capability | Task completed correctly and well | Off-topic, rampant hallucinations |
| Experience | Fast, stable, human-like | 30-second loading, repetitive rambling |
| Trust | Honest, explainable, guarded | Fabricated data, over-refusing users |
| Cost | Per-conversation cost is controllable | Burning several yuan per conversation |
These four constrain each other: adding guardrails reduces "satisfaction"; reducing cost may hurt capability. The product manager's daily job is finding balance in this quadrant.
2. How to Evaluate a Dialogue's Quality
| Check Item | Example | Notes |
|---|---|---|
| Goal achieved | Did the user get what they wanted? | Core |
| Factual correctness | Numbers/dates/names accurate? | Errors can be fatal |
| Multi-turn consistency | Doesn't contradict previous statements? | Common in long conversations |
| Safety and compliance | No jailbreaks/harmful output? | Launch baseline |
| Efficiency | Doesn't go in circles? | Affects feel |
Evaluation system setup is in Evaluations in Practice.
3. Balancing Cost and Scale
Cost formula for scaling dialogue products:
Daily cost ≈ daily requests × avg tokens × unit price
Cost reduction paths:
① Routing: simple tasks go to small models (can save 80%+)
② Caching: hit and reuse (similar questions don't recompute)
③ Prompt slimming: compress system prompt and history
④ Distillation: replace with small models in high-frequency scenarios4. Recommended Path for Enterprise Deployment
- PoC (week-level): API + prompt engineering to validate value;
- Add RAG: connect enterprise knowledge base for factual needs (see RAG: Retrieval-Augmented Generation);
- Add tools and Agents: connect business systems for "working" needs (see LLM-Based Agents);
- Fine-tuning and distillation: lock in style, reduce cost;
- Private deployment / open-source: for data-sensitive scenarios, eventually converge to self-deployment (see Llama and the Open-Source Ecosystem).
Set an evaluation set at each step to avoid "feels better" — verify with data.
5. Common Misconceptions and Anti-Patterns
| Misconception | Correct Approach |
|---|---|
| Using a dialogue model as a database | Use RAG/search |
| Making "more human-like" the #1 goal | Prioritize facts and safety first |
| Endlessly stacking prompts | Simplify, structure, version |
| Only looking at Arena leaderboards for model selection | Build a custom golden set for evaluation |
| Ignoring the product after launch | Build feedback loops, iterate continuously |
6. Case Study: Building an Enterprise Customer Service Assistant from Scratch
A recommended path for a 4-person team to build a usable customer service assistant in 2 weeks:
Day 1: Define scope (only answer product/after-sales questions), organize 50 typical Q&A pairs
Day 2–3: Split FAQ and manual into chunks, build vector DB (RAG retrieval)
Day 4–5: Write system prompt (role + rules + citation format), connect API
Day 6–7: Evaluate with 50 real questions, iterate prompts and retrieval
Day 8–10: Add human fallback (transfer to human), logging, safety guardrails
Day 11–14: Small-traffic gray release, collect feedback, expand evaluation setKey principles: start small, rules before models, retrieval before fine-tuning. Most customer service scenarios can reach usability within "prompt + RAG + evaluation loop"; only when style and format have hard requirements should you consider LoRA fine-tuning (see Fine-tuning Practice). Keep daily cost in the tens to hundreds of yuan range; introduce routing and caching once scale increases.
The key to this case isn't how advanced the tech stack is, but "scope convergence + feedback loop": scope creep is the most common reason customer service assistants fail (the model gets asked about sales, complaints, complex business processes); the feedback loop (daily review of failure cases and improvement) is the only mechanism for continuous improvement. Getting 50 questions right at 90% accuracy before expanding to 500 is more realistic than pursuing full coverage from the start.
The core metrics worth remembering for customer service assistants: human transfer rate, first-contact resolution rate, average handling time, and user satisfaction. Model capability improvements typically manifest as reduced transfer rates and increased first-contact resolution; these online metrics, linked with offline evaluation sets, tell you whether a prompt change or model upgrade is real.
The ultimate moat of dialogue products
Not the model, but the data loop: user questions → quality annotation → evaluation set expansion → prompt/fine-tuning/retrieval improvement → launch → flow back. Models iterate, but the data loop is the mechanism for sustained lead.
XI. From ChatGPT to the Future
The next decade of dialogue models has three main threads: multimodal dialogue (voice/image/video as dialogue content, see Multimodal LLMs), Agent-ified dialogue (dialogue evolving from "answering questions" to "executing tasks"), and personalization and memory (remembering users across sessions). ChatGPT defined the starting point, but these directions belong to the entire industry.
One dimension often overlooked: dialogue models are becoming the default interaction layer for all software products. Ten years ago, apps interacted via buttons and forms; today, more and more products use a "dialog box" as the entry point (customer service, data analysis, dev assistants). This means "dialogue" is no longer a specific product but a kind of infrastructure — like databases and APIs were for web applications. For developers, instead of asking "should we build a dialogue product," ask "where can dialogue capability be embedded in my product?" Embedding approaches include: document Q&A (RAG), operation assistant (Agent), content generation (writing/reports), and review assistant (human reviewing AI drafts). Each embedding needs to answer three questions: where does the data come from, what happens when it fails, and who is responsible — the engineering and accountability questions this page has emphasized throughout.
Finally, a boundary to clarify: ChatGPT is often mistaken for approaching artificial general intelligence, but it lacks sustained autonomous goals, a stable world model, and true causal understanding. "Strong dialogue capability ≠ intelligence" — dialogue is the interface of intelligence, not intelligence itself. See What Are Large Language Models? for related discussion.
XII. Social and Commercial Impact of Dialogue Models
1. Impact on the Labor Market
| Job Category | How Affected | Response Direction |
|---|---|---|
| Customer service / clerical | Dialogue models replace routine responses | Shift to complex complaints and experience design |
| Junior writing / translation | Generated drafts greatly accelerated | Shift to review, curation, and personalization |
| Programmers | Coding assistants boost efficiency | Shift to architecture, review, and requirements |
| Analysts | Reports and drafts automated | Shift to insight and decision-making |
The historical pattern is "technology eliminates positions, not occupations": dialogue models replace the "typing and retrieval" phase, amplifying judgment and communication skills.
2. Education, Academia, and Trust
- Academic integrity: AI-written papers/assignments spark an AI detection arms race;
- Information environment: AI-generated content makes truth harder to distinguish, making provenance and labeling (watermarking) a must-have;
- Education paradigm: shifting from "teaching retrieval" to "teaching questioning and critical thinking."
3. Regulation and Compliance
| Region | Key Regulation / Practice | Key Points |
|---|---|---|
| EU | AI Act (tiered regulation) | High-risk applications need assessment and transparency |
| China | Interim Measures for Generative AI Services | Filing, content safety, labeling |
| US | Executive orders and industry self-regulation | Self-assessment and voluntary commitments |
Compliance requirement details are subject to local official texts; products must undergo specific assessment before launch (see Safety and Risk).
4. Industry Penetration
Dialogue models have entered customer service, education, legal research, medical consultation, and financial advisory scenarios — but at varying depth: low-risk, high-certainty, verifiable scenarios are adopting fastest (e.g., customer service, knowledge base Q&A); high-risk scenarios (diagnosis, legal advice, financial decisions) still rely on "human-in-the-loop + assistant positioning."
5. A Balanced Perspective
Discussions about the impact of dialogue models often swing to two extremes: either "AI replaces everything" panic or "AI can't do anything" dismissal. Neither is accurate. A more realistic assessment: dialogue models are becoming a "universal productivity interface." They won't eliminate occupations, but they will redistribute lower-level tasks within occupations — automating information retrieval, draft writing, and formatting, pushing human effort toward judgment, negotiation, and creativity. For individuals, the key skill shifts from "knowing how to use tools" to "knowing how to define problems and evaluate results"; for organizations, the cost structure shifts from "human-scale" to "model-scale + human review." Regulators worldwide are building transparency, labeling, and accountability rules at different paces. Overall, dialogue models are an "accelerator" rather than a "replacement" — they amplify existing advantages and risks, neither creating nor destroying out of nothing.
This means that whether you're an engineer, product manager, or researcher, treating "how to collaborate with dialogue models" as a trainable skill is more valuable than debating "will it replace me." The skill list includes: writing clear instructions, designing evaluation sets, judging output quality, and integrating model capability into business processes. These skill-building pathways are exactly the combination of chapters in this handbook (prompting, RAG, Agents, evaluation, deployment).
XIII. FAQ Quick Answers
| Question | Quick Answer |
|---|---|
| What's the difference between ChatGPT free and paid? | Paid unlocks stronger models, longer context, and priority access |
| Why do some dialogues "confidently talk nonsense"? | Hallucination; mitigate with RAG/citations and "say I don't know" |
| How does multi-turn dialogue maintain context? | Product-layer history window management + summarization (see Context and Long Context) |
| Can dialogue models replace search engines? | For factual queries, pair with search/RAG; don't ask naked |
| How do I judge if a dialogue model is good? | Arena blind tests + custom evaluation set + safety testing |
| How can small teams deploy dialogue products? | Start with API + prompt + RAG, then consider fine-tuning |
Note: ChatGPT's other moat is the data flywheel — every like, dislike, and correction becomes a preference signal flowing back into the reward model and safety training. This is why "get it running first, collect feedback" follows the evolution pattern of dialogue products better than "hold back for a big launch." Methods for building up user feedback into evaluation sets and preference data are in Evaluations in Practice.
Note: Whether ChatGPT or competitors, the "usable" bar for dialogue products is being repeatedly reset — response speed and accuracy that are good enough today may be just passing grade a year from now. The one constant is the fundamentals of "evaluation sets + feedback loops + cost control." Do these three well, and you'll stay composed when models change.
Additionally: when iterating dialogue products, distinguish "model upgrades" from "configuration changes" — change only one variable per release (model, prompt, or retrieval), otherwise performance fluctuations are hard to attribute.
The right way to use dialogue models
Treat them as "extremely capable draft machines + interaction interfaces," not authorities. Every output should have a "verification path" (citations, sources, human confirmation), especially in professional domains.
XIV. Further Reading
- The GPT Series: From GPT-1 to GPT-4o — ChatGPT's technical foundation evolution
- Alignment: RLHF and DPO — ChatGPT's key training technology
- Safety and Risk — dialogue product guardrails and jailbreak arms race
- Llama and the Open-Source Ecosystem — open-source dialogue models catching up
- LLM-Based Agents — the next stop from dialogue to "working"
- Evolution Overview — the historical inevitability of the 2022 explosion
- Model Compendium — horizontal comparison of dialogue models
References
- OpenAI. Introducing ChatGPT (2022.11.30) — ChatGPT official release blog
- Ouyang et al. Training language models to follow instructions with human feedback (InstructGPT, 2022) — RLHF original paper (arXiv)
- Zheng et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (2023) — MT-Bench and Arena methodology (arXiv)
- Microsoft. Reinventing search with a new AI-powered Microsoft Bing (2023.2) — New Bing official blog
- Google. An important next step on our AI journey (Bard, 2023.3) — Google Bard official blog
- Anthropic. Claude 2 (2023.7) — Anthropic official release
- OpenAI. Hello GPT-4o (2024.5) — GPT-4o official release
- OpenAI. Learning to reason with LLMs (o1, 2024.9) — o1 reasoning model release notes