Appearance
Sora and Video Generation
Late on February 15, 2024, OpenAI released a demo video running under two minutes: a stylish woman walks through a neon-lit Tokyo street, the camera tracking her along the block, with lighting, reflections, facial expression, and camera movement all indistinguishable from reality. This was not live footage, and it was not traditional CGI — it was a continuous shot of roughly 60 seconds generated directly from a plain text prompt. Overnight, the demo swept the globe — Sora had arrived.
Sora pushed text-to-video from a laboratory toy to the level of "footage that looks like real film material," and it lit the fuse for the video-generation explosion of 2024–2025: Kuaishou's Kling, Runway Gen, Google Veo, Pika, and other products entered the market one after another, and the competition turned white-hot. This article reviews everything that happened in the two years from Sora's teaser to its launch: the technical foundation (video diffusion models and DiT), why video generation is harder than image generation, a comparison of the product landscape, and the industry impact and limitations. For background, start with Diffusion Models and Generative AI and Midjourney and Image Generation.
1. February 2024: Sora Goes Viral Overnight
| Time | Event |
|---|---|
| 2024.02 | OpenAI releases the Sora demo video and technical report, igniting a global sensation |
| 2024.06 | Kuaishou releases Kling; Runway releases Gen-3 Alpha, with video generation racing to catch up |
| 2024.12 | Sora officially launches (Sora.com, for ChatGPT Plus/Pro subscribers) |
| 2025.05 | Google I/O unveils Veo 3 / Veo 3.1 |
| 2025.06 | Kling 2.5 and Runway Gen-4 ship in succession, escalating the parameter and duration race |
Why did Sora's teaser hit so hard? Three reasons:
- Duration and coherence: Mainstream video generation models before it could only produce a few seconds of looping or short clips, and motion frequently "drifted"; Sora generated roughly 60 seconds of video with coherent camera movement in a single pass — a whole step up in quality;
- The world model narrative: In its technical report, OpenAI outright claimed that video generation models are "world simulators" — if a model can predict future frames, doesn't that mean it is learning physical laws and causality? This grand narrative elevated video generation from a "VFX tool" to "a step toward AGI";
- Big tech endorsement: OpenAI had just climbed to the crest of the AI wave on the strength of ChatGPT, so a new capability from it was naturally read as "the next generation of technology."
An important clarification
What was released in February 2024 was only a teaser: a small selection of curated demos, no public model, no open beta, and no full paper details disclosed; some demos were later called out as likely "cherry-picked/edited." Real productization had to wait until Sora's official launch in December 2024. The lesson: a technical preview is not technical reality — between media hype and engineering usability there often lies anywhere from a few months to a year.
2. The Technical Foundation: Video Diffusion Models, DiT, and Spacetime Patches
1. From image diffusion to video diffusion
Sora's technical roots are still the diffusion model, but the object being processed changes from "a single image" to "a video." Image diffusion models are a "spatial" task: add noise to an image, then learn to remove it. Video diffusion models must handle "space + time" simultaneously: an entire video (a stack of frames) is treated as the object to be generated — noise is added to every frame, and during denoising the neural network must keep each frame internally plausible while staying continuous from frame to frame. Early work such as Video Diffusion Models (Ho et al., 2022) had already validated the feasibility of "running diffusion directly on video," but compute constraints limited it to extremely short clips.
2. DiT: handing "denoising" to the Transformer
At the core of Sora's architecture is DiT (Diffusion Transformer), proposed by Peebles & Xie in 2023 (ICCV 2023). Before it, the backbone of diffusion models was the convolutional UNet; DiT instead cuts an image/video into a pile of patches and feeds them straight into a Transformer. The significance of this swap:
- Scalability: The Transformer's scaling law holds for diffusion tasks too — the bigger the model and the more training data, the higher the generation quality. This let video generation ride the "scale conquers all" fast lane for the first time;
- Long-range dependencies: Attention is naturally good at capturing relationships between "distant pixels/frames," which is exactly what video consistency requires.
Sora's pipeline can be summarized in three steps:
text
① Video compression network:
Raw video ──▶ lower-dimensional latent space ──▶ cut into spacetime patches
② DiT denoising:
Noisy patch sequence + text condition (large-capacity text encoder) ──▶ iterative denoising
③ Decoding:
Fully denoised patch sequence ──▶ decoder reconstructs the video framesThe key insight is "patches are tokens": once a video is compressed and cut into "spacetime blocks," it becomes something like a token sequence in a language model and can be handled uniformly with attention. OpenAI's technical report states explicitly that Sora's patching is conceptually the same move as LLM tokenization — video generation, Transformers and Attention, and Large Language Models converge at the architectural level.
3. The relationship to LLMs: extending the scaling law
Sora is not "an image model tweaked to do video"; many of its design choices borrow directly from LLM experience:
- Data scale: Sora was trained on unprecedented volumes of "video + text" data, mixing images and videos during training (images are treated as single-frame videos);
- Condition injection: The prompt is encoded by a large-capacity text encoder and injected into the generation process, with stronger semantic understanding than the CLIP used by early image models;
- A unified token view: Because images count as "single-frame videos," Sora can handle text-to-video, image-to-video, video continuation, and editing in one model — what OpenAI calls "a foundation for understanding and simulating the dynamic world."
3. Video Generation vs. Image Generation: What Makes It Hard
Video is not simply "a sequence of images." Going from image generation to video generation adds at least four hurdles:
| Hurdle | How it shows up | Impact |
|---|---|---|
| Temporal consistency | The same object must not drift, deform, or suddenly change outfits between frames | Characters, backgrounds, and lighting must be locked down across frames |
| Motion physics | Physical laws — gravity, collisions, occlusion, fluids — must behave sensibly | Physically absurd "ghostly deformations" are the classic failure point |
| Duration and resolution | Duration × resolution × frame rate jointly determine the "amount of information" | Every extra frame and every resolution notch adds a step change in compute |
| Compute cost | Training and inference are one to two orders of magnitude more expensive than for images | Directly determines product pricing and how widely a product opens up |
Compute cost deserves elaboration: a 60-second 720p video contains roughly 1,800 frames, whereas a 1024×1024 image is just one "frame." Video generation bundles a "spatial task" and a "temporal task" together, and its training and inference overhead grows roughly with "spatial resolution × duration" — the fundamental reason text-to-video products generally bill by the "second" and impose monthly limits. Early on after launch, Sora users could only generate a limited number of videos per month, precisely because of cost. For inference-side optimization options (distillation, quantization, sparsification), see Inference Optimization and Quantization.
Physical common sense: video generation's biggest failure zone
Counting objects, pouring water, walking through walls, the number of fingers changing mid-wave — problems that "single-frame stillness" hides in image models get exposed in video as "dynamic bloopers." Sora's demos and later user tests surfaced plenty of such errors, showing that the model has learned "data correlations" rather than "physical laws." This is the widest gap between the "world model" narrative and engineering reality.
4. The 2024–2025 Video Generation Product Landscape
After Sora lit the fuse, vendors in China and abroad turned video generation into an "arms race" within 18 months. As of mid-2025, the main products are roughly:
| Product | Vendor | Release cadence | Resolution & duration (approx.) | Availability |
|---|---|---|---|---|
| Sora | OpenAI | Teased 2024.02, launched 2024.12, multiple iterations through 2025 | 1080p, ~20s per clip (extendable to 60s+) | Subscription (ChatGPT Plus/Pro), Sora.com |
| Kling | Kuaishou | Released 2024.06, Kling 2.5 in 2025.06 | 1080p/4K, up to ~3 minutes (multi-shot) | App/web, free tier available |
| Runway Gen | Runway | Gen-3 (2024.06), Gen-4 (2025.06) | 1080p, ~10s per clip | Web subscription + API, film-oriented toolchain |
| Veo | Google DeepMind | Veo (2024.05), Veo 3/3.1 (2025.05) | 1080p, from ~8s per clip | Integrated into the Gemini app/Flow, limited availability |
| Pika | Pika Labs | 1.0 (2023.12), 2.0 (2024.11) | From 720p, ~10s | Web, focused on fun and ease of use |
Several competitive dimensions worth noting:
- Duration: Kling pushed maximum duration to 3 minutes (multi-shot stitching), Sora emphasizes single-clip coherence, and Runway goes the "cinematic single-shot" route — "long" and "stable" each involve trade-offs;
- Consistency: Runway Gen-4 and Kling 2.x both make "character/scene consistency across shots" their headline selling point — the key step from "generating a single clip" toward "generating usable content";
- Ecosystem play: Runway builds a creator toolchain (editing, editing APIs), Google embeds Veo into the Gemini multimodal suite, and OpenAI ties Sora into the ChatGPT subscription — big tech is all using "video generation" to feed its own main battleground. For a quick model capability reference, see Model and Leaderboard Quick Reference.
5. Applications and Impact: Film, Advertising, and the Authenticity Crisis
1. A game-changing disruption for the content industry
Video generation hits content production more directly than image generation, because video is the content industry's primary medium:
- Advertising and e-commerce: Product videos, voiceover footage, cinematic themed ads — a single shot used to cost several thousand yuan; now text-to-video can mass-produce candidates;
- Film pre-production: Storyboards, previsualization (previs), concept tests — directors can quickly "see" how a shot will look from text, sharply cutting communication costs;
- Short video and UGC: Ordinary creators gained, for the first time, the ability to produce cinematic shots without a film crew, and the volume of content supply exploded;
- Gaming and the metaverse: Generative production of cutscenes, ambient video, dynamic NPC expressions, and other assets is entering game pipelines.
2. Deepfake risks and governance
The other side of video generation is the collapse of authenticity: a 60-second video that passes as real further shakes "seeing is believing," a baseline of social trust. AI face-swap videos, fabricated celebrity statements, forged incident footage — the distribution cost of such deepfakes has been driven down to nearly nothing, and since 2024 "AI-synthesized celebrity/executive videos" have fueled multiple real fraud cases. Industry and regulatory responses operate on several layers:
- Proactive watermarking and provenance: OpenAI embeds C2PA (Content Credentials) metadata in Sora output, marking it "AI-generated"; Google's SynthID writes an invisible watermark directly into the pixels, which stays detectable after cropping and compression;
- Platform labeling: Major video platforms are progressively requiring mandatory labels on AI-generated content;
- Legislation and labeling: The EU AI Act requires explicit disclosure of deepfakes; China's Measures for Labeling AI-Generated Synthetic Content took effect in September 2025, requiring both explicit and implicit labels on AI-generated content — governments everywhere are turning "traceability of generative content" into a compliance obligation.
For a more systematic discussion of watermarking, detection, and governance, see AI Safety and Governance.
Detection always lags generation
Once generation quality gets high enough, "after-the-fact detection" is doomed to lose — AI-generated content gets screened by another AI, and the detector is soon bypassed by newer generators. Hence the industry consensus: proactive labeling (tagging at publication) beats passive detection (forensics after the fact), which is exactly why C2PA and SynthID have been adopted across the board by mainstream vendors.
6. Limitations and the Evaluation Problem
1. The model's hard weaknesses
As of mid-2025, even the most advanced video generation models still exhibit several recurring failure modes:
- Physical common-sense errors: Object counts, gravity direction, fluids, and occlusion relationships frequently go wrong, especially in longer videos;
- Cross-scene consistency: The same character's appearance drifts across multiple videos, and props and costumes cannot be locked down reliably;
- Text semantics: Complex instructions, negations, and cause-and-effect ("because... therefore...") relationships remain hard to follow strictly;
- Generation speed and cost: High-quality video generation is measured in "minutes," so "instant video" remains unrealistic in commercial settings.
2. Evaluation is hard: video has no "ground truth"
"Is this video good?" is harder to answer than "is this image good?" because quality spans at least three orthogonal dimensions: per-frame image quality, temporal consistency, and motion plausibility. On automatic metrics, the video field has borrowed the image playbook: FVD (Fréchet Video Distance) compares the distributions of real and generated videos in feature space, while CLIP-style metrics measure video–text alignment. But just as with image evaluation, automatic metrics cannot reflect subjective dimensions like "is the motion natural" or "does the narrative make sense," so the field ultimately still relies on large-scale human blind testing (for example, LMArena's video Arena, where users vote between two models' outputs). To place video generation within the broader framework of "how generative AI gets evaluated," see LLM Evaluation and Benchmarks.
A quick heuristic
When evaluating a video model, first check "do objects stay the same over a long video," then "does the motion match physical intuition," and only then look at per-frame image quality — reverse the order and you can easily be fooled by pretty static frames.
7. Trade-offs
- Duration vs. stability: The longer the video, the harder consistency becomes. Models generally adopt a "short single clip (≤20s) + multi-shot stitching" strategy, forcing users to choose between "long" and "stable";
- Quality vs. cost: Resolution, duration, and number of iterations directly determine inference cost; for commercial use, budget on a "per thousand generations" basis instead of judging by demo results;
- Openness vs. control: Web-subscription black-box products work out of the box; the API and open-source routes (community follow-up on open video models like VideoCrafter, CogVideoX, and LTX-Video) give you control but carry high engineering costs;
- Capability vs. responsibility: The stronger the generation capability, the bigger the investment required in watermarking, moderation, and compliance — video is "higher-risk" than images because it persuades more and spreads farther.
Further Reading
- Diffusion Models and Generative AI: the technical foundation of video generation;
- Midjourney and Image Generation: the image revolution came first, then the video revolution;
- Transformers and Attention: how DiT, spacetime patches, and attention underpin video generation;
- Large Language Models: the shared lineage between Sora's tokenization, scaling laws, and LLMs;
- Multimodal Models: video as a member of the multimodal family;
- Inference Optimization and Quantization: why video inference is expensive and how to bring the cost down;
- LLM Evaluation and Benchmarks: FVD, human blind testing, and video evaluation methodology;
- AI Safety and Governance: deepfakes, C2PA/SynthID watermarking, and global regulation;
- A Brief History and Learning Paths: placing text-to-video within the full landscape of hot AI concepts;
- Model and Leaderboard Quick Reference: a quick reference to mainstream video models' versions and capabilities.
References
- Brooks, Peebles, Holmes, DePue, et al. Video generation models as world simulators (OpenAI technical report, 2024) — Sora's official technical report
- Peebles, Xie. Scalable Diffusion Models with Transformers (ICCV 2023) — DiT, the architectural cornerstone of Sora
- Ho et al. Video Diffusion Models (NeurIPS 2022) — early foundational work on video diffusion models
- Blattmann et al. Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets and Long Video Generation (2023) — SVD, a major milestone for open-source video generation
- Unterthiner et al. Towards Accurate Generative Models of Video: A New Metric & Challenges (2018) — the original FVD metric paper
- OpenAI. Sora product page — Sora features, pricing, and version notes
- Runway. Gen-4 official introduction — Runway's video generation product page
- Kuaishou Kling official website — Kling product entry point and version notes
- Google DeepMind. Veo — Veo's official technical page
- Pika official website — Pika product entry point
- C2PA. Content Credentials — the Coalition for Content Provenance and Authenticity standard
- Google DeepMind. SynthID — invisible watermarking technology