Skip to content

Sora and Video Generation

At a glance A retrospective on how OpenAI's Sora, from its early-2024 teaser to its official launch, ignited the text-to-video explosion, covering video diffusion models and the DiT architecture, the spacetime patching technical foundation, the product landscape spanning Sora, Kling, Runway, Veo, and Pika, and the physical limitations, deepfake governance, and evaluation challenges that remain.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

Sora and Video Generation ​

Late on February 15, 2024, OpenAI released a demo video running under two minutes: a stylish woman walks through a neon-lit Tokyo street, the camera tracking her along the block, with lighting, reflections, facial expression, and camera movement all indistinguishable from reality. This was not live footage, and it was not traditional CGI — it was a continuous shot of roughly 60 seconds generated directly from a plain text prompt. Overnight, the demo swept the globe — Sora had arrived.

Sora pushed text-to-video from a laboratory toy to the level of "footage that looks like real film material," and it lit the fuse for the video-generation explosion of 2024–2025: Kuaishou's Kling, Runway Gen, Google Veo, Pika, and other products entered the market one after another, and the competition turned white-hot. This article reviews everything that happened in the two years from Sora's teaser to its launch: the technical foundation (video diffusion models and DiT), why video generation is harder than image generation, a comparison of the product landscape, and the industry impact and limitations. For background, start with Diffusion Models and Generative AI and Midjourney and Image Generation.

1. February 2024: Sora Goes Viral Overnight ​

TimeEvent
2024.02OpenAI releases the Sora demo video and technical report, igniting a global sensation
2024.06Kuaishou releases Kling; Runway releases Gen-3 Alpha, with video generation racing to catch up
2024.12Sora officially launches (Sora.com, for ChatGPT Plus/Pro subscribers)
2025.05Google I/O unveils Veo 3 / Veo 3.1
2025.06Kling 2.5 and Runway Gen-4 ship in succession, escalating the parameter and duration race

Why did Sora's teaser hit so hard? Three reasons:

  1. Duration and coherence: Mainstream video generation models before it could only produce a few seconds of looping or short clips, and motion frequently "drifted"; Sora generated roughly 60 seconds of video with coherent camera movement in a single pass — a whole step up in quality;
  2. The world model narrative: In its technical report, OpenAI outright claimed that video generation models are "world simulators" — if a model can predict future frames, doesn't that mean it is learning physical laws and causality? This grand narrative elevated video generation from a "VFX tool" to "a step toward AGI";
  3. Big tech endorsement: OpenAI had just climbed to the crest of the AI wave on the strength of ChatGPT, so a new capability from it was naturally read as "the next generation of technology."

An important clarification

What was released in February 2024 was only a teaser: a small selection of curated demos, no public model, no open beta, and no full paper details disclosed; some demos were later called out as likely "cherry-picked/edited." Real productization had to wait until Sora's official launch in December 2024. The lesson: a technical preview is not technical reality — between media hype and engineering usability there often lies anywhere from a few months to a year.

2. The Technical Foundation: Video Diffusion Models, DiT, and Spacetime Patches ​

1. From image diffusion to video diffusion ​

Sora's technical roots are still the diffusion model, but the object being processed changes from "a single image" to "a video." Image diffusion models are a "spatial" task: add noise to an image, then learn to remove it. Video diffusion models must handle "space + time" simultaneously: an entire video (a stack of frames) is treated as the object to be generated — noise is added to every frame, and during denoising the neural network must keep each frame internally plausible while staying continuous from frame to frame. Early work such as Video Diffusion Models (Ho et al., 2022) had already validated the feasibility of "running diffusion directly on video," but compute constraints limited it to extremely short clips.

2. DiT: handing "denoising" to the Transformer ​

At the core of Sora's architecture is DiT (Diffusion Transformer), proposed by Peebles & Xie in 2023 (ICCV 2023). Before it, the backbone of diffusion models was the convolutional UNet; DiT instead cuts an image/video into a pile of patches and feeds them straight into a Transformer. The significance of this swap:

  • Scalability: The Transformer's scaling law holds for diffusion tasks too — the bigger the model and the more training data, the higher the generation quality. This let video generation ride the "scale conquers all" fast lane for the first time;
  • Long-range dependencies: Attention is naturally good at capturing relationships between "distant pixels/frames," which is exactly what video consistency requires.

Sora's pipeline can be summarized in three steps:

text
① Video compression network:
   Raw video ──▶ lower-dimensional latent space ──▶ cut into spacetime patches

② DiT denoising:
   Noisy patch sequence + text condition (large-capacity text encoder) ──▶ iterative denoising

③ Decoding:
   Fully denoised patch sequence ──▶ decoder reconstructs the video frames

The key insight is "patches are tokens": once a video is compressed and cut into "spacetime blocks," it becomes something like a token sequence in a language model and can be handled uniformly with attention. OpenAI's technical report states explicitly that Sora's patching is conceptually the same move as LLM tokenization — video generation, Transformers and Attention, and Large Language Models converge at the architectural level.

3. The relationship to LLMs: extending the scaling law ​

Sora is not "an image model tweaked to do video"; many of its design choices borrow directly from LLM experience:

  • Data scale: Sora was trained on unprecedented volumes of "video + text" data, mixing images and videos during training (images are treated as single-frame videos);
  • Condition injection: The prompt is encoded by a large-capacity text encoder and injected into the generation process, with stronger semantic understanding than the CLIP used by early image models;
  • A unified token view: Because images count as "single-frame videos," Sora can handle text-to-video, image-to-video, video continuation, and editing in one model — what OpenAI calls "a foundation for understanding and simulating the dynamic world."

3. Video Generation vs. Image Generation: What Makes It Hard ​

Video is not simply "a sequence of images." Going from image generation to video generation adds at least four hurdles:

HurdleHow it shows upImpact
Temporal consistencyThe same object must not drift, deform, or suddenly change outfits between framesCharacters, backgrounds, and lighting must be locked down across frames
Motion physicsPhysical laws — gravity, collisions, occlusion, fluids — must behave sensiblyPhysically absurd "ghostly deformations" are the classic failure point
Duration and resolutionDuration × resolution × frame rate jointly determine the "amount of information"Every extra frame and every resolution notch adds a step change in compute
Compute costTraining and inference are one to two orders of magnitude more expensive than for imagesDirectly determines product pricing and how widely a product opens up

Compute cost deserves elaboration: a 60-second 720p video contains roughly 1,800 frames, whereas a 1024×1024 image is just one "frame." Video generation bundles a "spatial task" and a "temporal task" together, and its training and inference overhead grows roughly with "spatial resolution × duration" — the fundamental reason text-to-video products generally bill by the "second" and impose monthly limits. Early on after launch, Sora users could only generate a limited number of videos per month, precisely because of cost. For inference-side optimization options (distillation, quantization, sparsification), see Inference Optimization and Quantization.

Physical common sense: video generation's biggest failure zone

Counting objects, pouring water, walking through walls, the number of fingers changing mid-wave — problems that "single-frame stillness" hides in image models get exposed in video as "dynamic bloopers." Sora's demos and later user tests surfaced plenty of such errors, showing that the model has learned "data correlations" rather than "physical laws." This is the widest gap between the "world model" narrative and engineering reality.

4. The 2024–2025 Video Generation Product Landscape ​

After Sora lit the fuse, vendors in China and abroad turned video generation into an "arms race" within 18 months. As of mid-2025, the main products are roughly:

ProductVendorRelease cadenceResolution & duration (approx.)Availability
SoraOpenAITeased 2024.02, launched 2024.12, multiple iterations through 20251080p, ~20s per clip (extendable to 60s+)Subscription (ChatGPT Plus/Pro), Sora.com
KlingKuaishouReleased 2024.06, Kling 2.5 in 2025.061080p/4K, up to ~3 minutes (multi-shot)App/web, free tier available
Runway GenRunwayGen-3 (2024.06), Gen-4 (2025.06)1080p, ~10s per clipWeb subscription + API, film-oriented toolchain
VeoGoogle DeepMindVeo (2024.05), Veo 3/3.1 (2025.05)1080p, from ~8s per clipIntegrated into the Gemini app/Flow, limited availability
PikaPika Labs1.0 (2023.12), 2.0 (2024.11)From 720p, ~10sWeb, focused on fun and ease of use

Several competitive dimensions worth noting:

  • Duration: Kling pushed maximum duration to 3 minutes (multi-shot stitching), Sora emphasizes single-clip coherence, and Runway goes the "cinematic single-shot" route — "long" and "stable" each involve trade-offs;
  • Consistency: Runway Gen-4 and Kling 2.x both make "character/scene consistency across shots" their headline selling point — the key step from "generating a single clip" toward "generating usable content";
  • Ecosystem play: Runway builds a creator toolchain (editing, editing APIs), Google embeds Veo into the Gemini multimodal suite, and OpenAI ties Sora into the ChatGPT subscription — big tech is all using "video generation" to feed its own main battleground. For a quick model capability reference, see Model and Leaderboard Quick Reference.

5. Applications and Impact: Film, Advertising, and the Authenticity Crisis ​

1. A game-changing disruption for the content industry ​

Video generation hits content production more directly than image generation, because video is the content industry's primary medium:

  • Advertising and e-commerce: Product videos, voiceover footage, cinematic themed ads — a single shot used to cost several thousand yuan; now text-to-video can mass-produce candidates;
  • Film pre-production: Storyboards, previsualization (previs), concept tests — directors can quickly "see" how a shot will look from text, sharply cutting communication costs;
  • Short video and UGC: Ordinary creators gained, for the first time, the ability to produce cinematic shots without a film crew, and the volume of content supply exploded;
  • Gaming and the metaverse: Generative production of cutscenes, ambient video, dynamic NPC expressions, and other assets is entering game pipelines.

2. Deepfake risks and governance ​

The other side of video generation is the collapse of authenticity: a 60-second video that passes as real further shakes "seeing is believing," a baseline of social trust. AI face-swap videos, fabricated celebrity statements, forged incident footage — the distribution cost of such deepfakes has been driven down to nearly nothing, and since 2024 "AI-synthesized celebrity/executive videos" have fueled multiple real fraud cases. Industry and regulatory responses operate on several layers:

  • Proactive watermarking and provenance: OpenAI embeds C2PA (Content Credentials) metadata in Sora output, marking it "AI-generated"; Google's SynthID writes an invisible watermark directly into the pixels, which stays detectable after cropping and compression;
  • Platform labeling: Major video platforms are progressively requiring mandatory labels on AI-generated content;
  • Legislation and labeling: The EU AI Act requires explicit disclosure of deepfakes; China's Measures for Labeling AI-Generated Synthetic Content took effect in September 2025, requiring both explicit and implicit labels on AI-generated content — governments everywhere are turning "traceability of generative content" into a compliance obligation.

For a more systematic discussion of watermarking, detection, and governance, see AI Safety and Governance.

Detection always lags generation

Once generation quality gets high enough, "after-the-fact detection" is doomed to lose — AI-generated content gets screened by another AI, and the detector is soon bypassed by newer generators. Hence the industry consensus: proactive labeling (tagging at publication) beats passive detection (forensics after the fact), which is exactly why C2PA and SynthID have been adopted across the board by mainstream vendors.

6. Limitations and the Evaluation Problem ​

1. The model's hard weaknesses ​

As of mid-2025, even the most advanced video generation models still exhibit several recurring failure modes:

  • Physical common-sense errors: Object counts, gravity direction, fluids, and occlusion relationships frequently go wrong, especially in longer videos;
  • Cross-scene consistency: The same character's appearance drifts across multiple videos, and props and costumes cannot be locked down reliably;
  • Text semantics: Complex instructions, negations, and cause-and-effect ("because... therefore...") relationships remain hard to follow strictly;
  • Generation speed and cost: High-quality video generation is measured in "minutes," so "instant video" remains unrealistic in commercial settings.

2. Evaluation is hard: video has no "ground truth" ​

"Is this video good?" is harder to answer than "is this image good?" because quality spans at least three orthogonal dimensions: per-frame image quality, temporal consistency, and motion plausibility. On automatic metrics, the video field has borrowed the image playbook: FVD (Fréchet Video Distance) compares the distributions of real and generated videos in feature space, while CLIP-style metrics measure video–text alignment. But just as with image evaluation, automatic metrics cannot reflect subjective dimensions like "is the motion natural" or "does the narrative make sense," so the field ultimately still relies on large-scale human blind testing (for example, LMArena's video Arena, where users vote between two models' outputs). To place video generation within the broader framework of "how generative AI gets evaluated," see LLM Evaluation and Benchmarks.

A quick heuristic

When evaluating a video model, first check "do objects stay the same over a long video," then "does the motion match physical intuition," and only then look at per-frame image quality — reverse the order and you can easily be fooled by pretty static frames.

7. Trade-offs ​

  • Duration vs. stability: The longer the video, the harder consistency becomes. Models generally adopt a "short single clip (≤20s) + multi-shot stitching" strategy, forcing users to choose between "long" and "stable";
  • Quality vs. cost: Resolution, duration, and number of iterations directly determine inference cost; for commercial use, budget on a "per thousand generations" basis instead of judging by demo results;
  • Openness vs. control: Web-subscription black-box products work out of the box; the API and open-source routes (community follow-up on open video models like VideoCrafter, CogVideoX, and LTX-Video) give you control but carry high engineering costs;
  • Capability vs. responsibility: The stronger the generation capability, the bigger the investment required in watermarking, moderation, and compliance — video is "higher-risk" than images because it persuades more and spreads farther.

Further Reading ​

References ​