Skip to content

Midjourney and Image Generation

At a glance Using Midjourney as an entry point to unpack the 2022-2023 text-to-image explosion: the diffusion-model and CLIP foundations behind the three-way race of DALL·E 2, Stable Diffusion, and Midjourney, the product iteration methodology, the open-source ecosystem, and the copyright controversy.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

Midjourney and Image Generation ​

In August 2022, the digital art category of the Colorado State Fair's annual art competition produced a bombshell: an entry titled Théâtre D'opéra Spatial, generated with Midjourney and then enhanced with Photoshop and Gigapixel AI by a contestant named Jason Allen, took first place. As the news spread, the debate over "whether AI art counts as art" broke out of niche circles, and for the first time truly pushed text-to-image technology in front of the general public.

If the rise of text-to-image is seen as an image-making revolution, Midjourney was one of its most closely watched products: it was not the first to launch (DALL·E 2 came earlier), nor the most technically open (Stable Diffusion is fully open source), but with a product methodology centered on "aesthetics" it turned AI image generation into a subscription business worth hundreds of millions of dollars. Using Midjourney as the entry point, this article retraces the full story of the 2022–2023 text-to-image explosion: the three-way industry landscape, the diffusion-model foundation, Midjourney's product iterations, and the revolution's profound impact on design, copyright, and the creator economy. For the underlying technical principles, start with Diffusion Models and Generative AI.

1. 2022–2023: The Three-Way Race in Text-to-Image ​

Text-to-image was not an idea born in 2022. As early as around 2015, image captioning and conditional generation models were experimenting at the pixel level, but the quality fell far short of usable. The real inflection point came after diffusion models matured in 2021–2022 — they pushed "turning text into images" from "barely recognizable outlines" to "indistinguishable from photographs."

2022 was the year that decided the industry landscape, with three players making their moves at the same time:

ProductVendorKey datesOpennessSignature strengths
DALL·E 2OpenAIReleased 2022.04Closed source, waitlist, pay per imageLeap in realism, strong semantic understanding
Stable DiffusionStability AIOpen-sourced 2022.08 (v1.4)Open weights, runs locallyLargest ecosystem, controllable and hackable
MidjourneyMidjourney Inc.Public beta 2022.07Subscription-based, used inside DiscordBest-in-class aesthetics, hottest community

The positioning differences among the three routes were stark:

  • DALL·E 2 stood for "research-driven": OpenAI spent more than a year polishing realism from DALL·E 1 (2021) to DALL·E 2; it stunned the industry on release, but closed source plus a waitlist meant the public could only watch from the sidelines;
  • Stable Diffusion stood for "open-source-driven": trained on the LAION-5B dataset (roughly 5.85 billion image-text pairs), with weights released under a relatively permissive license, anyone could download, modify, and deploy it locally — directly spawning a vast community ecosystem;
  • Midjourney stood for "product-driven": it published no research papers and had no open-source ambitions; instead it pushed "how good the results look" to the extreme and used the ultra-low friction of a Discord bot to capture the mass market.

Why did it explode precisely in 2022?

Three elements happened to mature at the same time in those two years: diffusion models completed their theoretical foundations in 2020–2021 (DDPM, DDIM, LDM), CLIP solved "text-image alignment" in 2021, and massive image-text datasets like LAION-5B plus compute at scale made "actually trainable" a reality. With technology, data, and compute converging, text-to-image finally moved from the lab to the masses. For the full arc of this evolution, see A Brief History of AI.

2. The Technical Foundation: How Diffusion Models Turn Text into Images ​

Whatever the product form, the technical foundation of the 2022 wave of text-to-image systems was remarkably uniform: a diffusion model + a text encoder + (optionally) control and customization plugins.

1. From DDPM to LDM: Denoising in Latent Space ​

The core idea of diffusion models is "break first, then learn to restore": during training, Gaussian noise is added to an image step by step until it becomes pure noise, while a neural network learns to denoise in reverse; at generation time, the model starts from random noise and "restores" an image that matches the training data distribution, step by step. DDPM (Denoising Diffusion Probabilistic Models, 2020) proved this path can produce high-quality images, but it had one fatal flaw — slowness: it required a thousand or more iterative steps in pixel space.

Latent Diffusion Models (LDM, 2022) solved it with a clever piece of engineering: run the diffusion in latent space instead of pixel space. An autoencoder first compresses a 512×512 image into a 64×64×4 latent (roughly a 48× compression), noise is added and removed in that low-dimensional space, and a decoder finally restores the latent back into an image. Compute drops by an order of magnitude — Stable Diffusion is precisely the productization of LDM, which is why it can run on consumer GPUs. For the fuller mathematics, see Diffusion Models and Generative AI.

2. CLIP: Aligning Text and Images ​

An image generator alone is not enough; the system also has to "understand" the prompts humans write. The key component here is CLIP (Contrastive Language-Image Pre-training), proposed by OpenAI in 2021. CLIP uses about 400 million image-text pairs for contrastive learning, mapping text and images into a single semantic space: the sentence "a fox sitting in the snow" and a corresponding photo should sit close together in that space, and far away from unrelated images.

Stable Diffusion 1.x and early Midjourney versions both borrowed CLIP's text encoder directly: the user's prompt is encoded into a semantic vector and injected into every layer of the diffusion model via cross-attention, so that the "denoising" process is always steered toward "matching the text." CLIP's limitations are equally obvious — a short context window (77 tokens), weak support for non-English languages such as Chinese, and limited understanding of abstract concepts (like "a melancholy dusk"). Those shortcomings became the starting point for the large body of prompt engineering tricks. For a fuller treatment of multimodal alignment, see Multimodal Models.

3. ControlNet and LoRA: From Gacha Pulls to Painting ​

Diffusion models are natural "random generators": with the same prompt, every output has a different composition, and early users jokingly dubbed this uncertainty "gacha pulls." Starting in 2023, two technical routes turned "gacha" into "painting":

  • ControlNet (2023.02): by "cloning and freezing the original UNet while training a trainable copy," it injects structural information — line art, depth maps, human pose skeletons, semantic segmentation maps — into the generation process. From then on, users could say "colorize according to this line drawing" or "keep this composition but change the style" — precise structural control brought text-to-image to the point where designers could actually use it;
  • LoRA (Low-Rank Adaptation): originally proposed as a parameter-efficient fine-tuning (PEFT) method for large language models, it updates only a small number of parameters using low-rank matrices. The community quickly ported it to text-to-image: with just a few dozen images, you can train a LoRA file of a few dozen to a few hundred MB that "customizes" a character, art style, or object into the model, and multiple LoRAs can even be stacked. To train your own, see Fine-Tuning and PEFT (LoRA) and Fine-Tuning in Practice.

The one-line takeaway

If you keep only one lesson: output quality is determined by the base model, while "usability" is determined by controllability. After 2023, the competitive focus of text-to-image was no longer "who paints prettier pictures" but "who follows your instructions more precisely."

3. Inside the Product: Why Midjourney Looks the Best ​

Midjourney has published no technical papers, so outsiders can only speculate about its architecture (the consensus is that it is based on the diffusion model family, with heavy use of fine-tuning and aesthetic data calibration). But its product methodology can be publicly reconstructed in three parts: Discord interaction, version iteration, and aesthetic tuning.

1. Building the Product Inside Discord ​

Midjourney is one of very few companies to build its product entirely inside Discord: users type the /imagine command followed by a prompt in a server, the bot returns four candidate images, and users interact with U1–U4 (upscale one) and V1–V4 (generate variations of one). The brilliance of this interaction lies in:

  • Zero installation cost: no app to download, no configuration to understand; five minutes with a community tutorial and you are up and running;
  • Community as traffic: everything anyone generates is public in shared channels by default, creating an endless waterfall of content — "browsing other people's generations" is itself a retention mechanism;
  • Free advertising and data: the massive volume of real feedback in public channels (which images get upscaled, which get liked) becomes a natural signal for model tuning.
text
/imagine prompt: a red fox in the snow, golden hour, photorealistic --ar 16:9 --v 6
# Common suffixes: --ar aspect ratio, --stylize stylization strength, --v version number, --seed random seed

Not without a cost

Discord interaction means prompts and images are visible to other users on the same server by default — hardly friendly to privacy-sensitive business users (Midjourney only later introduced Stealth Mode and a web app). The tension between "product choice" and "privacy" persists to this day.

2. Version Iterations: V1 → V6 → V7 ​

Midjourney's version history is the best window into "how text-to-image evolved" and a living textbook of "aesthetic tuning" (as of mid-2025):

VersionTime (approx.)Key changes
V12022.02Lab prototype, coarse grainy style
V22022.04Better composition and detail
V32022.07Public beta, heavy stylization, community takes off
V42022.11Big jump in realism, handles more complex prompts
V52023.03Improved hand rendering, near-photographic detail
V5.22023.06Stronger prompt understanding and style control
V62023.12Major gains in natural language understanding and realism
V6.12024.07Further refinement of skin and hair detail
V72025.04Personalization profiles, Draft Mode, speed and quality combined

Two turning points are worth noting: V3→V4 moved Midjourney from a "strongly stylized, illustration-like feel" toward "near-photographic texture," cementing its reputation as "the best-looking"; V6→V7 shifted the focus from "better image quality" to "understanding who you are" — V7 introduced personalization, letting users build their own profile through likes and favorites so the model's output fits their personal taste. That is already building a "user model."

3. Aesthetic Tuning: Midjourney's Core Competence ​

Midjourney's biggest difference from other models is that it deliberately biases the model toward "pretty": through carefully curated training data, human scoring, and fine-tuning, the model leans toward "aesthetically pleasing output" over "strict fidelity to the prompt." This directly fueled the rise of prompt engineering as a craft — with the same description, adding or omitting parameters like --ar 16:9, --stylize, or --v 6 makes a world of difference. For common prompt techniques, parameter explanations, and anti-patterns, see Prompt Engineering and The Prompt Playbook.

Why is Midjourney "good-looking" while SD is "hard to tune"?

A widely shared analogy: Midjourney bakes "taste" into the model, so users only need to describe the content; Stable Diffusion makes "taste" a knob the user turns themselves — the same prompt, with a different base model, a different sampler, or a different CFG value, produces wildly different styles. The former is effortless but less controllable; the latter is flexible but has a high barrier. It also explains why one became a mass-market product and the other the darling of the technical community.

4. The Stable Diffusion Ecosystem: An Open-Source World ​

If Midjourney is "one company polishing one product," Stable Diffusion is "one set of weights setting a whole world on fire." After the SD 1.4 weights were open-sourced, a three-layer ecosystem grew around them:

  1. Weights and model marketplaces: platforms like Civitai and Hugging Face host a massive collection of community fine-tuned models — distributed as .ckpt / .safetensors files and split into "base models" (such as SDXL, DreamShaper, Realistic Vision) and "LoRAs" (characters, styles, objects), covering nearly every art style from anime to photorealistic portraits. For a model quick reference, see Models and Leaderboards;
  2. Toolchains: WebUI (AUTOMATIC1111) turned SD into a visual interface; ComfyUI exposes the entire generation pipeline (text-to-image, image-to-image, ControlNet, LoRA, upscaling, masked inpainting) as editable, reusable node-graph "workflows," becoming the standard for professional users and automated production;
  3. Production pipeline integration: designers wired SD/ComfyUI into Photoshop, Blender, and game asset pipelines, batch-producing images with a "ControlNet for composition + upscale and refine + LoRA to hold the style" assembly line.

The cost of the open-source ecosystem unfolds in the next section: controllability brings complexity, and freely available weights bring abuse risk. For a deeper discussion of open-source vs closed-source trade-offs, see Common Pitfalls and Anti-Patterns.

5. Applications and Impact: From Toy to Productivity ​

Text-to-image was still "just for fun" in 2022; by 2024 it had permeated production workflows across multiple industries:

  • Design: "drawing from scratch" became "prompt a draft → curate → refine," letting one brief produce multiple versions in batch; e-commerce uses image-to-image to swap backgrounds, change scenes, and generate model shots for product photos;
  • Game concept art: concept designers use text-to-image to mass-produce character and scene candidates, then integrate and refine them by hand, dramatically shortening early exploration cycles;
  • Photography and retouching: restoring old photos and conjuring missing detail out of nothing; some commercial shoots have started replacing real studio sessions with AI "digital models";
  • The prompt economy: prompt marketplaces (PromptBase and the like), prompt engineer job listings, and prompt-focused tutorials and courses have proliferated — around 2023, "knowing how to write prompts" was briefly hailed as "the copywriting of the new era." As models' understanding improves, the value of prompts is shifting from "mystical incantation" back to "plain expressive ability."

Just as impossible to ignore is the dark side: the copyright and training-data controversy is the single biggest landmine in the text-to-image industry. Stable Diffusion was trained on LAION-5B, scraped from the internet, which contains large amounts of copyrighted images — triggering two landmark lawsuits:

  • Getty Images v. Stability AI (2023): Getty alleged that the model output images bearing traces of the Getty watermark, amounting to direct infringement;
  • A class action by three artists (2023): claiming the model "learned" their art styles without permission.

The core legal questions — "does training itself constitute infringement" and "are outputs derivative works" — remain unresolved. The industry has responded along two paths: "licensed data" (Adobe Firefly, for instance, trains only on licensed images) and "AI content labeling/watermarking." For a more systematic view of safety and compliance, see AI Safety and Governance.

6. Milestone: The Colorado State Fair Incident ​

Back to the opening story. In August 2022, Théâtre D'opéra Spatial — generated by Jason Allen with Midjourney, then enhanced with Photoshop and Gigapixel AI — won first place in the Colorado State Fair's digital art category, and the judges had no idea the piece came from AI. Allen later disclosed his creation process voluntarily, and the controversy exploded:

  • Supporters argued: "brush or prompt, both are simply choices of creative tool" — and Allen did indeed go through extensive prompt selection and post-editing; it was hardly "one-click output";
  • Critics argued: AI "digested" the styles of countless human artists during training, so the award was unfair competition against human creators;
  • Neutral observers pointed out: the rules themselves were not ready for AI tools — once the definition of "what makes an artist" starts to loosen, contest rules and copyright law both need rewriting.

The aftermath of the controversy lasted a long time. In September 2023, the U.S. Copyright Office rejected Allen's copyright application for the piece on the grounds of "insufficient human authorship"; subsequent appeals were repeatedly denied — as of mid-2025, the work still has not received copyright protection. This incident, together with the 2023 "Zarya of the Dawn" comic case (the AI-generated images were denied protection while the text and arrangement were protected), forms the early case law for the legal puzzle of "can AI-generated content hold copyright" — and it is the best case study for understanding the "authenticity crisis" in AI Safety and Governance.

7. Trade-offs ​

After using a few generations of text-to-image products, practitioners keep running into four trade-offs:

  • Open-source self-hosting vs closed-source subscription: SD is free and controllable, but you have to run your own GPUs, safety filters, and version upgrades; Midjourney works out of the box with consistent quality, but the underlying stack is out of your control and your data leaves your environment. Team decisions are fundamentally about "cost structure," not technical preference;
  • Quality vs controllability: Midjourney is "good-looking" by default but hard to control precisely; SD + ControlNet can reproduce a composition exactly but demands endless parameter tuning. Whether the project calls for "style exploration" or "precise delivery" decides which route to take;
  • Generation vs memorization: models may "memorize" images from the training set, creating copyright and plagiarism risks (automatic metrics cannot even tell generation from copying), so human review is required before delivering commercial work;
  • Speed vs quality: for the same model, sampling in 20 steps vs 50 steps, or using the base model vs a distilled one, is an explicit trade between time and quality. For inference-side optimization, see Inference Optimization and Quantization.

One practical recommendation

For personal learning: play with Midjourney first to build "aesthetic intuition," then learn SD + ComfyUI to gain "control," and finally use LoRA to customize your own style — this path saves the most time. Whenever terminology is unclear, consult the Glossary.

Further Reading ​

References ​