Skip to content

Case Study: Dify

At a glance A teardown of Dify's harness design — packaging the agent loop, the RAG pipeline, tool plugins, and a model-adaptation layer into "self-hostable middleware" — and where it sits on the spectrum between LangGraph (a code framework) and Coze (hosted SaaS).

This page contains time-sensitive content. Data is current as of 2026-08; job listings, pricing, and product details may have changed — verify against the original source before citing.

Case Study: Dify ​

An LLM application development platform open-sourced in April 2023, founded by Zhang Luyu (formerly a product manager at Tencent's CODING); the name stands for "Do It For You." With over 150,000 GitHub stars, it ranks among the most-starred open-source LLM application platforms today. It represents a third shape of the harness: not a code framework, and not hosted SaaS, but agent middleware with a graphical interface that you can move into your own data center.

What It Is: Middleware Between Models and Applications ​

Let's position Dify in this site's terms. In What Is an Agent Harness? we defined a harness as "the complete software system, beyond the model itself, that turns a model into an agent that can actually get work done." Most of the case studies in this series — Claude Code, SWE-agent, LangGraph — are harnesses built by developers writing code. Dify starts from a different question: it assumes that most teams putting LLMs into production don't want — and shouldn't have — to write that system from scratch themselves.

Dify's official self-description is an "open-source LLM app development platform," and it ships every component of the harness as an out-of-the-box product:

  • Agent loop and orchestration: two modes on a visual canvas — Workflow and Agent (dissected below);
  • Context engineering: a Prompt IDE, variable injection, and conversation-history management, all surfaced in the UI;
  • Tool system: 50+ built-in tools (Google Search, DALL·E, WolframAlpha, and more) plus a plugin marketplace;
  • RAG pipeline: a built-in pipeline spanning document ingestion (text extraction from PDF, PPT, and other formats) through retrieval;
  • Observability (LLMOps): application logs, annotation, and continuous improvement driven by production data;
  • Model adaptation layer: hundreds of models from dozens of inference providers, including any OpenAI-API-compatible endpoint;
  • Backend-as-a-Service: everything above comes with an API — once an application is published, it is a callable backend service.

In other words, Dify takes the component diagram from Anatomy of the Harness and productizes the whole thing. Using Dify essentially means renting (or self-hosting) a harness someone else wrote, compressing your own work down to three things: orchestrating the flow, writing prompts, and hooking up data.

In one sentence

If LangGraph gives you the primitives for building a harness (State/Node/Edge), Dify gives you the finished harness itself — database, queues, admin console, and permission system included — spun up with a single docker compose command.

Two Orchestration Modes: Workflow and Agent ​

Two orchestration paradigms coexist on Dify's canvas, and the split is no accidental product decision: it maps almost word-for-word onto the famous distinction from Anthropic's "Building effective agents" — workflow (the LLM is orchestrated along a predefined path) versus agent (the LLM dynamically directs its own process).

text
Two orchestration paradigms on the Dify canvas:

  Workflow (Chatflow)                  Agent / Agent node
 ┌──────────┐                         ┌──────────────┐
 │ Start    │                         │     LLM      │
 │    ↓     │                         │    ↓      ↑  │
 │ Knowledge│   Nodes run in the      │ Think → Tool │  The loop's direction
 │ retrieval│   human-drawn topology, │    ↓      ↑  │  is decided by the
 │    ↓     │   in sequence; models   │ Observe →    │  model at runtime
 │ LLM      │   work inside nodes,    │ Think again  │
 │ generate │   never route           │              │
 │    ↓     │                         │  Function    │
 │ Branch   │  → human-preset forks   │  Calling or  │
 │    ↓     │                         │  ReAct       │
 │ End      │                         └──────────────┘
 └──────────┘
  Topology fixed by humans:           Loop topology improvised by
  controllable, auditable             the model: flexible

This comparison is isomorphic to the "controlled loop vs. free loop" tension in the LangGraph case study; only the medium differs — LangGraph expresses the spectrum in code, Dify turns it into two buttons on a canvas. In practice, Dify users most often land on a hybrid architecture: an outer Workflow runs the deterministic flow, with an Agent node embedded in the middle to handle unpredictable subtasks — the layered idea that "the graph governs what can be governed, and the loop handles what can't be predicted."

One evolution detail is worth noting: Dify's flagship feature at launch in 2023 was knowledge-base Q&A (RAG support chatbots); Workflow became the core abstraction only in 2024, with the Agent node added inside Workflow afterward. That path is a miniature history of harness design in itself: from a DAG (retrieve → generate), to an explicit flowchart, to carving a controlled "cage" for the free loop inside the flowchart.

Built-in RAG: Making the Most Fraught Component a Default ​

Dify's founding skill is RAG, and it remains the number-one reason people adopt the platform. It turned RAG from "an architecture you have to assemble" into "a feature you check a box for":

  • Ingestion: built-in text extraction for PDF, PPT, and other common formats; chunking, cleaning, and indexing are all configured in the UI;
  • Retrieval: vector search, full-text search, hybrid search, and reranking as options; the knowledge base is a node type you drag straight onto the Workflow;
  • Operations: hit testing, inline editing of chunk content, and continuous tuning based on annotation data — this is RAG through an LLMOps lens, not a one-off script.

Viewed against Context Engineering: RAG is the single largest engineering effort in "provisioning context for the model," and it is the part most likely to be abandoned half-finished in a DIY harness (extraction quality, chunking strategy, retrieval evaluation — every one is a pitfall). Dify's bet: rather than have a thousand teams each write their own mediocre RAG, the platform ships one above the passing grade built in. The bet holds for RAG, and for the harness as a whole — it is the reason the entire product exists.

The Plugin System and MCP: Two Leaps in the Tool Layer ​

Dify's tool layer has been through one major architectural overhaul, and understanding it is key to understanding the boundaries of a "platform-style harness."

Phase one (February 2025, v1.0): the plugin architecture. Model providers and tools were carved out of the main repository and migrated into independently distributed plugins, with new plugin types such as Agent Strategy and Extension, plus an official Marketplace. The official blog states the motivation plainly: the core repository can't possibly carry adapter code for every model and tool in the world; for the platform to survive, the ecosystem has to grow itself. This is the rite of passage for every platform-style harness — the adaptation layer's workload is destined to outlast the core team's development capacity.

Phase two (July 2025, v1.6): built-in two-way MCP. The community had previously hooked into MCP through third-party plugins (the MCP SSE plugin, for instance); v1.6 made it a native capability — and a two-way one:

  • Consuming MCP: tools from any MCP server can be called inside a Workflow — either placed in an Agent node for the model to pick dynamically (the dynamic path) or wired in as a standalone node at a fixed position (the precise path);
  • Exposing MCP: conversely, Dify applications and Workflows can be published as MCP servers, for external agents like Claude Code to call.

The meaning of two-way MCP deserves a pause: Dify no longer just swallows external capabilities — the flows it orchestrates have themselves become tools inside someone else's harness. Platforms are starting to interoperate over a single shared protocol — the clearest signal yet that "harnesses are becoming middleware."

Model Neutrality: Making Model Swaps Routine ​

Dify connects to hundreds of models from dozens of providers and works with any OpenAI-API-compatible endpoint (including local inference via vLLM, Ollama, and similar). Within Dify, the model is a configuration item you can swap at runtime: the same application changes engines with a provider switch in the admin console, and the Prompt IDE supports comparing outputs from multiple models side by side.

This is a live demonstration of the boundary this site keeps drawing: the model is not part of the harness (see Model vs. Harness). Dify's business model is built precisely on that boundary — it binds to no model vendor and instead lives off "helping you swap models anytime." For enterprises, this neutrality also carries compliance weight: confidential data can switch to a locally deployed open-source model without changing a line of the flow or application layer.

Deployment Options: Cloud, Community Edition, and Enterprise Edition ​

Dify itself practices what the spectrum preaches — one codebase, three delivery modes:

  • Dify Cloud: the official hosted version, zero installation, with a free Sandbox plan — good for kicking the tires and prototyping. Features largely match the self-hosted edition.
  • Community Edition: the self-hosted open-source edition — what most people mean by "Dify." One docker compose up -d spins it up, with a floor of just 2 CPU cores and 4 GB of RAM; for heavy production use, the community has contributed Helm Charts (Kubernetes), Terraform (Azure/GCP), AWS CDK, and a full set of deployment options.
  • Enterprise Edition: enterprise-grade capabilities layered on top of Community (stronger security and administration, for example), under a commercial license — exactly the monetization outlet of the license's additional terms.

This three-tier structure descends directly from the open-source commercialization playbook of the GitLab and Mattermost generation: the open-source edition takes developers and on-prem deployments, the hosted edition captures teams who dread operations, and the enterprise edition monetizes compliance and scale. The lesson for harness research: once the harness itself becomes a commodity, "open-source core plus paid hosting" may be a more sustainable business model than the model itself — models depreciate, while operations, compliance, and ecosystem stickiness depreciate far more slowly.

The Platform Spectrum: LangGraph / Dify / Coze ​

Put the three "platform-type" entries from the case-study chapter side by side and a clear spectrum emerges — one problem (how to get agents into production without rewriting the harness), three delivery forms:

DimensionLangGraph (code framework)Dify (self-hosted middleware)Coze (hosted SaaS)
DeliverablePython/TS libraryA complete system deployed via DockerA cloud account
UsersEngineers writing codeEngineers + business users dragging on a canvasMostly business users
Control-flow expressionGraphs in codeWorkflow/Agent on a canvasWorkflows on a canvas
Where data and runtime liveYour processYour data center / VPCThe vendor's cloud
Customization ceilingNone (it's just code)Plugins + APIs; the core is unmodifiableWithin platform limits
Ops burdenYours to carryYours to carry (PostgreSQL/Redis/vector store and a whole stack)Zero
Lock-in riskLow (Apache 2.0 codebase)Medium (license addendum, see below)High (data and flows live in someone else's hands)
Typical buyerAI companies with strong engineering teamsEnterprises and regulated organizations that must self-hostTeams and individuals optimizing for speed

No point on the line is better or worse; what changes is the trade-off between control and operational burden as you move along the spectrum. Dify sits in the middle: it spares you LangGraph's "build the system yourself" workload while keeping Coze's non-negotiable — your data and runtime stay on your own turf.

Enterprise Self-Hosting: Why Dify Specifically ​

Dify's penetration among government-affiliated organizations, financial institutions, and manufacturers runs well ahead of what its technical novelty would explain. The reasons aren't on the feature list; they're several very practical structural factors:

  1. "Data never leaves the domain" is a hard constraint. Compliance requirements for private deployment rule out hosted SaaS outright, and building a framework in-house is too expensive — "self-hostable finished middleware" slots exactly into that gap. Docker Compose runs on a floor of 2 cores and 4 GB of RAM, and the community has contributed Helm Charts, Terraform, AWS CDK, and a complete production deployment toolkit.
  2. Business people need to be in the loop. In enterprises, knowledge of prompts and processes lives with domain experts, not engineers. The canvas lets the business side edit flows directly while engineers step back to "integrate systems and write plugins" — the collaboration structure alone is worth a lot.
  3. API-first, embeds into existing systems. Every application publishes as a REST API the moment it ships. Dify plays the role of a shared AI capability layer: portals, approval systems, and support systems call its APIs; it never competes for the frontend.
  4. Friendly to Chinese domestic models. First-class support for China-based model providers and local inference endpoints, combined with a Chinese-language community and documentation, leaves it with virtually no like-for-like rival in environments built around a domestic Chinese tech stack.

License caveat: not pure Apache 2.0

Dify ships under the "Dify Open Source License" — Apache 2.0 with two key restrictions added (which is why GitHub flags it as NOASSERTION rather than a standard OSI license):

  1. Multi-tenancy restriction: without written permission, you may not operate a multi-tenant environment externally using the Dify source code (in Dify terms, one workspace equals one tenant). In other words: internal enterprise use, or backing your own applications, is fine; turning it into a SaaS product serving multiple customers requires a commercial license.
  2. Frontend branding preservation: when using the Dify frontend, you may not remove or modify its logo and copyright notices.

For the typical enterprise self-hosting scenario, both clauses are essentially invisible; for teams building a paid service on top of Dify, this is the first document to read before signing anything.

Where It Fits — and Where It Doesn't ​

Where Dify shines: knowledge-base Q&A and customer support (RAG comes standard out of the box); internal process automation (approvals, routing, report generation — Workflow scenarios with deterministic steps); teams that need business people directly in the orchestration; government-affiliated and regulated organizations with private-deployment compliance requirements; and teams that want to validate many ideas fast, using Dify as an AI prototyping factory.

Where it grates: open-ended exploratory coding-agent tasks — the canvas wasn't designed for "the model improvises freely," and the Claude Code-style free loop actually has no room to move inside Dify; teams needing extreme customization of harness behavior — the platform's core loop can't be changed, and changing it hurts more than writing your own with LangGraph; minimal-embedding scenarios — standing up the full PostgreSQL + Redis + vector database dependency stack for a single LLM call is delivering a parcel with an aircraft carrier.

A quick decision rule

When choosing, ask yourself: is my bottleneck "I can't write a harness" or "I can't afford the operations"? Can't write it → Dify (or Coze). Can't afford it → a library like LangGraph, and you raise the loop yourself. Can do both → go back to Build Your Own: A Minimal Harness — that may be the answer that fits best.

Further Reading ​

References ​