Appearance
Knowledge Graphs and Knowledge Injection
The One-Line Definition
A knowledge graph (KG) is a graph-structured database that describes world knowledge with "entity — relation — property" triples, giving AI a structured, reasoning-ready representation of knowledge.
The one-line version: a knowledge graph turns "what the world looks like" into a net of connected points and lines — the points are entities (people, companies, drugs, proteins), the lines are relations (works at, treats, inhibits), and each point can carry properties (founding year, molecular formula). A vector database stores "similar meaning"; a knowledge graph stores "verified fact." The former finds approximate answers by similarity; the latter derives certain answers by walking relations. Each covers different ground, and they complement each other (comparison below, and in Vector Databases and Semantic Search).
Why does this matter? Because the weaknesses of large language models (LLMs) are exactly the strengths of knowledge graphs: hallucination, staleness, untraceable answers, no multi-hop reasoning. A knowledge graph provides facts that are traceable, updatable, and logically inferable, while the LLM understands natural language and tolerates the graph's noise. From Google's 2012 knowledge panel in search results to Microsoft's 2024 GraphRAG bringing "global questions" to private documents, this symbolic line has never been broken — and it is now merging deeply with neural networks.
Symbolism: entities/relations/logic → precise, explainable, hard to build
Connectionism: vectors/attention/params → flexible, strong generalization, hallucinates
Knowledge graph × LLM = the two routes converge: precision × flexibility1. Basic Concepts: A World Model in Graph Form
1. Nodes, Edges, Properties
A graph consists of three things:
| Element | Alias | Corresponds to | Examples |
|---|---|---|---|
| Node | Entity | An independent thing | Zhang Wei, Tsinghua University, Beijing, aspirin |
| Edge | Relation | A link between two entities | works at, graduated from, treats, located in |
| Property | Attribute | A descriptive value attached to an entity | Zhang Wei's birth year, a company's founding date |
Nodes and edges both have types: node types (person, company, drug) are called concepts, and edge types (works at, treats) are called predicates. What governs the types is the ontology (see below). The biggest difference between a knowledge graph and an "ordinary graph database" is that it explicitly declares types and semantics — so it isn't just a data structure, it's a "representation of knowledge."
2. Triples: The Minimal Unit of Knowledge
The basic storage unit of a graph is the triple — (subject, predicate, object) — equivalent to a directed edge:
(Zhang Wei) ──[works at]──▶ (DeepIntelligence Inc.)
│ │
│[graduated from] │[headquartered in]
▼ ▼
(Tsinghua University) ────[located in]────▶ (Beijing)Written as triples:
<Zhang Wei> <graduated from> <Tsinghua University>
<Tsinghua University> <located in> <Beijing>
<Zhang Wei> <works at> <DeepIntelligence Inc.>
<DeepIntelligence Inc.> <headquartered in> <Beijing>From just these four triples you can answer "Where does Zhang Wei work?", "Which city is Tsinghua in?", "Where is the headquarters of the company Zhang Wei works for?" — that last one is two-hop reasoning: Zhang Wei → DeepIntelligence Inc. → Beijing. This is the core difference from vector retrieval: a vector store can only answer "what is this similar to"; a graph can answer "how do you get from A to B in N steps."
3. The Ontology: Rules for the Graph
The ontology is the knowledge graph's schema layer, declaring "which concepts exist in the world and which relations are allowed between them." A graph without an ontology is just a pile of nodes and edges; only with an ontology can you call it "knowledge."
class Person {
properties: name, date of birth, nationality;
relations: works at → Company, graduated from → School;
}
class Company {
properties: name, founding year;
relations: headquartered in → City, hires → Person;
}
class Drug { properties: formula, indications; relations: treats → Disease; }The ontology buys you three things: consistency constraints (a person shouldn't "graduate from" a city), automatic inference (A works at B, B is a multinational, therefore A works for a multinational), and cross-store alignment (different systems can only be merged if they follow the same ontology). A field-tested lesson: keep the ontology simple before making it complete — start with a lightweight ontology covering 80% of the business and evolve it; projects that try to model "everything in the universe" on day one usually die in the modeling phase.
4. RDF and Property Graphs: Two Storage Paradigms
| Paradigm | Representatives | Query language | Traits |
|---|---|---|---|
| RDF (Resource Description Framework) | Jena, Virtuoso, RDFLib | SPARQL | W3C standard, built for open-data interoperability, everything is a triple |
| Property graph | Neo4j, NebulaGraph, JanusGraph | Cypher / nGQL / Gremlin | Properties embedded on nodes/edges, strong engineering performance, the industry mainstream |
RDF (Resource Description Framework) is the W3C's open data standard, emphasizing semantic interoperability — RDF data from anywhere in the world can be joined with one query language (SPARQL). Wikidata, which underlies Wikipedia (over 110 million items as of 2024), is the most famous open RDF graph. The property graph is the engineer's camp: types and properties live directly on nodes and edges, with better query performance and more convenient modeling, and Neo4j's Cypher is the de facto industrial standard.
2. Knowledge Graph vs. Vector Database: Symbols vs. Vectors
The two are constantly compared, because both tackle "how do machines understand knowledge." But their underlying logic is exactly opposite:
| Dimension | Knowledge graph | Vector database |
|---|---|---|
| Knowledge form | Symbolic triples, facts are exact | Dense vectors, continuous semantics |
| Precision | High: a fact is right or wrong, and verifiable | Low: ranked by similarity, can miss the question |
| Explainability | High: answers trace back to a concrete path | Low: similar vectors can't explain "why" |
| Fuzzy matching | Weak: say "Lao Wang" and it won't match "Wang Jianguo" | Strong: paraphrases and colloquialisms all hit |
| Multi-hop reasoning | Strong: follow relation chains N steps, provably | Weak: no compositional constraints or logical jumps |
| Build cost | High: ontology design, extraction, cleaning | Low: chunk documents + embed |
| Update cost | High: edge edits need consistency management | Low: recompute the vectors |
| Typical scenarios | Anti-fraud, medicine, enterprise knowledge, KGQA | Semantic search, similar-item recommendation, RAG recall |
The one-line verdict
Adopt a knowledge graph only if you have a hard requirement for "exact constraints + multi-hop reasoning." If all you need is "find similar by semantics," ride the vector store all the way. The golden combo: the vector store handles fuzzy recall, the graph refines and reasons over the entities inside the recalled results — which is exactly the logic underneath GraphRAG.
3. Knowledge Graphs × LLMs: Four Integration Patterns
This section is the flagship of neuro-symbolic fusion: let the neural network (the LLM), strong at statistical perception, and the symbolic system (the graph), strong at logical inference, each do what it does best — the LLM handles natural-language understanding and absorbs the graph's noise, the graph provides a reasoning-ready, traceable skeleton of facts. Four patterns, ordered from loose to tight coupling:
This is the heart of this page. Graphs and LLMs are not substitutes but mutual patches: the LLM patches the graph's construction and Q&A; the graph patches the LLM's precision and traceability.
| Pattern | Direction | One-line description | Typical techniques |
|---|---|---|---|
| GraphRAG | Graph → retrieval | Build documents into a graph, retrieve along it, feed the LLM | Community summaries, graph traversal recall |
| Knowledge injection | Graph → LLM | Compile graph knowledge into prompts or training corpora | Subgraph serialization, KG-corpus pretraining |
| KGQA | Natural language → graph query | Translate the question into Cypher/SPARQL and execute on the graph DB | Text2Cypher, semantic parsing |
| LLM-assisted construction | LLM → graph | Extract entities and relations from unstructured text to build the graph | Information extraction (IE), entity linking |
1. KG-Enhanced Retrieval: GraphRAG
GraphRAG is the graph-flavored variant of RAG, proposed by Microsoft Research in April 2024 (paper: From Local to Global: A Graph RAG Approach to Query-Focused Summarization). Traditional RAG chunks the text and retrieves chunks, so it can only answer "local" questions; GraphRAG first builds the whole corpus into a knowledge graph, then uses community detection to cluster the graph into "topic communities" and generate summaries; when a question arrives, it locates the relevant communities and synthesizes an answer.
Document set → LLM extracts entities/relations → knowledge graph
→ community detection (e.g. Leiden) → community summaries
→ user question → locate relevant communities → global synthesized answerThe pain point it solves is corpus-level questions (query-focused summarization): "How do the character relationships evolve across this novel?" — questions like this are scattered through the whole book; no single chunk can answer them, while graph structure naturally gathers the dispersed entity relations. Field experience: GraphRAG's indexing cost is high (LLM extraction on every chunk costs money), so it fits private knowledge bases with a hard requirement for global understanding; if you just want "document Q&A," traditional RAG is more economical. To build one yourself, see Build a RAG App from Scratch.
2. Knowledge Injection: Into the LLM's Input or Parameters
Knowledge injection has two routes, both boiling down to "turn graph triples into something the LLM can consume":
- Prompt-side injection: serialize the query-relevant subgraph into text and append it to the prompt. For example, after retrieving the local subgraph for "Zhang Wei," generate
[Zhang Wei] -works at-> [DeepIntelligence Inc.] -headquartered in-> [Beijing]and hand it to the LLM. Pros: zero training, explainable, graph updates take effect immediately. Cons: context length is finite, so subgraphs must be curated. This is prompt engineering in the field. - Parameter-side injection: convert triples to text (e.g.
"Zhang Wei graduated from Tsinghua University") and add them to the corpus for pretraining or fine-tuning, so the knowledge "grows into" the parameters. Classic examples from 2019: K-BERT and ERNIE fed graph knowledge into BERT-family models; in the LLM era this evolved into using graph data for Fine-Tuning and PEFT (LoRA). The cost: knowledge is frozen into parameters and hard to update — once a fact changes you must re-fine-tune or override it via prompts.
Common misconception
Parameter-side injection is not a guarantee. Parametric knowledge still conflicts, forgets, and hallucinates. The robust approach is layering: common, stable knowledge goes into parameters; dynamic, precise, strongly-constrained knowledge goes into prompt injection or GraphRAG — treat the parameters as a "scratchpad" and the graph as the "official archive."
3. Graph Q&A (KGQA): From Natural Language to Graph Queries
Knowledge Graph Question Answering (KGQA) lets users ask the graph in natural language while the system handles the translation into graph queries. Two mainstream routes:
Route A (semantic parsing): question → Text2Cypher/SPARQL → execute on the graph DB → results
Route B (retrieve-and-read): question → retrieve subgraph → serialize subgraph → LLM generates the answerRoute A is precise but brittle: one wrong syntax token or wrong field in NL2Cypher and the whole thing falls apart. Route B is more robust (the LLM reads a "graph snippet" rather than being forced to write a query), but graph retrieval quality caps its ceiling. The engineering recommendation is a hybrid: let the LLM first extract entities and intent from the question, use graph retrieval to recall candidate subgraphs, then let the LLM generate the answer from the subgraph — this downgrades "write a query" to "read the graph and speak," sharply narrowing the hallucination surface. KGQA is also a staple of knowledge-graph interview questions; review it alongside the interview question bank.
4. LLM-Assisted Construction: Automated Entity/Relation Extraction
Traditional graph construction relied on manual labeling — expensive and slow; LLMs turn information extraction into a "prompt task." A typical pipeline includes:
- Entity recognition: find named entities in the text — people, organizations, drugs, diseases;
- Relation extraction: decide what relation holds between two entities (works at, treats, inhibits...);
- Entity linking: align "Lao Wang" to the canonical entity "Wang Jianguo" (disambiguation) — the make-or-break step of automated construction;
- Deduplication and denoising: the same entity under multiple spellings and hallucinated triples from the LLM all need cleaning;
- Graph merging: fold incremental triples into the existing graph and check consistency with the ontology.
The number-one trap in LLM-based construction is hallucinated triples — the model writes "possible" as "fact." Countermeasures: multi-model voting, cross-validation, whitelisting allowed relation types, and manual sampling audits. Which is why automated construction ≠ maintenance-free — the cleaning pipeline is never optional (failure stories in Common Pitfalls and Anti-Patterns).
4. Classic Case Studies
1. The Google Knowledge Graph
In May 2012, Google launched the Google Knowledge Graph, showing entity cards beside search results: "Einstein's birthplace, children, awards." The phrasing of then-VP of Engineering Amit Singhal became the famous motto — "things, not strings." Google used it to upgrade search from "matching keywords" to "understanding entities," and it later supplied the knowledge foundation for Google Assistant. Around 2020, Google disclosed that its graph held roughly 500 million entities and over 35 billion facts. This case also directly seeded the graph usage in AI search — read it alongside Graphs in AI Search.
2. Medical Knowledge Graphs
Medicine is the most fertile soil for knowledge graphs: the relation network among diseases, symptoms, drugs, genes, and targets is a natural graph, and the field demands extreme precision. Representative projects:
- UMLS (Unified Medical Language System): the super-ontology maintained by the US National Library of Medicine, covering 1M+ concepts;
- Drug–target–disease graphs: pharmaceutical companies use them for drug repurposing — finding "candidate drugs connected to known indication pathways" on the graph;
- Integrated Chinese–Western medicine knowledge graphs: mapping the relations among symptoms, syndrome patterns, formulas, and herbs to support prescription assistance and clinical decision-making.
The lesson from medical graphs: a domain graph = a strong ontology + human review, because one wrong triple can point straight to a wrong treatment plan.
3. Enterprise Knowledge Management
Internal enterprise knowledge is highly fragmented: wikis, email, tickets, code comments, and the heads of departed colleagues. An enterprise knowledge graph models all of these sources into one "people–projects–documents–customers–systems" network, supporting:
- Intelligent Q&A: a new hire asks "what have we delivered for customer X" — the graph walks project → customer → deliverable and returns a sourced answer;
- Expert location: "who knows the payment logic of that legacy system" — find the people connected to that system on the graph;
- Knowledge lineage: derivation relations among documents, code, models, and datasets, supporting compliance audits.
This is also the knowledge foundation for putting agents into production: for an AI agent to act reliably, it needs more than fluency — it needs to know where the facts live and whom to trust.
5. Application Scenarios
| Scenario | What the graph does | Key capability |
|---|---|---|
| Q&A augmentation | Multi-hop reasoning: company A acquired B — who is B's CEO? | Jump 2–5 steps along relation chains |
| Recommendation | User–item–category–brand graph; explain recommendations via paths | Meta-paths, graph embeddings (KGAT etc.) |
| Risk control | Account–device–phone–IP connectivity graph; spot fraud rings | Connected subgraphs, graph propagation |
| Research literature | Protein structures, citation networks, paper–method–dataset links | Graph queries + subgraph summaries |
| AIOps | Service–dependency–alert–owner topology; root-cause analysis | Impact analysis, causal chains |
- Q&A augmentation: vanilla RAG handles "who" and "what" reasonably well but falls apart on multi-hop compositional questions (glaringly obvious during evaluation). Only with a graph layered on top does a triple-constrained query like "find all engineers who joined in 2019, graduated from Tsinghua, and currently work at the company" get a reliable answer.
- Recommendation: knowledge-graph-based recommendation uses external knowledge about users and items for features and explanations. The classic work KGAT (KDD 2019) models high-order connectivity with graph attention. For how recommendation evolved in the LLM era, see Recommendation Systems in the LLM Era.
- Risk control: anti-fraud graphs connect seemingly unrelated accounts through "shared devices, shared phone numbers"; a ring's connected subgraph jumps out at a glance. Link mining and community detection are home turf for graph algorithms.
- Research literature: protein structure databases are essentially giant "molecule–structure–function" graphs. The AlphaFold Protein Structure Database holds roughly 200 million predicted structures, and the protein-interaction network behind it is the knowledge graph's close cousin — see AlphaFold and AI for Science.
6. Engineering Challenges: Building the Graph Is Hard; Keeping It Alive Is Harder
| Challenge | Manifestation | Mitigation |
|---|---|---|
| Build cost | Manual ontology design is expensive; LLM auto-extraction brings hallucination and noise | Start with a lightweight ontology; multi-model cross-validation + spot checks |
| Maintenance and updates | Entities vanish and relations change (people leave, companies rename) | Incremental update pipeline + freshness annotations + scheduled crawls to reconcile |
| Scaling | Storage and query latency at hundreds of millions of nodes and billions of edges | Graph DB sharding, graph indexes, caching, batch graph algorithms |
| Quality evaluation | Can triple accuracy, coverage, and consistency be quantified? | Sampled labeling audits + graph consistency checkers |
Three rules of thumb:
- Small before big: build a million-element graph in a single business domain and close the loop before talking about company-wide integration — "one big unified graph" is a project graveyard;
- Decouple ontology from data: ontology changes must migrate data automatically, or one type change means reworking the whole graph;
- Stamp knowledge with a shelf life: attach "founding year, data source, last updated" to triples, so dynamic knowledge can keep up with reality.
7. Trade-offs and Common Misconceptions
The trade-off checklist
- Precision vs. coverage: prefer a missing fact to a wrong one — an absent fact can be patched by an LLM, a wrong fact poisons everything downstream;
- Manual vs. automated construction: humans vet the core ontology and key entities; the long tail goes to LLM extraction plus spot checks;
- Parameters vs. prompt injection: stable common knowledge into parameters, dynamic precise knowledge into prompts — route by update frequency;
- Graph vs. vectors: split duties by "precise/fuzzy" rather than either-or (full comparison in Vector Databases and Semantic Search).
Common misconceptions
| Misconception | Reality |
|---|---|
| "Knowledge graphs are dead" | Quite the opposite — GraphRAG brought them back in the RAG era |
| "With a vector store, graphs are unnecessary" | Vector stores can't do constrained queries or multi-hop reasoning |
| "LLM-extracted graphs are ready to use" | Hallucinated triples and failed entity disambiguation are the norm; cleaning is mandatory |
| "A graph = a visualization dashboard" | Visualization is just the shell; querying, reasoning, and updating are the core |
| "KGQA is just wrapping an LLM" | Graph retrieval quality and query translation decide success; both need grinding |
The one-line verdict
Ask two questions before adopting a knowledge graph: Does the business demand traceable, precise facts? Does it need multi-step reasoning along relation chains? Two "yes" answers make the graph worth it; otherwise vector store + RAG is faster and cheaper.
Further Reading
- RAG — GraphRAG is the graph variant of RAG; understand RAG before the graph version
- Vector Databases and Semantic Search — the graph's "twin," in charge of fuzzy recall
- Prompt Engineering — concrete techniques for knowledge-injection prompts
- Fine-Tuning and PEFT (LoRA) — the other road: baking graph knowledge into parameters
- Large Language Models (LLM) — the protagonist of all four integration patterns
- AI Agents — reliable agent action rests on a precise knowledge foundation
- Graphs in AI Search — graph retrieval in Perplexity and AI search
- Recommendation Systems in the LLM Era — knowledge-graph-enhanced recommendation in production
- AlphaFold and AI for Science — the protein structure database, the graph's close cousin
- Build a RAG App from Scratch — get GraphRAG running hands-on
- Common Pitfalls and Anti-Patterns — construction and retrieval failure stories
- Glossary — quick reference for entity, ontology, triple, and related terms
References
- Edge et al. From Local to Global: A Graph RAG Approach to Query-Focused Summarization (arXiv:2404.16130, 2024) — the original GraphRAG paper, Microsoft Research
- Microsoft Research blog: GraphRAG: Unlocking LLM discovery on narrative private data (2024) — the GraphRAG announcement and design motivation
- Google: Introducing the Knowledge Graph: things, not strings (2012) — the official Google Knowledge Graph launch blog post
- W3C: Resource Description Framework (RDF) — the RDF data model standard
- W3C: SPARQL 1.1 Query Language — the RDF graph query language spec
- Wikidata official site — the world's largest open knowledge graph
- Neo4j official documentation — the property graph database and the Cypher query language
- Wang et al. KGAT: Knowledge Graph Attention Network for Recommendation (KDD 2019) — the classic work on knowledge-graph-enhanced recommendation