Skip to content

Skills & Knowledge Injection

At a glance The three ways agents acquire domain knowledge — static injection (system prompts and rules files), dynamic retrieval (RAG), and capability packages (Agent Skills) — with a close look at Anthropic's progressive disclosure design, rule precedence and conflict resolution, and how to pick the right approach for the job.

Skills & Knowledge Injection ​

Knowledge in model weights is frozen on the day training ends, and the model has never seen your project: it doesn't know your coding conventions, how your internal SDK works, or which command has to run before every deploy. One of a harness's core jobs is injecting the knowledge the model lacks but the task requires at runtime.

How you inject determines the cost structure of that knowledge: do you pay a token tax on every message, pay only on demand, or keep it standing by at zero cost? This page walks through the three mainstream approaches — static injection, dynamic retrieval, and capability packages (skills) — and how Agent Skills, launched by Anthropic in October 2025, used progressive disclosure to turn the third path into an open standard.

Why Knowledge Injection Matters ​

The naive idea is to stuff every piece of relevant knowledge into the system prompt. That road hits a wall fast:

  • The context budget is finite. The more you stuff in, the less room is left for conversation history and tool results — and latency and cost climb with every call. See Context Engineering.
  • Attention gets diluted. When irrelevant knowledge mixes with the task at hand, the model's adherence to key instructions degrades — a long context is not an effective context.
  • Different knowledge has different lifecycles. Coding conventions hold steady for six months, API docs refresh weekly, and an incident runbook only matters when something breaks. Injecting everything the same way guarantees waste, staleness, or both.

So the real engineering problem is: giving each kind of knowledge an injection channel with the right balance of cost and freshness.

The Three Approaches at a Glance ​

ApproachTypical carrierWhen it's injectedShape of the knowledgeBest for
Static injectionSystem prompt, CLAUDE.md, .cursor/rules/*.mdcSession start / every messageRules, conventions, personaSmall, stable constraints that apply every time
Dynamic retrieval (RAG)Vector store + retriever + assembly templateRetrieved live per queryVoluminous, fast-changing document factsCorpora too large for the context window, where freshness matters
Capability package (skill)SKILL.md + scripts/resources directoryMetadata always resident; body loaded on triggerProcedures, domain know-howStep-by-step procedural knowledge — idle most days, but you can't afford to lose it

The dividing line among the three is the shape of the knowledge: declarative constraints belong in static injection, vast quantities of facts belong in retrieval, and procedural workflows belong in a packaged skill. Each is covered in turn below.

Approach 1: Static Injection ​

System Prompts and Rules Files ​

The most direct injection: write rules into the system prompt, or into rules files in the repo that the harness reads into context at session start. Claude Code reads CLAUDE.md, Cursor reads .mdc files under .cursor/rules/, and the community has settled on the cross-tool AGENTS.md convention.

The semantics of CLAUDE.md are blunt: the file's contents load into context at every session start, serving as a project-level onboarding manual — build commands, directory conventions, testing discipline, prohibited actions. It supports multiple levels, from the project root to user scope, with the nearest one taking precedence.

Cursor's rules system goes a step further: each rule carries frontmatter declaring when it should be injected — already a primitive form of progressive disclosure:

mdc
---
description: "Guidelines for writing database migration files"
globs: "migrations/**/*.sql"
alwaysApply: false
---

All migrations must be reversible: every up migration ships with a matching down migration.
Never backfill data inside a migration; backfills run as a separate job.
Rule modeConfigurationBehavior
AlwaysalwaysApply: trueInjected into every conversation
Auto AttachedSet globsAuto-attached when a matching file appears in context
Agent RequestedSet description onlyThe model judges relevance and pulls it in itself
ManualNeitherThe user explicitly references it with @ in conversation

The inflation trap of static injection

alwaysApply: true and sprawling CLAUDE.md files are the most common misuse. Every "inject on every message" rule taxes every turn of the conversation, and the more rules you add, the lower the odds that any single one gets followed. Rule of thumb: keep resident rules to the scale of "what a new hire must remember on day one," and sink the rest into glob-triggered rules or skills.

When Static Injection Fits ​

  • The content is small (a few hundred tokens), changes slowly, and is almost always relevant → inject it statically.
  • It's strongly tied to specific files or directories → use a glob-triggered rule, not a globally resident one.
  • It's an operating procedure longer than a screen, useful only for particular tasks → it doesn't belong in static injection; see skills below.

Approach 2: Dynamic Retrieval (RAG) ​

When the volume of knowledge makes residency in context impossible — tens of thousands of pages of internal docs, the full ticket history, an API reference that updates constantly — retrieval-augmented generation (RAG) is the only option: chunk and index the corpus against the current task, then retrieve the most relevant fragments at runtime and stitch them into context.

For agents, RAG rarely appears as a one-shot lookup. It's usually wrapped as a tool (say, search_docs(query)) that the agent calls repeatedly inside its loop — what's known as agentic RAG. Retrieval quality hinges not on which vector store you pick, but on three things:

  1. Whether the chunking granularity aligns with the knowledge's natural boundaries (a function, a doc section — not fixed-length truncation);
  2. Whether recall and reranking hold up against an evaluation set;
  3. Whether assembly into context preserves provenance and structure, so the model can cite it and flag doubts.

The essential difference between RAG and rules files

Rules files inject instructions ("here's how you should act"), which the model is expected to obey; RAG injects evidence ("here's what the material says"), which the model must weigh for relevance and timeliness. Funneling both through one channel is a classic way for retrieval noise to become behavior noise.

For a fuller treatment of RAG, see Context Engineering and Memory. All you need here is its position: factual knowledge that answers "what is it" — large in volume, fast-changing, retrievable.

Approach 3: Capability Packages (Agent Skills) ​

From MCP to Skills ​

In October 2025, Anthropic released Agent Skills and soon after published it as an open standard (spec at agentskills.io), open-sourcing the official examples repo anthropics/skills alongside it, with built-in document-processing skills for docx, xlsx, pptx, pdf, and more. The format has since been adopted by several mainstream coding tools beyond Claude Code, including Cursor and GitHub Copilot / VS Code.

If MCP answers "what tools can the agent call," Skills answer "does the agent know how to do it": they package a domain's workflows, checklists, companion scripts, and templates into one portable folder. In its minimal form, a skill is just a Markdown file with frontmatter; in full form, it's a directory:

pdf-processing/
├── SKILL.md          # Required: frontmatter (metadata) + body (operating instructions)
├── scripts/          # Optional: executable scripts the agent runs rather than reads line by line
│   └── extract_tables.py
├── references/       # Optional: reference docs loaded into context on demand
│   └── FORMAT_SPEC.md
└── assets/           # Optional: templates and other resources used directly, never entering context
    └── report_template.docx

The Core Design: Progressive Disclosure ​

Skills' most elegant design choice is the loading strategy — nothing enters context all at once. Instead there are three tiers, cost rising at each level and depth available on demand:

┌──────────────────────────────────────────────────────────┐
│ Tier 1: Metadata (always resident)                       │
│   SKILL.md YAML frontmatter: name + description          │
│   Tens to ~100 tokens per skill; loaded in full          │
│   into the system prompt at startup                      │
│   Purpose: model knows "this skill exists, when to use"  │
├──────────────────────────────────────────────────────────┤
│ Tier 2: Body instructions (loaded on trigger)            │
│   SKILL.md Markdown body: workflows, checklists,         │
│   constraints                                            │
│   Read in only when the task matches the description;    │
│   typically several thousand tokens                      │
├──────────────────────────────────────────────────────────┤
│ Tier 3: Bundled resources (referenced on demand)         │
│   references/ docs: the model Reads the file itself      │
│     when it needs details                                │
│   scripts/ files: executed like commands —               │
│     the code never enters context, only the output does  │
│   assets/ templates: copied/filled in as files,          │
│     never entering context                               │
└──────────────────────────────────────────────────────────┘

The three tiers map neatly onto three costs of knowledge: metadata is the catalog index, cheap enough to keep fully resident; the body is the operating manual, opened only when needed; resources are the reference library — scripts never even have to be "read." Running python scripts/extract_tables.py beats having the model read the script into context and mentally simulate it: it saves tokens and it's more reliable.

A real SKILL.md looks like this (the field requirements in the frontmatter follow Anthropic's official authoring guide):

markdown
---
name: pdf-processing
description: Extract text and tables from PDFs, split and merge documents.
  Use when the user needs to work with PDF files, extract table data,
  or fill in PDF forms.
---

# PDF processing

## Quick start
To extract text: run `python scripts/extract_text.py <input.pdf>`.

## Table extraction
When you need to preserve table structure, use `scripts/extract_tables.py`,
which outputs CSV. For complex layout rules, see references/FORMAT_SPEC.md.

## Don'ts
- Don't screenshot pages and rely on visual recognition, unless the PDF is a scan.

Key fields:

  • name: kebab-case, and it must match the directory name;
  • description: what the model triggers on. It must spell out both what the skill does and when to use it (capped at 1024 characters, no XML tags allowed) — it isn't documentation; it's a fuzzy-matching trigger;
  • Claude Code also supports fields such as allowed-tools, which restrict the tools available while the skill is active, wired into the permission system.

A bad description means the skill doesn't exist

Whether the model activates a skill depends almost entirely on semantic matching between the description and the user's request. "Helps with document tasks" will never fire; "Use when the user asks to extract tables from PDFs, merge/split PDFs, or fill in PDF forms" will. When writing a skill, write the description first, validate the trigger rate against real tasks, then write the body.

Skills vs. Adjacent Concepts ​

  • vs tools/MCP: tools are capability interfaces ("what can be done"); skills are procedural knowledge ("how it should be done"). A skill often orchestrates several tools to carry out a workflow.
  • vs rules files: rules are constraints that apply every single time; a skill is a manual activated per task. The same team's deploy process, written as a rule, taxes every message; written as a skill, it stands by at zero cost.
  • vs subagents: a skill hands knowledge to the current agent; a subagent outsources the task together with its context. A skill can serve as a subagent's source of capability — the two are orthogonal.

Precedence and Conflict Resolution ​

Once knowledge comes from many sources, conflict is inevitable: the user message says "just edit the prod config," CLAUDE.md says "never edit prod directly," and the skill spells out a deploy procedure. The precedence order commonly used in practice (highest to lowest) is roughly:

  1. System prompt (built into the harness; defines identity and safety floors; users cannot override it);
  2. Permission policies (deterministic enforcement that bypasses model judgment; see Permissions & Human-in-the-Loop);
  3. The user's current-turn instructions (highest authority on what the task should be, but cannot punch through 1 or 2);
  4. Project-level rules (among CLAUDE.md and .cursor/rules, the rule closer to the current directory wins over the global one);
  5. User-level global rules;
  6. Skill bodies (procedural advice; they yield when they conflict with higher-level rules);
  7. Retrieved documents (evidentiary content, lowest priority, and they should carry source and timestamp).

Two practical recommendations:

  • Make conflicts explicit; don't let them silently overlap. Writing "when X and Y conflict, X wins" in the rules file is far more reliable than hoping the model guesses right.
  • Never leave safety-critical constraints in natural language alone. Rules files rely on the model complying, and models do err; real hard constraints belong in the permission layer (tool allowlists, path blocking), with natural-language rules acting only as guidance.

Decision Guide: Rules, RAG, or Skill ​

Before writing anything, ask three questions: how often does this knowledge change? How likely is it to be relevant to the current task? And is it a constraint, a fact, or a procedure?

ScenarioChoiceWhy
"Commit messages follow Conventional Commits"Rules file (resident)Short, stable, relevant every time
"Code changes under api/ must update the OpenAPI docs in step"Rules file (glob-triggered)Only relevant when touching specific files
"Our incident runbook runs three hundred pages"RAGHigh volume; pull fragments on demand
"The internal SDK's API docs are regenerated daily"RAG / tool-based lookupFreshness beats structure
"A release runs a checklist: cut a tag, update the changelog, canary 5%"SkillProcedural flow — zero cost at rest, complete when invoked
"Codify this 8-step data-cleaning workflow for the whole team"Skill (with scripts/)Portable, versionable, deterministic script execution
"Company security red line: no secrets in code"Hard block at the permission layer + a rules-file reminderNatural-language constraints alone aren't enough

In one sentence: constraints go into rules, facts go to retrieval, procedures become skills, and hard lines land in permissions.

What a Good Skill Looks Like ​

A counterexample (assembled from failure patterns common in the wild):

markdown
---
name: helper
description: A useful assistant skill that helps users with all kinds of tasks.
---

# Assistant
You can use me to help work with files. Try to do a good job.
There are plenty of other features — explore as needed.

The problems are plain: the description contains no trigger conditions, so the model will never activate it; the body is empty slogans with no executable steps; there are no scripts, so every step is left to model improvisation.

A good example:

markdown
---
name: changelog-release
description: Generate a release changelog for this repo and cut the git tag.
  Use when the user asks to "cut a release", "generate the changelog",
  or "prepare a release".
---

# Release process

1. Confirm the working tree is clean: `git status` shows no uncommitted
   changes; otherwise stop and ask the user.
2. Collect commits since the last tag: `scripts/commits_since_last_tag.sh`.
3. Classify by Conventional Commits (feat/fix/chore) and write them to the
   top of CHANGELOG.md, following references/CHANGELOG_FORMAT.md.
4. Version bump rule: any feat bumps minor; fixes only bump patch.
5. Show the changelog draft to the user and get confirmation
   before running `git tag`.

What good skills have in common:

  • The description is the trigger: state what it does + when to use it, including the actual words users would say;
  • The body is an executable process: numbered steps, ordered, with failure branches ("otherwise stop and ask") — not prose;
  • Deterministic work goes to scripts: anything that can live in scripts/ shouldn't be hand-rolled by the model; the model only makes decisions and chains steps together;
  • Details sink into references/: keep the body within a few thousand tokens; format specs and full checklists go to tier 3, loaded on demand;
  • There are human checkpoints: an explicit pause for confirmation before irreversible operations (tagging, releasing, writing to prod) — consistent with the thinking behind permission design.

Trade-offs ​

  • Skills aren't free. Each piece of metadata is small, but a few hundred descriptions add up to a meaningful slice of context and dilute one another's trigger precision. Keep the skill count matched to your task domain, and periodically prune skills that never fire.
  • Model-triggered invocation is inherently fuzzy. Description matching is a semantic judgment; it can miss (should have fired, didn't) or misfire (shouldn't have, did). Don't rely on automatic triggering alone for critical workflows — keep an explicit invocation path (such as /skill-name).
  • Static injection wins on determinism. Rules files are present on every message; there is no "forgot to trigger" failure mode. For constraints that must hold every single time, that's a feature, not waste.
  • Skills are a double-edged sword on security. Scripts inside a skill run with the agent's permissions, which makes third-party skills effectively supply-chain dependencies; before adopting an external skill, audit its scripts/ and body for prompt-injection risk.
  • Don't turn a skill into a second documentation library. If a "skill" body is mostly reference material with no operational procedure, it belongs in RAG, not in a trigger slot.

Further Reading ​

References ​