Theme
Prompt Engineering in Practice
Prompt engineering isn't "writing a pretty prompt." It's a controlled experiment design around the model's context window: you design the input distribution, observe the output distribution, and lock in improvements through regression testing.
Prompts are the cheapest "programmable interface" for LLM applications. This article answers three questions: what validated template patterns exist? how to tune prompts systematically rather than by feel? and when to stop tuning prompts and switch to fine-tuning or RAG? The whole piece assumes you already understand the basics of prompting (roles, few-shot, chain-of-thought) and the decoding parameters from Inference Fundamentals.
Prompt engineering is "easy to start, hard to master"
You can write a few lines and get the model to work, but getting the model to consistently output the way you want is engineering. The core discipline of this page: prompt changes without regression testing are like code changes without commit messages — neither can be rolled back or attributed.
One, Prompt Engineering's Position in the System
About 80% of failures in a typical LLM application come from "outside the prompt" (data, context, evaluation, routing), yet 80% of people's first reaction is "just tweak the prompt." Set the decision order first:
Problem arises
│
▼
① Prompt design / fix (cheapest, do first) ──── Topic of this page
▼
② Context supply (retrieval, tools, multi-turn memory) ── RAG / Agent
▼
③ Model upgrade (stronger foundation) ──────────────── Model selection
▼
④ Fine-tuning (style / format / domain adaptation) ── Fine-tuning in PracticeEach level is an order of magnitude more expensive than the previous. Never touch the model if a prompt can solve it — this principle will be expanded further in the decision tree below.
Two, Prompt Template Design Patterns
Here are seven patterns validated by public practice. Any pattern may work or fail — the key is judging with evaluation, not gut feeling (methods in Section Four).
| Pattern | One-line Description | Applicable Scenarios | Example Skeleton |
|---|---|---|---|
| Role-playing | Give the model an identity and stance | Unified style, professional tone | You are a senior X expert... |
| Instruction + Constraints | Clear task + boundaries / format | Most tasks | Please do X, don't do Y, output... |
| Few-shot Examples | Provide 2~5 input-output examples | Need to mimic format / tone | Input:...\nOutput:... |
| Chain-of-Thought (CoT) | Ask to reason before answering | Math / logic / multi-step tasks | Please think step by step |
| Tree-of-Thought (ToT) | Explicitly explore multiple paths and self-evaluate | Complex planning, search-type tasks | Score branches and backtrack |
| Self-Reflection | Let the model check and correct its own output | Q&A, summarization, code review | Check the above answer, point out errors and fix them |
| Structured Output Template | Constrain output shape via template / JSON | Programmatic consumption of output | Output as JSON: {"conclusion":...} |
1. Instruction + Constraints: Most Basic and Most Overlooked
text
Task: Classify the following user feedback into [complaint, suggestion, inquiry, other], and output one sentence of reasoning.
Constraints:
- Only output JSON, no extra text
- Reasoning must be under 20 characters
- If unsure, classify as "other"
User feedback: Your app crashed again, I'm so mad!The essence of constraints is narrowing the output distribution. Observation shows: expanding "constraints" from one sentence to four (format / length / boundary / fallback) significantly improves stable output ratio — but constraints should be moderate. Overly long prompts dilute attention (see Section Six failure modes).
2. Few-shot Examples: The Strongest "Capability Plug-in"
Few-shot (in-context learning) lets the model imitate new formats and tones without updating parameters. Three key points:
- Examples must cover edge cases: Good examples = 2 normal + 1 edge case + 1 counterexample.
- Examples must match real input format exactly: Including punctuation, newlines, prefixes.
- Order-sensitive: The model has the strongest tendency to imitate the most recent examples. Put hard cases last.
text
Please rewrite the following sentences in a more polite version.
Sentence: Give me the report.
Rewrite: Could you please send me a copy of the report? Thank you.
Sentence: This bug is way too basic.
Rewrite: This bug is a bit unexpected, could you please confirm the root cause?
Sentence: Hurry up.
Rewrite:3. Chain-of-Thought (CoT): Let the Model "Show Its Work"
Let's think step by step-style prompts have been widely validated for their improvement on multi-step reasoning tasks (GSM8K math problems, etc.) (References). Engineering-style CoT writing:
text
Please answer the question below. Before giving the final answer, first write out your reasoning steps:
1. List known conditions
2. Derive step by step
3. Give the final answer (marked with ①)Two engineering details:
- Separate "reasoning" from "answer": Have the model output
reasoning:... \n answer:..., so you can take only the answer at the application layer while keeping auditable reasoning. - Self-consistency: Sample N times for the same question (higher temperature), then majority-vote the answers. This significantly improves accuracy (exchanging 3~5x inference cost for ~10~20% relative improvement, numbers vary by task). This is the simplest form of "inference-time compute," see the test-time scaling discussion in Scaling Laws.
Boundaries of chain-of-thought
CoT works for tasks that "require multi-step logic," but fails for tasks the "model simply won't know" (high hallucination risk) — it cannot create knowledge from nothing. For knowledge questions, go through retrieval (RAG in Practice).
4. System Prompt Design Guidelines
The system prompt is "the constitution of the conversation": it's injected every request, and models typically comply with the system section at a higher rate than the user section (most models give the system more weight). Four design guidelines:
| Guideline | Practice | Anti-pattern |
|---|---|---|
| Define role and boundaries | One-line positioning + clear "what NOT to do" | Just writing "You are an assistant" |
| Front-load key instructions | Put the most important constraints at the start | Buried after line 8 |
| Fewer rules, more examples | Format expressed via examples is more stable than rules | Twenty "pleaseessential..." directives → prefer examples over repetitive rules |
| Fixed version, version-managed | System prompts are code — go through review and regression | Changing casually in a chat box |
System prompts are an attack surface
Content in system prompts (especially permissions, tool descriptions, internal info) will be attempted to be extracted via prompt injection. Never put secrets or internal system details in a system prompt; keep untrusted inputs isolated, see Safety & Risks.
Three, Structured Output: Getting the Model into Your Program
An LLM's output must be parsable by a program — this is a hard requirement for production deployment. Three routes, from weak to strong:
| Route | Approach | Reliability | Implementation Cost |
|---|---|---|---|
| Prompt constraints + manual parsing | Specify JSON format in prompt, json.loads + fallback retry in code | Low~medium (~90% or less) | Low |
| JSON mode | API forces valid JSON output (schema not guaranteed) | Medium (~99% valid JSON) | Medium |
| Structured output / function calling | API generates per given JSON Schema, or declare functions for model to fill parameters | High (~99% schema-compliant) | Medium |
1. Guide JSON Output via Few-shot (No Platform Feature Required)
text
Convert the following email to JSON. Fields: type (inquiry/complaint/other), urgency (low/medium/high), summary.
Email: Hello, when will the refund I submitted yesterday be processed? Kind of urgent.
Output: {"type": "inquiry", "urgency": "high", "summary": "Asking about refund arrival time"}
Email: Your membership price is too expensive, the competitor charges half.
Output:2. JSON Mode & Function Calling (OpenAI-Style API Example)
python
from openai import OpenAI
client = OpenAI() # reads env var OPENAI_API_KEY
# Approach A: response_format forces valid JSON
resp = client.chat.completions.create(
model="gpt-4o-mini",
response_format={"type": "json_object"},
messages=[
{"role": "system", "content": "Extract information, output JSON."},
{"role": "user", "content": "Email: Your order O-1024 has shipped, expected delivery in 3 days."},
],
)
print(resp.choices[0].message.content)
# {"order_id": "O-1024", "status": "shipped", "eta_days": 3}
# Approach B: Function calling, let the model "fill parameters" instead of "write format"
tools = [{
"type": "function",
"function": {
"name": "extract_order_info",
"description": "Extract order information from email",
"parameters": {
"type": "object",
"properties": {
"order_id": {"type": "string"},
"status": {"type": "string", "enum": ["shipped", "pending", "cancelled"]},
},
"required": ["order_id", "status"],
},
},
}]
resp = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Email: Order O-1024 has shipped."}],
tools=tools,
tool_choice="auto",
)
print(resp.choices[0].message.tool_calls) # Structured parameters, consumed directly by codeDon't hand-roll JSON parsing as a brute-force fix
Hand-rolled parsing + retry is fine for "demo projects," but in production format errors cause cascading exceptions. JSON mode / function calling are platform-guaranteed interfaces — use them first. Open-source self-hosted models can achieve equivalent effects with constraint decoding libraries like outlines, guidance, xgrammar (see Framework & Tool Selection).
Four, Systematic Prompt Tuning: Turning Black Magic into Experimentation
1. Five-Step Methodology
The most efficient path for tuning prompts isn't "change whatever comes to mind," but:
① Set baseline ── Run a minimal version first, record its metrics
② Form hypothesis ── Clearly state "I'm changing X to improve Y"
③ Isolate variable ── Change only one variable at a time
④ Measure ── Quantitative comparison on a fixed test set
⑤ Regress ── Confirm the new prompt doesn't break existing capabilitiesEach step produces a reproducible artifact. Especially step ④: prompt optimization without evaluation is self-hypnosis. See Evaluations in Practice for building an evaluation system.
2. A Reproducible A/B Experiment
python
import json, random
# Fixed test set (golden set): structure = [{"input": ..., "expected": ...}, ...]
with open("test_cases.json") as f:
test_cases = json.load(f)
def run_prompt(prompt_template: str, cases, sample_n=20) -> dict:
"""Given a prompt template, score on a test set subset, return accuracy"""
random.seed(42)
sample = random.sample(cases, min(sample_n, len(cases)))
correct = 0
for case in sample:
prompt = prompt_template.replace("{{INPUT}}", case["input"])
output = call_model(prompt) # your own call wrapper
if normalize(output) == normalize(case["expected"]):
correct += 1
return {"accuracy": correct / len(sample), "n": len(sample)}
baseline = run_prompt(TEMPLATE_V1)
candidate = run_prompt(TEMPLATE_V2)
print(f"v1: {baseline} v2: {candidate}")
# Only switch if candidate is significantly higher than baseline; note random fluctuations with small samplesMinimum discipline for variable control
Change only one variable at a time. Mixing "added examples + changed wording + switched format" into a single change means you'll never know which change did what. Retest on the same test set after each change — comparing across test sets is meaningless.
3. Decoding Parameters Like Temperature Are Part of "the Prompt"
The same prompt behaves completely differently at temperature=0.7 vs temperature=0. Engineering best practice: fix decoding parameters first, then tune the prompt text. Write decoding parameters into experiment records — otherwise nothing is reproducible.
Five, Prompt Version Management
Prompts are code and must be version-managed:
| Approach | Practice | Problem Solved |
|---|---|---|
| Code-based storage | Prompts written into repo (.py/.json/.md), not scattered in chat windows | Diffable, rollable back |
| Version numbers | Each prompt has v1, v2 or semantic version numbers | Attribution, communication |
| Regression testing | Critical prompts in CI: only merge if golden set passes | Prevent "fix A, break B" |
| Change log | Record the hypothesis and result of each change | Avoid re-tripping over the same pit |
text
# prompts/extract_order_info/v3.md
### Change Log
- v1: Initial version, few-shot + hand-rolled parsing. accuracy 0.82
- v2: Switched to JSON mode. accuracy 0.91 (fixed invalid JSON errors)
- v3: Added enum-constrained field status. accuracy 0.95 (fixed non-standard status values)Boundary between prompts and code
A common conclusion from one post-incident review: "the prompt was quietly modified in a chat by someone." Managing prompts as code (review, test, rollback) is the watershed of LLM application engineering maturity.
Six, Typical Failures and Fix Reference Table
| Failure | Root Cause | Fix | Example |
|---|---|---|---|
| Output format unstable, occasional extra rambling text | Insufficient constraints / not enough examples | Add format constraints, few-shot, JSON mode | See Section Three |
| Refusals too frequent | System prompt overly defensive / over-emphasizes safety | Distinguish "dangerous" from "sensitive" boundaries, tighten safety phrasing | Related discussion in Safety & Risks |
| Longer prompt = worse performance | Context dilution: key instructions drowned out | Key instructions at start and end; delete redundancy | See example below |
| Fabricated facts | Model knowledge cutoff / weak long-tail knowledge | Retrieval-augmented generation (RAG), see RAG in Practice | — |
| Few-shot performance fluctuates | Poor example quality/order | Cover edge cases, hard cases last | See Section Two |
| User input steers the model (prompt injection) | Untrusted input not isolated | Input boundary isolation, output filtering | See Safety & Risks |
| Behavior changes drastically when switching models | Prompt overfitted to a specific model | Cross-model regression testing, use more robust constraint methods | See Section Four |
Context Dilution: The Cost of Ever-Longer Prompts
A typical anti-pattern example — piling twenty legacy rules into the system prompt, and the model ends up ignoring key instructions:
text
# ❌ Anti-pattern: Overly long system prompt, key instructions drowned in the middle
You are a customer service bot. Please be polite. Our refund policy is... (500 words)
Please classify user feedback as A/B/C. Please output in JSON...text
# ✅ Fixed: Key instructions front-loaded + trimmed
Classify user feedback as A/B/C, output JSON, no extra text.
(Policy details go in a separate paragraph, moved to the end of the prompt when necessary, to avoid competing for attention with key instructions.)"Put key instructions at the start and end" is an engineering application of the "attention focusing" principle from Prompting.
Engineering Defenses Against Prompt Injection
Prompt injection is when "a user or third-party text tries to rewrite your instructions." Tiered engineering defenses:
| Tier | Approach |
|---|---|
| Input isolation | System prompt / instructions and user content, retrieval content placed separately and clearly labeled |
| Content filtering | De-sensitize and strip instruction markers from retrieval content, web content |
| Output validation | Schema/content validation on model output, critical actions require secondary confirmation |
| Principle of least privilege | Tool permissions designed on least-privilege basis, critical operations require human confirmation |
Principle: Treat "untrusted input" as code injection — never concatenate directly, never implicitly execute. Full threat model in Safety & Risks.
Seven, "Prompt vs. Fine-tuning vs. RAG" Decision Tree
When "the model doesn't meet requirements," judge in the following order without jumping ahead:
Model output doesn't match expectations?
├─ Knowledge/fact problem (doesn't know, fabricates) ──→ RAG (supplement external knowledge)
├─ Reasoning/format problem (calculates wrong, doesn't follow format) → Try prompt first (CoT, few-shot, structured output)
│ └─ Prompt at its limit (unstable, expensive) ──→ Consider fine-tuning (format/style/domain)
├─ Style/tone/domain language problem ──────────────→ Fine-tuning or more detailed few-shot
└─ Capability ceiling problem (model fundamentally can't) ──→ Switch to a stronger model (or accept the status quo)| Dimension | Prompt | RAG | Fine-tuning |
|---|---|---|---|
| Cost | Lowest (only inference tokens) | Medium (retrieval pipeline + inference) | Highest (training + inference) |
| Timeliness | Seconds to update | Minutes to update (index rebuild) | Requires re-training |
| Knowledge injection | Weak (limited context) | Strong (any document) | Strong but solidified |
| Stability | Model-dependent, drifts easily | Dominated by retrieval quality | Stable |
| Suitable for | General tasks, rapid iteration | Private knowledge, time-sensitive knowledge | Output format, domain style, task tone |
Summary: Solve through input first (knowledge, examples, format) with prompts and RAG; only fine-tune when you need to change "the model's internal behavior" (style, domain-specific tasks). Specific costs and processes for fine-tuning are in Fine-tuning in Practice.
Further Reading
- Prompting (Concepts) —— Prompt composition, zero-shot vs. few-shot, theoretical CoT
- Inference Fundamentals: Autoregression & Sampling —— How temperature, top-k, top-p fundamentally affect prompt effectiveness
- Evaluations in Practice —— Turning "is my prompt good" into quantifiable metrics
- RAG in Practice —— The complete path from prompt to retrieval for knowledge-type problems
- Fine-tuning in Practice: LoRA Full Process —— The next step after prompts hit their limit
- Common Pitfalls and Anti-Patterns —— Pitfalls related to prompting, like "the longer the prompt gets"
- Safety & Risks —— Defenses against prompt injection and jailbreaks
References
- OpenAI: Prompt engineering guide —— Official practice guide with six strategies
- Anthropic: Prompt engineering overview —— Prompt engineering docs for the Claude family
- Chain-of-Thought Prompting (arXiv:2201.11903) —— Original paper on chain-of-thought prompting
- Self-Consistency Improves Chain of Thought (arXiv:2203.11171) —— Original paper on self-consistency majority voting
- Tree of Thoughts (arXiv:2305.10601) —— Original paper on tree-of-thought prompting
- OpenAI: Function calling guide —— Official docs for structured tool calling
- OpenAI: Structured outputs guide —— Official docs for JSON Schema-forced output