Skip to content

Prompt Engineering in Practice

At a glance om prompt template design patterns to structured output, from systematic prompt-tuning methodology to version management and regression testing, to the "prompt vs. fine-tuning vs. RAG" decision tree: turning prompt engineering from "black magic" into reproducible, measurable engineering capability.

Prompt Engineering in Practice ​

Prompt engineering isn't "writing a pretty prompt." It's a controlled experiment design around the model's context window: you design the input distribution, observe the output distribution, and lock in improvements through regression testing.

Prompts are the cheapest "programmable interface" for LLM applications. This article answers three questions: what validated template patterns exist? how to tune prompts systematically rather than by feel? and when to stop tuning prompts and switch to fine-tuning or RAG? The whole piece assumes you already understand the basics of prompting (roles, few-shot, chain-of-thought) and the decoding parameters from Inference Fundamentals.

Prompt engineering is "easy to start, hard to master"

You can write a few lines and get the model to work, but getting the model to consistently output the way you want is engineering. The core discipline of this page: prompt changes without regression testing are like code changes without commit messages — neither can be rolled back or attributed.

One, Prompt Engineering's Position in the System ​

About 80% of failures in a typical LLM application come from "outside the prompt" (data, context, evaluation, routing), yet 80% of people's first reaction is "just tweak the prompt." Set the decision order first:

Problem arises
    │
    ▼
① Prompt design / fix (cheapest, do first)  ──── Topic of this page
    ▼
② Context supply (retrieval, tools, multi-turn memory) ── RAG / Agent
    ▼
③ Model upgrade (stronger foundation) ──────────────── Model selection
    ▼
④ Fine-tuning (style / format / domain adaptation) ── Fine-tuning in Practice

Each level is an order of magnitude more expensive than the previous. Never touch the model if a prompt can solve it — this principle will be expanded further in the decision tree below.

Two, Prompt Template Design Patterns ​

Here are seven patterns validated by public practice. Any pattern may work or fail — the key is judging with evaluation, not gut feeling (methods in Section Four).

PatternOne-line DescriptionApplicable ScenariosExample Skeleton
Role-playingGive the model an identity and stanceUnified style, professional toneYou are a senior X expert...
Instruction + ConstraintsClear task + boundaries / formatMost tasksPlease do X, don't do Y, output...
Few-shot ExamplesProvide 2~5 input-output examplesNeed to mimic format / toneInput:...\nOutput:...
Chain-of-Thought (CoT)Ask to reason before answeringMath / logic / multi-step tasksPlease think step by step
Tree-of-Thought (ToT)Explicitly explore multiple paths and self-evaluateComplex planning, search-type tasksScore branches and backtrack
Self-ReflectionLet the model check and correct its own outputQ&A, summarization, code reviewCheck the above answer, point out errors and fix them
Structured Output TemplateConstrain output shape via template / JSONProgrammatic consumption of outputOutput as JSON: {"conclusion":...}

1. Instruction + Constraints: Most Basic and Most Overlooked ​

text
Task: Classify the following user feedback into [complaint, suggestion, inquiry, other], and output one sentence of reasoning.
Constraints:
- Only output JSON, no extra text
- Reasoning must be under 20 characters
- If unsure, classify as "other"
User feedback: Your app crashed again, I'm so mad!

The essence of constraints is narrowing the output distribution. Observation shows: expanding "constraints" from one sentence to four (format / length / boundary / fallback) significantly improves stable output ratio — but constraints should be moderate. Overly long prompts dilute attention (see Section Six failure modes).

2. Few-shot Examples: The Strongest "Capability Plug-in" ​

Few-shot (in-context learning) lets the model imitate new formats and tones without updating parameters. Three key points:

  • Examples must cover edge cases: Good examples = 2 normal + 1 edge case + 1 counterexample.
  • Examples must match real input format exactly: Including punctuation, newlines, prefixes.
  • Order-sensitive: The model has the strongest tendency to imitate the most recent examples. Put hard cases last.
text
Please rewrite the following sentences in a more polite version.

Sentence: Give me the report.
Rewrite: Could you please send me a copy of the report? Thank you.

Sentence: This bug is way too basic.
Rewrite: This bug is a bit unexpected, could you please confirm the root cause?

Sentence: Hurry up.
Rewrite:

3. Chain-of-Thought (CoT): Let the Model "Show Its Work" ​

Let's think step by step-style prompts have been widely validated for their improvement on multi-step reasoning tasks (GSM8K math problems, etc.) (References). Engineering-style CoT writing:

text
Please answer the question below. Before giving the final answer, first write out your reasoning steps:
1. List known conditions
2. Derive step by step
3. Give the final answer (marked with ①)

Two engineering details:

  • Separate "reasoning" from "answer": Have the model output reasoning:... \n answer:..., so you can take only the answer at the application layer while keeping auditable reasoning.
  • Self-consistency: Sample N times for the same question (higher temperature), then majority-vote the answers. This significantly improves accuracy (exchanging 3~5x inference cost for ~10~20% relative improvement, numbers vary by task). This is the simplest form of "inference-time compute," see the test-time scaling discussion in Scaling Laws.

Boundaries of chain-of-thought

CoT works for tasks that "require multi-step logic," but fails for tasks the "model simply won't know" (high hallucination risk) — it cannot create knowledge from nothing. For knowledge questions, go through retrieval (RAG in Practice).

4. System Prompt Design Guidelines ​

The system prompt is "the constitution of the conversation": it's injected every request, and models typically comply with the system section at a higher rate than the user section (most models give the system more weight). Four design guidelines:

GuidelinePracticeAnti-pattern
Define role and boundariesOne-line positioning + clear "what NOT to do"Just writing "You are an assistant"
Front-load key instructionsPut the most important constraints at the startBuried after line 8
Fewer rules, more examplesFormat expressed via examples is more stable than rulesTwenty "pleaseessential..." directives → prefer examples over repetitive rules
Fixed version, version-managedSystem prompts are code — go through review and regressionChanging casually in a chat box

System prompts are an attack surface

Content in system prompts (especially permissions, tool descriptions, internal info) will be attempted to be extracted via prompt injection. Never put secrets or internal system details in a system prompt; keep untrusted inputs isolated, see Safety & Risks.

Three, Structured Output: Getting the Model into Your Program ​

An LLM's output must be parsable by a program — this is a hard requirement for production deployment. Three routes, from weak to strong:

RouteApproachReliabilityImplementation Cost
Prompt constraints + manual parsingSpecify JSON format in prompt, json.loads + fallback retry in codeLow~medium (~90% or less)Low
JSON modeAPI forces valid JSON output (schema not guaranteed)Medium (~99% valid JSON)Medium
Structured output / function callingAPI generates per given JSON Schema, or declare functions for model to fill parametersHigh (~99% schema-compliant)Medium

1. Guide JSON Output via Few-shot (No Platform Feature Required) ​

text
Convert the following email to JSON. Fields: type (inquiry/complaint/other), urgency (low/medium/high), summary.

Email: Hello, when will the refund I submitted yesterday be processed? Kind of urgent.
Output: {"type": "inquiry", "urgency": "high", "summary": "Asking about refund arrival time"}

Email: Your membership price is too expensive, the competitor charges half.
Output:

2. JSON Mode & Function Calling (OpenAI-Style API Example) ​

python
from openai import OpenAI

client = OpenAI()   # reads env var OPENAI_API_KEY

# Approach A: response_format forces valid JSON
resp = client.chat.completions.create(
    model="gpt-4o-mini",
    response_format={"type": "json_object"},
    messages=[
        {"role": "system", "content": "Extract information, output JSON."},
        {"role": "user", "content": "Email: Your order O-1024 has shipped, expected delivery in 3 days."},
    ],
)
print(resp.choices[0].message.content)
# {"order_id": "O-1024", "status": "shipped", "eta_days": 3}

# Approach B: Function calling, let the model "fill parameters" instead of "write format"
tools = [{
    "type": "function",
    "function": {
        "name": "extract_order_info",
        "description": "Extract order information from email",
        "parameters": {
            "type": "object",
            "properties": {
                "order_id": {"type": "string"},
                "status": {"type": "string", "enum": ["shipped", "pending", "cancelled"]},
            },
            "required": ["order_id", "status"],
        },
    },
}]
resp = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Email: Order O-1024 has shipped."}],
    tools=tools,
    tool_choice="auto",
)
print(resp.choices[0].message.tool_calls)   # Structured parameters, consumed directly by code

Don't hand-roll JSON parsing as a brute-force fix

Hand-rolled parsing + retry is fine for "demo projects," but in production format errors cause cascading exceptions. JSON mode / function calling are platform-guaranteed interfaces — use them first. Open-source self-hosted models can achieve equivalent effects with constraint decoding libraries like outlines, guidance, xgrammar (see Framework & Tool Selection).

Four, Systematic Prompt Tuning: Turning Black Magic into Experimentation ​

1. Five-Step Methodology ​

The most efficient path for tuning prompts isn't "change whatever comes to mind," but:

① Set baseline ── Run a minimal version first, record its metrics
② Form hypothesis ── Clearly state "I'm changing X to improve Y"
③ Isolate variable ── Change only one variable at a time
④ Measure ── Quantitative comparison on a fixed test set
⑤ Regress ── Confirm the new prompt doesn't break existing capabilities

Each step produces a reproducible artifact. Especially step ④: prompt optimization without evaluation is self-hypnosis. See Evaluations in Practice for building an evaluation system.

2. A Reproducible A/B Experiment ​

python
import json, random

# Fixed test set (golden set): structure = [{"input": ..., "expected": ...}, ...]
with open("test_cases.json") as f:
    test_cases = json.load(f)

def run_prompt(prompt_template: str, cases, sample_n=20) -> dict:
    """Given a prompt template, score on a test set subset, return accuracy"""
    random.seed(42)
    sample = random.sample(cases, min(sample_n, len(cases)))
    correct = 0
    for case in sample:
        prompt = prompt_template.replace("{{INPUT}}", case["input"])
        output = call_model(prompt)          # your own call wrapper
        if normalize(output) == normalize(case["expected"]):
            correct += 1
    return {"accuracy": correct / len(sample), "n": len(sample)}

baseline = run_prompt(TEMPLATE_V1)
candidate = run_prompt(TEMPLATE_V2)
print(f"v1: {baseline}  v2: {candidate}")
# Only switch if candidate is significantly higher than baseline; note random fluctuations with small samples

Minimum discipline for variable control

Change only one variable at a time. Mixing "added examples + changed wording + switched format" into a single change means you'll never know which change did what. Retest on the same test set after each change — comparing across test sets is meaningless.

3. Decoding Parameters Like Temperature Are Part of "the Prompt" ​

The same prompt behaves completely differently at temperature=0.7 vs temperature=0. Engineering best practice: fix decoding parameters first, then tune the prompt text. Write decoding parameters into experiment records — otherwise nothing is reproducible.

Five, Prompt Version Management ​

Prompts are code and must be version-managed:

ApproachPracticeProblem Solved
Code-based storagePrompts written into repo (.py/.json/.md), not scattered in chat windowsDiffable, rollable back
Version numbersEach prompt has v1, v2 or semantic version numbersAttribution, communication
Regression testingCritical prompts in CI: only merge if golden set passesPrevent "fix A, break B"
Change logRecord the hypothesis and result of each changeAvoid re-tripping over the same pit
text
# prompts/extract_order_info/v3.md
### Change Log
- v1: Initial version, few-shot + hand-rolled parsing. accuracy 0.82
- v2: Switched to JSON mode. accuracy 0.91 (fixed invalid JSON errors)
- v3: Added enum-constrained field status. accuracy 0.95 (fixed non-standard status values)

Boundary between prompts and code

A common conclusion from one post-incident review: "the prompt was quietly modified in a chat by someone." Managing prompts as code (review, test, rollback) is the watershed of LLM application engineering maturity.

Six, Typical Failures and Fix Reference Table ​

FailureRoot CauseFixExample
Output format unstable, occasional extra rambling textInsufficient constraints / not enough examplesAdd format constraints, few-shot, JSON modeSee Section Three
Refusals too frequentSystem prompt overly defensive / over-emphasizes safetyDistinguish "dangerous" from "sensitive" boundaries, tighten safety phrasingRelated discussion in Safety & Risks
Longer prompt = worse performanceContext dilution: key instructions drowned outKey instructions at start and end; delete redundancySee example below
Fabricated factsModel knowledge cutoff / weak long-tail knowledgeRetrieval-augmented generation (RAG), see RAG in Practice—
Few-shot performance fluctuatesPoor example quality/orderCover edge cases, hard cases lastSee Section Two
User input steers the model (prompt injection)Untrusted input not isolatedInput boundary isolation, output filteringSee Safety & Risks
Behavior changes drastically when switching modelsPrompt overfitted to a specific modelCross-model regression testing, use more robust constraint methodsSee Section Four

Context Dilution: The Cost of Ever-Longer Prompts ​

A typical anti-pattern example — piling twenty legacy rules into the system prompt, and the model ends up ignoring key instructions:

text
# ❌ Anti-pattern: Overly long system prompt, key instructions drowned in the middle
You are a customer service bot. Please be polite. Our refund policy is... (500 words)
Please classify user feedback as A/B/C. Please output in JSON...
text
# ✅ Fixed: Key instructions front-loaded + trimmed
Classify user feedback as A/B/C, output JSON, no extra text.
(Policy details go in a separate paragraph, moved to the end of the prompt when necessary, to avoid competing for attention with key instructions.)

"Put key instructions at the start and end" is an engineering application of the "attention focusing" principle from Prompting.

Engineering Defenses Against Prompt Injection ​

Prompt injection is when "a user or third-party text tries to rewrite your instructions." Tiered engineering defenses:

TierApproach
Input isolationSystem prompt / instructions and user content, retrieval content placed separately and clearly labeled
Content filteringDe-sensitize and strip instruction markers from retrieval content, web content
Output validationSchema/content validation on model output, critical actions require secondary confirmation
Principle of least privilegeTool permissions designed on least-privilege basis, critical operations require human confirmation

Principle: Treat "untrusted input" as code injection — never concatenate directly, never implicitly execute. Full threat model in Safety & Risks.

Seven, "Prompt vs. Fine-tuning vs. RAG" Decision Tree ​

When "the model doesn't meet requirements," judge in the following order without jumping ahead:

Model output doesn't match expectations?
├─ Knowledge/fact problem (doesn't know, fabricates) ──→ RAG (supplement external knowledge)
├─ Reasoning/format problem (calculates wrong, doesn't follow format) → Try prompt first (CoT, few-shot, structured output)
│   └─ Prompt at its limit (unstable, expensive) ──→ Consider fine-tuning (format/style/domain)
├─ Style/tone/domain language problem ──────────────→ Fine-tuning or more detailed few-shot
└─ Capability ceiling problem (model fundamentally can't) ──→ Switch to a stronger model (or accept the status quo)
DimensionPromptRAGFine-tuning
CostLowest (only inference tokens)Medium (retrieval pipeline + inference)Highest (training + inference)
TimelinessSeconds to updateMinutes to update (index rebuild)Requires re-training
Knowledge injectionWeak (limited context)Strong (any document)Strong but solidified
StabilityModel-dependent, drifts easilyDominated by retrieval qualityStable
Suitable forGeneral tasks, rapid iterationPrivate knowledge, time-sensitive knowledgeOutput format, domain style, task tone

Summary: Solve through input first (knowledge, examples, format) with prompts and RAG; only fine-tune when you need to change "the model's internal behavior" (style, domain-specific tasks). Specific costs and processes for fine-tuning are in Fine-tuning in Practice.

Further Reading ​

References ​