Appearance
Fine-Tuning and PEFT (LoRA)
Fine-tuning (FT) means continuing to train a pretrained model on labeled data from a specific task or domain so it adapts to a new objective; parameter-efficient fine-tuning (PEFT) is a training paradigm that reaches similar results by updating only a small subset of parameters (or even by adding just a few new ones).
Pretrained models (Llama, Qwen, DeepSeek, etc.) picked up general language ability from trillion-scale corpora — they are "generalists who know a little about everything." But the model doesn't know whether you want JSON output or customer-service phrasing, doesn't know your industry jargon, and has no idea where your company's knowledge base lives. Fine-tuning is the step that turns the generalist into a specialist — it doesn't touch the underlying language ability itself, but adjusts the model's "behavioral habits" and "output preferences." Almost every major LLM product you can name today (ChatGPT, Claude, DeepSeek) went through the full path of "pretraining → post-training (including fine-tuning)"; see Large Language Models (LLM) and A Brief History.
1. Why Fine-Tuning: Three Paths from Generalist to Specialist
An LLM that already speaks and reasons becomes one that "answers your domain questions in your format" through exactly three mainstream routes:
| Approach | What it does | Changes model parameters? | Typical cost | Best for |
|---|---|---|---|---|
| Prompt engineering | Spell out rules, examples, and persona in the input | No | Nearly zero | Quick validation, rule-based tasks, things the model can already do |
| Retrieval-Augmented Generation (RAG) | Hook up an external knowledge base and stitch retrieved content into the context | No | Building the index + retrieval pipeline | Knowledge-intensive work, facts that must be fresh and accurate, content that needs to be traceable |
| Fine-tuning (FT) | Keep training on task data to change model behavior/format/style | Yes | GPU training cost | Fixed output formats, a distinctive house voice, behavior problems prompts can't fix |
Rule of thumb: fine-tuning is for the problems prompting can't fix. If the model knows how but answers wrong → adjust the prompt; if the answer needs external facts → bring in RAG; if the format just won't come out right, the voice never sounds like yours, or everything won't fit into a prompt → then consider fine-tuning. The three are not mutually exclusive; the most common production combination is "fine-tuning for format + RAG for knowledge." For a detailed comparison, see Prompt Engineering and Retrieval-Augmented Generation (RAG).
A Practical Screening Framework
Ask three questions: ① Is the model capable enough (if not → move to a bigger model or add knowledge)? ② Is it a "what to know" problem (→ RAG) or a "how to talk" problem (→ fine-tuning)? ③ Can prompting cover it (if so → don't train yet)? Run a baseline first; talk fine-tuning later.
2. The Fine-Tuning Spectrum: From Full Fine-Tuning to QLoRA
"Fine-tuning" is a big umbrella; inside it stretches a full spectrum from "touching every parameter" to "touching only a few matrices." The two key metrics are how many parameters get updated (which drives training memory and storage) and how close the result gets to full fine-tuning.
| Approach | Full name | What gets updated | Relative memory | Training speed | Quality reference | Typical use |
|---|---|---|---|---|---|---|
| Full fine-tuning (FFT) | Full Fine-Tuning | All model parameters | Baseline (highest) | Slowest | 100% (the upper-bound benchmark) | Top teams with ample data and compute |
| Layer-wise FT | Layer-wise FT | Unfreeze only the last few layers / certain modules | Slightly lower | Slightly faster | Close to full FT (when the task maps to top-layer semantics) | Legacy approach, rarely used today |
| Adapter tuning | Adapter Tuning | Insert small networks between Transformer layers and train only those | Down ~1/2 | Faster | Close to full FT on single tasks | Multi-task with a shared backbone (one small parameter set per task) |
| LoRA | Low-Rank Adaptation | Freeze the backbone; train only the injected low-rank matrices (~0.1%–1% of parameters) | Down ~2/3 | Fast | ≈ full FT on most tasks, slightly lower on a few | Today's mainstream default choice |
| QLoRA | Quantized LoRA | Quantize the backbone to 4-bit and freeze it, then train the LoRA matrices | Another ~1/3 down (fine-tunes 70B-class models on a single GPU) | Fast | On par with LoRA | Consumer GPUs / personal research / tight budgets |
How to Read This Table
The far right of the spectrum doesn't mean "best quality" — only "lowest cost." Full fine-tuning remains the reference point for the quality ceiling, but fully fine-tuning a 7B model needs roughly 4×14GB of memory (Adam states and gradients all take memory), whereas QLoRA at the same scale runs on a single 16GB consumer GPU. For the underlying principles and engineering trade-offs, see Inference Optimization and Quantization (the quantization idea) and Deployment and Inference Optimization in Practice.
One more word on adapters: each Transformer layer gets a small Bottleneck network inserted (project down, then project back up), which adds one extra computation at inference time. The upside is that many tasks can share the same frozen backbone, each attaching its own small parameter set — this made adapters quite popular for multi-task settings. But LoRA wins across the board on both quality and implementation simplicity, so today the rule is basically "if LoRA works, skip Adapters."
3. How LoRA Works: Splitting One Big Update into Two Small Matrices
LoRA (Low-Rank Adaptation) comes from the 2021 paper LoRA: Low-Rank Adaptation of Large Language Models. It rests on a key observation: in full fine-tuning, the weight update itself can be approximated with a very low rank — pretrained weights W encode vast general knowledge, and task adaptation only needs to perturb them along a few small directions.
3.1 Low-Rank Decomposition
Suppose a linear layer (say, an attention projection matrix) has original weights W of shape d×d. Full fine-tuning learns the complete update ΔW (also d×d — hundreds of millions of parameters); LoRA instead forces ΔW to decompose into the product of two small matrices:
ΔW ≈ B × A (A: d×r, B: r×d, with r ≪ d)
Pretrained weights W (frozen) Delta branch (only this is trained)
┌─────────────────┐ ┌─────────┐ ┌─────────┐
│ W : d×d │ + │ B: d×r │×│ A: r×d │
│ (kept as-is) │ │ │ │ │
└─────────────────┘ └─────────┘ └─────────┘
↓ mergeable at forward/inference time
W' = W + (alpha / r) × B × ADuring training only A and B are updated, so the parameter count drops from d×d to d×r + r×d. With d=4096 and r=8, the delta parameters are only about 1/256 of full fine-tuning. A is initialized with Gaussian random values, B with zeros — so at the start of training B×A=0, the model behaves exactly like the pretrained model, and training is more stable.
python
import torch
import torch.nn as nn
import torch.nn.functional as F
class LoRALinear(nn.Module):
"""Inject a low-rank branch into a linear layer: freeze the main weight, train only A and B"""
def __init__(self, in_dim: int, out_dim: int, r: int = 8, alpha: int = 16):
super().__init__()
self.weight = nn.Parameter(torch.randn(out_dim, in_dim) * 0.02) # simulates pretrained weights
self.weight.requires_grad = False # freeze
self.A = nn.Parameter(torch.randn(in_dim, r) * 0.01) # Gaussian init
self.B = nn.Parameter(torch.zeros(r, out_dim)) # zero init
self.scaling = alpha / r
def forward(self, x):
base = F.linear(x, self.weight) # frozen backbone: plain forward pass
delta = (x @ self.A) @ self.B * self.scaling # low-rank delta: only this is trained
return base + deltaThe above is a minimal sketch of the idea; in production just use the Hugging Face PEFT library: LoraConfig(r=8, lora_alpha=16, target_modules=["q_proj","v_proj"]) injects LoRA into the specified modules. For the complete engineering workflow, see Fine-Tune Your Own LLM.
3.2 Choosing the Rank r
| r value | Delta parameter share (7B model reference) | Characteristics |
|---|---|---|
| r=1–4 | Tiny | Very few update directions; suits ultra-light adaptation on very little data; risk of underfitting |
| r=8 | ~0.1% | The safe default for most tasks |
| r=16–32 | ~0.2%–0.4% | More stable for complex tasks with larger datasets; higher quality ceiling |
| r=64+ | Large | Clearly diminishing returns; the parameter count approaches full fine-tuning, which defeats the purpose |
Rule of thumb: start small (r=8), double and compare once the pipeline runs (16/32), and don't jump straight to 64. Doubling r grows memory and storage roughly linearly, while quality gains usually flatten after 8→16.
Zero Overhead at Inference
The A and B matrices from a finished LoRA run can be merged back into the backbone before inference: W_merged = W + scaling × (B @ A). After merging, the model architecture is identical to the original, with zero extra inference latency or memory (for quantized merging, see Inference Optimization and Quantization). You can also skip merging and store A and B separately — one LoRA weight set per task (usually tens to a few hundred MB), mounted on demand. That is exactly why LoRA is so widely used in multi-task and multi-tenant scenarios.
4. Training Data: The Data Science of Instruction Tuning
4.1 Anatomy of Instruction Data (SFT Data)
The most mainstream form of fine-tuning is instruction tuning, also called supervised fine-tuning (SFT): you feed the model triplets of "instruction + input + expected output" so it learns to follow instructions and act on them. This is the paradigm that kicked off the fine-tuning boom after ChatGPT launched, originating from the InstructGPT paper.
A typical instruction sample looks like this:
json
[
{
"instruction": "Rewrite the following sentence in a more formal, professional tone",
"input": "This plan looks pretty good — let's give it a try.",
"output": "The plan is feasible; a trial run is recommended."
},
{
"instruction": "Classify the sentiment of the text below. Output only: positive / neutral / negative",
"input": "The new client app crashed yet again, and support never responds.",
"output": "negative"
}
]4.2 Scale and Quality
| Data scale | Typical outcome | Notes |
|---|---|---|
| A few hundred samples | Observable behavior change | Only suits ultra-light "format constraints"; don't expect new capabilities |
| Thousands to tens of thousands | The quality bar for the vast majority of tasks | 200–2,000 characters per sample, covering the task's main forms |
| Hundreds of thousands | Clear capability gains | Approaching "rebuilding a general-purpose assistant"; expensive |
Three lessons from practice:
- Quality matters far more than quantity. 1,000 hand-curated samples often beat 100,000 auto-generated dirty ones. The industry has the "LIMA phenomenon": fine-tuning on just 1,000 high-quality instructions can significantly improve alignment and usability.
- Diversity beats repetition. Covering different forms, difficulty levels, and edge cases of the task is far more useful than copying the same sample ten times.
- Outputs must be "what the model should answer." SFT is essentially imitating the expected output distribution; if outputs contain wrong formats or hallucinations, the model will absorb them wholesale.
Data Leakage and Contamination
Public instruction datasets (chat data scraped from the web, shared SFT datasets, etc.) often mix in test-set answers, private text, even profanity. Fine-tuning data must be cleaned, deduplicated, and privacy-scrubbed; otherwise the model will memorize answers or pick up bad habits. For dataset tools and catalogs, see Datasets and Tools.
5. The Complete Workflow: From Data to Production
The end-to-end process of fine-tuning a model boils down to four steps:
① Data preparation → ② Base model selection → ③ Training → ④ Evaluation & iteration
Collect/clean/format into instruction JSON Pick a model for the task LoRA hyperparameter experiments Ship/regress/back to ①① Data preparation: Collect real business samples → clean and deduplicate → write instruction JSON → split into train/val (keep 5%–10% for val). This step eats up over 60% of the project's total time.
② Base model selection: Bigger isn't automatically better; follow three rules of thumb — match the task language to the base model's language (for Chinese tasks, prefer base models with strong Chinese corpora); the base model's knowledge and capability floor must be high enough (fine-tuning changes behavior, not the boundaries of what the model knows — see the risks section below); prefer open-source base models to reduce closed-source dependence. For a quick model reference, see Models and Leaderboards at a Glance.
③ Training: The hyperparameter window for LoRA fine-tuning is narrow; common starting points:
learning_rate = 2e-4 # LoRA typically uses 1e-4 ~ 5e-4, one notch higher than full fine-tuning
num_epochs = 3 # 1~3 epochs is enough for small data; more will overfit
batch_size = 4 ~ 16 # bounded by GPU memory; can pair with gradient accumulation
lora_r = 8 # commonly 8 / 16 / 32
lora_alpha = 16 # usually 1~2× r
max_seq_len = 2048 # set to your task's longest input; don't blindly max it out
warmup_steps = 100 # a short warmup stabilizes training④ Evaluation and iteration: After training, absolutely do not look only at the training loss — check validation-set performance and business metrics (format compliance rate, task accuracy, style similarity). Regress every historical fine-tuned version against a fixed eval set to prevent "fixing A and breaking B." For evaluation methodology, see LLM Evaluation and Benchmarks; for the tooling, see Building an LLM Evaluation Pipeline.
The One-Line Workflow Mantra
80% of a fine-tuning project's effort goes into the first two stages (data + evaluation design); actual GPU time is usually a small fraction. Build the eval set first; talk about training second.
6. Risks and Pitfalls
6.1 Catastrophic Forgetting
When a model trains hard on new-task data, it can "wash out" the general abilities learned during pretraining — questions it used to answer correctly start failing, and its common sense degrades. Usual triggers: data that is too small or too narrow, a learning rate that is too high, or too many epochs. Mitigations: keep epochs in check, mix in a small amount of general data (e.g., blend general instruction data at 5%–10%), and note that LoRA itself has a natural buffer against forgetting because it changes so few parameters.
6.2 Overfitting Small Datasets
Training for dozens of epochs on a few thousand samples, watching validation loss climb instead of fall, seeing perfect training-set performance while the real task falls apart — this is the most classic fine-tuning failure scene. Countermeasures: with small data, use a larger learning rate + fewer epochs + early stopping (monitor validation loss), and spot-check generation quality with random samples during training.
6.3 The Biggest Misconception: Fine-Tuning ≠ Learning New Knowledge
This is the point that bears repeating most. Fine-tuning mainly changes "behavior" (format, style, instruction-following) and almost never injects "knowledge" — during training the model may memorize scattered facts from the training set, but it cannot reliably learn a large body of new domain knowledge; real knowledge comes from the trillion-scale corpora of pretraining and is backed at runtime by Retrieval-Augmented Generation (RAG).
- Want the model to answer questions about new policies from after May 2026 → use RAG; fine-tuning can't fix recency;
- Want the model to master your company's 100,000 internal documents → use RAG; fine-tuning only helps it get the "way of answering questions about those documents" right;
- Want a stable output format and a voice that fits your brand → that's fine-tuning's comfort zone.
Rule of thumb: knowledge problems go to RAG, behavior problems go to fine-tuning, and combining the two is the production norm. For a fuller catalog of traps, see Common Pitfalls and Anti-Patterns.
7. Fine-Tuning and Alignment: Two Sides of Post-Training
Fine-tuning isn't an isolated step; it's one member of the "post-training" family. The full lifecycle of an industry LLM looks like this:
Pretraining (learning language from massive corpora) → Post-training (SFT fine-tuning → alignment via RLHF/DPO) → Deployment & inferenceAlignment is essentially the second half of fine-tuning: SFT first teaches the model to "follow instructions," then RLHF/DPO tunes it to be "safe, honest, and helpful" — both continue training on the pretrained model; only the objective function changes, from "imitating correct outputs" to "optimizing human preferences." So "fine-tuning" and "alignment" aren't two independent concepts but different stages of the same post-training pipeline; for the deep dive, see Alignment: RLHF and DPO.
Typical domain fine-tuning cases:
| Domain | Representatives | Key practices |
|---|---|---|
| Code | GitHub Copilot, DeepSeek-Coder, Code Llama | SFT on code-completion data (source code + comments + tests); for a complete breakdown of code-intelligence products, see GitHub Copilot and Code Intelligence |
| Healthcare | Med-PaLM, open-source medical base models | Fine-tune on curated data from medical records, medical QA, and clinical guidelines, while also doing safety alignment |
| Customer service / finance | Various industry LLMs | Fine-tune on historical tickets and standard scripts to build a "house voice," then layer RAG on top to supply product knowledge |
| Reasoning | DeepSeek-R1 | SFT on chain-of-thought reasoning data, then reinforcement learning to elicit deep thinking; see DeepSeek-R1 and Reasoning Models |
Beyond that, fine-tuned models are often combined with AI Agents — the stable output format produced by fine-tuning serves as the output constraint for tool calling and task planning. For the overall learning roadmap, see Learning Paths.
Further Reading
- Large Language Models (LLM) — what fine-tuning operates on: what pretrained models actually learn
- Prompt Engineering — the free option to try before fine-tuning
- Retrieval-Augmented Generation (RAG) — the right answer for knowledge problems; complements fine-tuning
- Alignment: RLHF and DPO — the second half of post-training: safety and human preferences
- Inference Optimization and Quantization — merging LoRA back into the backbone and quantized inference
- LLM Evaluation and Benchmarks — how to measure fine-tuning results
- Fine-Tune Your Own LLM — a complete hands-on walkthrough from data to code
- Common Pitfalls and Anti-Patterns — the traps fine-tuning projects hit most often
- Deployment and Inference Optimization in Practice — how to go to production after fine-tuning
- Datasets and Tools — a catalog of instruction datasets and fine-tuning tools
References
- Hu et al. LoRA: Low-Rank Adaptation of Large Language Models (2021) — the original LoRA paper, with the full mathematical derivation of low-rank decomposition
- Dettmers et al. QLoRA: Efficient Finetuning of Quantized LLMs (2023) — 4-bit quantization + LoRA; fine-tunes a 65B model on a single GPU
- Hugging Face PEFT official documentation — the unified library implementing LoRA, QLoRA, and Adapters
- Ouyang et al. Training language models to follow instructions with human feedback (InstructGPT, 2022) — the de facto standard for the SFT instruction-data paradigm
- Zhou et al. LIMA: Less Is More for Alignment (2023) — the classic evidence that quality comes first: 1,000 high-quality samples can achieve significant alignment
- Taori et al. Stanford Alpaca (2023) — the open-source milestone of fine-tuning on 50K self-generated instruction samples
- Rafailov et al. Direct Preference Optimization (DPO, 2023) — post-training directly on preference pairs; the simplified alternative to RLHF