Theme
Fine-Tuning: SFT and Parameter-Efficient Fine-Tuning
Fine-tuning is updating model parameters with a small amount of supervised data on top of a pre-trained large model, transforming a "can predict the next token" general-purpose language model into a "can follow instructions to complete tasks" usable product. Pretraining determines the model's knowledge ceiling; fine-tuning determines whether it can "connect the knowledge correctly" — pretraining answers "how smart the model is," fine-tuning answers "how usable the model is."
This article covers the post-training theme. See Pretraining: Data and Objectives for pretraining principles; after fine-tuning comes the even deeper Alignment: RLHF and DPO stage; for hands-on practice, see Fine-Tuning Practice: Full LoRA Pipeline.
1. Why Fine-Tuning Is Needed: The Separation of Knowledge and Behavior
A pre-trained model learns "statistical patterns of language": predict the next token given preceding context. This capability alone can already generate coherent text, but it has three clear productization flaws:
- Doesn't follow instructions: you ask it to "summarize in one sentence," it may output three paragraphs; you tell it to "output only JSON," it produces prose. "Follow instructions" was never in the pretraining objective.
- Wrong format: real products need dialogue-style, structured, role-constrained output, while the raw foundation only "continues writing from preceding context."
- No behavioral boundaries: the foundation model won't refuse harmful requests or admit it doesn't know.
These flaws aren't fundamentally about "insufficient knowledge" — they're about wrong behavior. Fine-tuning primarily changes behavior and style, not knowledge — the vast majority of a model's learned world knowledge comes from pretraining; "injecting new knowledge" (especially factual knowledge) with a few thousand fine-tuning samples is nearly impossible.
From the technical history perspective, the "pretrain + fine-tune" combination isn't new: GPT-1 (2018) already showed that general pretraining + task-specific fine-tuning beats training from scratch; BERT-era fine-tuning was NLP's standard. The change in the big model era is that with parameter explosion, full-parameter fine-tuning became expensive, making parameter-efficient fine-tuning (PEFT) mainstream, but the paradigm of "adjusting behavior with supervised data" remains continuous.
Knowledge vs Behavior
Think of pretraining as "general education" and fine-tuning as "on-the-job training": general education determines how much a person knows, on-the-job training determines how they perform on the job. To make a model "know more," go back to pretraining with more data (see Scaling Laws); to make it "behave more professionally," that's fine-tuning's home turf. This is also why evaluation distinguishes knowledge benchmarks from behavior benchmarks.
Thus fine-tuning's typical use cases include:
| Use Case | What to change | Typical data |
|---|---|---|
| Instruction following | Behavior: understand and execute instructions | Instruction-response pairs (Alpaca-style) |
| Chat-style | Behavior + style: multi-turn dialogue, roleplay | Dialogue trees (multi-turn conversation records) |
| Domain style | Style: legal docs, customer service scripts, code comment habits | Domain corpus |
| Output format | Behavior: JSON/table/summary formatting | Structured output samples |
| Task adaptation | Capability: classification/extraction/rewriting downstream tasks | Task-annotated data |
| Inject knowledge | Knowledge (limited, poor effect) | Domain Q&A pairs |
Fine-tuning can't replace pretraining
"Feeding" new facts to a model via fine-tuning (like internal company docs) usually works poorly: too few samples, too weak signal, and easy to mess up the model's existing knowledge. Such needs should prioritize retrieval-augmented generation (see RAG: Retrieval-Augmented Generation) over fine-tuning.
2. SFT: Supervised Fine-Tuning
Supervised Fine-Tuning (SFT) is the most basic, most commonly used type of fine-tuning: given a batch of "input → expected output" samples, use standard cross-entropy loss to teach the model to imitate the expected output. It's the first stop in post-training, and step one of alignment's RLHF (one of InstructGPT's three steps).
1. What SFT Data Looks Like
The core form of SFT data is the instruction-response pair:
User: Explain what backpropagation is in one sentence.
Assistant: Backpropagation is an algorithm that computes the gradient of loss with respect to network parameters using the chain rule, used to train neural networks.Higher quality is the dialogue tree: a multi-turn conversation record where each turn fully expands the historical context. This way the model learns not just "correct answers" but "remembering context across turns, no repetition or contradiction" dialogue behavior.
Famous open-source SFT data includes: Stanford Alpaca (~52K samples, 2023, generated instruction-response pairs using the strongest available API model at the time), ShareGPT (real user conversations with ChatGPT), UltraChat, etc. — full archive at Datasets & Benchmarks.
2. SFT Loss: Still Cross-Entropy
SFT's training objective is structurally identical to pretraining — just the data changed — pretraining is "continue writing," SFT is "write according to the standard answer." For each instruction-response pair, only the response tokens' cross-entropy is computed:
text
L_SFT = -Σ log P( y_t | x, y_<t ; θ )
↑ sum and average over each token y_t in the response y; x is the full instruction (including system prompt / dialogue history)In practice, this is usually implemented with "labels masking": the instruction tokens are masked in the loss (label set to -100), so the model only learns the response generation pattern.
3. Data Quality Far Outweighs Data Volume
SFT scale is far smaller than pretraining: a few thousand to tens of thousands of high-quality samples often suffice; hundreds of thousands of garbage samples can ruin the model. The main factors determining SFT effect, in order of importance:
- Quality and diversity: the instruction types, difficulty, and domain distribution covered matter more than count.
- Clean and consistent: errors, duplicates, contradictory samples directly teach bad behavior.
- Quantity: only after quality is met does increasing quantity yield positive returns.
The "quality > quantity" litmus test
A practical rule: if as a human you wouldn't consider a sample a "good response," don't put it in SFT data. After generating data with an API model, manually spot-check and deduplicate (Alpaca later discovered a large number of duplicates and meaningless instructions in its data).
4. Pretraining vs SFT Comparison
| Dimension | Pretraining | SFT |
|---|---|---|
| Data | Massive web text (trillions of tokens) | Small amount of human/synthetic data (thousands to millions of samples) |
| Objective | Predict next token (continuation) | Generate expected response to given instruction |
| Learns | Language, knowledge, reasoning foundation | Instruction following, style, format |
| Cost | Tens of thousands of GPU hours | A few GPUs suffice |
| Changes | Knowledge ceiling | Behavior and expression |
3. Full-Parameter Fine-Tuning vs Parameter-Efficient Fine-Tuning
1. The Cost and Risk of Full-Parameter Fine-Tuning
Full-parameter fine-tuning updates all model parameters. For 7B–70B models, this requires a large amount of memory (optimizer states, gradients, activations must all reside), and has two core issues:
- High cost: an 80GB A100/H100 can barely handle 7B full-parameter training; 70B requires multi-GPU or whole machines.
- Catastrophic forgetting: to fit fine-tuning data, the model may "overwrite" general capabilities learned in pretraining — typical manifestation: math/code/commonsense reasoning gets worse after fine-tuning.
Catastrophic forgetting is a real risk
Many teams only look at business metrics after fine-tuning (it went up), not general capabilities (they quietly dropped). Before deployment, always run a set of general benchmarks for regression — precisely what Evaluation & Benchmarks repeatedly emphasizes: fine-tuning must accompany an evaluation loop, or you won't know what got "better" alongside what got "worse."
2. Parameter-Efficient Fine-Tuning: Change Only a Small Part
Parameter-Efficient Fine-Tuning (PEFT) freezes most model parameters, training only a small number of added/selected parameters, reducing fine-tuning cost by one to two orders of magnitude. Mainstream methods compared:
| Method | Representative Paper/Year | What to train | Memory/Cost | Typical effect |
|---|---|---|---|---|
| Full fine-tuning | — | All parameters | High | Highest ceiling, but large forgetting risk |
| LoRA | Hu et al. 2021 | Injected low-rank matrices | Low | <1% trainable params, effect close to full |
| QLoRA | Dettmers et al. 2023 | Low-rank matrices + 4-bit base | Very low | 65B model fine-tunable on single GPU |
| Adapter | Houlsby et al. 2019 | Small networks between Transformer layers | Low | Adds params per layer, modular |
| P-Tuning / Prefix | Liu et al. 2022 / Li & Liang 2021 | Learnable continuous prefix vectors | Lowest | Only touches near-embedding layers |
| BitFit | Ben Zaken et al. 2021 | Bias terms only | Lowest | Extremely lightweight baseline |
3. LoRA Principle: Low-Rank Decomposition
LoRA (Low-Rank Adaptation)'s core assumption: the update ΔW that fine-tuning applies to pre-trained weights is low-rank — no need to train a full large matrix, just the product of two small matrices.
Let pre-trained weight be W₀ (shape d×d). LoRA freezes it and introduces two trainable matrices A (d×r) and B (r×d):
text
W = W₀ + ΔW = W₀ + (α/r) · B·A
↑ ↑ ↑
frozen scaling low-rank increment, rank r << dForward computation becomes h = W₀·x + (α/r)·B·A·x. Since r is small (commonly 8/16/32/64), trainable parameters drop from d×d to 2·d·r, typically less than 1% of full-parameter. The LoRA paper reported: on GPT-3 175B, trainable params reduced by ~10,000× with effect still close to full-parameter fine-tuning.
LoRA's key hyperparameters:
| Hyperparameter | Meaning | Practical experience |
|---|---|---|
| Rank r | Dimension of low-rank matrices | 8–64; simple tasks use 8, complex domains use 64; diminishing returns for r too large |
| α | Scaling coefficient | Often ~2r; α and r interact, affecting LR perception |
| target_modules | Which layers to inject (q_proj/k_proj/v_proj/o_proj, etc.) | Common: q/v or all attention projections |
| Learning rate | Fine-tuning-specific LR | Generally lower than pretraining (~1e-4 ~ 2e-4) |
| Merge | Add BA back to W₀ after training | Zero extra overhead at inference |
Why LoRA is effective and popular
Three reasons: ① few trainable params, drastically reducing memory and compute needs; ② frozen backbone, catastrophic forgetting significantly mitigated; ③ trained products are just two small matrices, enabling multiple LoRA stacks on one base model for on-demand switching (routing, multi-tenant adaptation).
QLoRA further quantizes the base weights to 4-bit (NF4 format) and freezes them, only dequantizing for LoRA parameter backprop computation, making 65B-scale models fine-tunable on a single high-end GPU (Dettmers et al., 2023). Quantization's system background is in Deployment & Serving.
4. Brief Overview of Other PEFT Methods
- Adapter: insert a small feedforward network within each Transformer layer (down-project then up-project), only updating these Adapters during training. Drawback: multi-layer insertion, slight extra latency at inference.
- P-Tuning v2 / Prefix Tuning: don't modify weights; only prepend a set of learnable continuous vectors (soft prompts) at the input. Extremely fast to train, but capability improvement on complex tasks is typically weaker than LoRA.
- BitFit: only train all bias parameters; surprisingly good effect on some tasks, the most lightweight baseline.
5. PEFT vs Full Fine-Tuning: The Effectiveness Gap
A question that must be faced honestly: does PEFT's "savings" come at the cost of "weakness"? From public experiments, the gap between PEFT and full is small in most scenarios, but not zero:
| Evidence | Conclusion |
|---|---|
| LoRA original paper (GPT-3 175B) | Close to full-parameter fine-tuning on most downstream tasks, even better on some |
| QLoRA paper (Guanaco series) | 4-bit base + LoRA on multiple benchmarks within acceptable range of full-precision fine-tuning |
| Community large-scale practice | On complex generation/reasoning tasks, full or larger-rank LoRA still has ceiling advantage |
Comprehensive conclusion: PEFT is fully adequate for simple tasks (classification, extraction, style adaptation); for complex tasks sensitive to ceiling, first use PEFT to validate data and direction, then decide whether to upgrade to full or bigger model. This is far steadier than betting on full-parameter fine-tuning from the start.
4. Four Key Issues in Fine-Tuning Engineering
1. Overfitting and Underfitting
Fine-tuning data volumes are small, making overfitting easy to trigger. Control methods: lower LR (start at 1e-4 ~ 2e-4), limit epochs (1–3 suffice; SFT gains little from extra rounds), early stopping (watch validation loss), add dropout. Underfitting is the other direction — LR too low or LoRA rank too small, making fine-tuning almost ineffective, performance indistinguishable from the base.
2. Mitigating Catastrophic Forgetting
- Use LoRA/QLoRA to freeze the backbone (most effective);
- Mix 5%–10% general/domain-preservation data into fine-tuning data (the "replay" idea from continual learning);
- Run general benchmark regression tests before/after fine-tuning (see Evaluation & Benchmarks);
- If needed, model merging (e.g., WEMIX-style parameter interpolation) to restore general capabilities.
3. Data Quality and Deduplication
Duplicate samples make the model overfit to "memorizing answers" for certain inputs; low-quality samples (irrelevant answers, grammar errors, harmful content) pollute behavior. SFT data preparation flow: format cleaning → deduplication (exact + semantic) → quality filtering → manual spot-check → mix by task type.
4. Evaluation Loop After Fine-Tuning
Fine-tuning isn't "train and done" — three questions must be answered: did the target behavior improve? Did general capabilities drop? Did harmful behavior increase? These correspond to business eval sets, general benchmarks (MMLU/GSM8K/HumanEval, etc., see Datasets & Benchmarks), and safety evals (see Safety & Risks). Full methodology in Evaluation in Practice.
5. Common Failure Modes and Troubleshooting
| Symptom | Possible Cause | Troubleshooting Direction |
|---|---|---|
| Training loss doesn't drop | LR too low, LoRA rank too small, data format wrong | Raise LR, increase r, check if instruction was masked |
| Training loss drops but output unchanged | Only learned instruction part, response part didn't participate in loss | Check if labels mask also masked the response |
| Business metrics up, general capabilities plummet | Catastrophic forgetting | Roll back to LoRA, mix in preservation data, run general benchmark regression |
| Output is all the training set's "standard answers" | Overfitting / too little/too-uniform data | Lower epochs, increase data diversity, add dropout |
| More hallucinations | Noisy data, or alignment didn't keep up | Data cleaning, add alignment stage |
| No difference before/after fine-tuning | LR too low / r too small / data overlaps with base capability | First run a small-scale full-parameter test to validate data effectiveness |
5. Positioning of Fine-Tuning vs Alignment vs Prompting vs RAG
1. Is Alignment Needed After SFT?
Yes. SFT solves "can imitate"; alignment solves "knows what to do, what not to do." InstructGPT's practice showed: SFT-only models still output harmful content and still make commonsense errors on simple questions; after SFT, going through alignment: RLHF and DPO (reward model + RL, or direct preference optimization), the model's helpfulness and safety improved significantly.
| Stage | Data/signal | Problem solved |
|---|---|---|
| Pretraining | Massive text | Knowledge and language capability |
| SFT | Instruction-response pairs | Instruction following and behavior format |
| RLHF/DPO | Human preference ranking | Values, safety, helpfulness |
| Evaluation | Benchmarks + red team | Verification and fallback |
2. Fine-Tuning vs Prompting vs RAG
| Method | What it changes | Cost | Applicable scenario |
|---|---|---|---|
| Prompt engineering | Doesn't change the model | Lowest | Behavior tweaking, format constraints, instant effect |
| RAG | Doesn't change model, connects external knowledge | Medium | Need new knowledge, traceable, knowledge updates frequently |
| Fine-tuning | Changes model behavior | High | Fixed style/format/task shape, prompting can't handle it |
Don't fine-tune right away
For most business problems, start with prompting (see Prompt Engineering); use RAG for new knowledge; only consider fine-tuning when both paths are exhausted — e.g., output style must be 100% stable, no amount of prompting makes it listen. Fine-tuning is the most expensive means and the last resort. Common anti-patterns see Common Pitfalls & Anti-Patterns.
3. The Fine-Tuning Cost Account: When It's Not Worth It
Fine-tuning's real cost goes far beyond GPU time — calculate the full account before deciding:
| Cost Item | Explanation |
|---|---|
| Data cost | Labeling/cleaning/spot-checking labor, often underestimated |
| Training cost | GPU rental, trial rounds (3–5 rounds of tuning is common) |
| Evaluation cost | Before/after regression, red team and safety evals |
| Ops cost | Model version management, canary rollout, monitoring, rollback |
When fine-tuning is clearly not worth it: the need is temporary; data volume insufficient to support stable behavior; business metric changes can't be quantified by evaluation; or the team lacks continuous iteration capability. In these cases, prompting or RAG are more rational choices.
4. How to Iterate After Fine-Tuning
A healthy fine-tuning workflow is: small data trial → evaluate → expand data → re-evaluate → canary deploy → online feedback feeds data back. Fine-tuning isn't a one-time action; it's a continuous iteration loop, like any product feature.
Further Reading
- Alignment: RLHF and DPO — the alignment stage after SFT, where RLHF's first step is SFT
- Pretraining: Data and Objectives — the modeling paradigm before fine-tuning, understand the boundaries between the two
- Evaluation & Benchmarks — the evaluation system that must accompany fine-tuning
- Fine-Tuning Practice: Full LoRA Pipeline — complete hands-on from data to deployment
- Prompt Engineering — the low-cost method tried before fine-tuning
- Datasets & Benchmarks — SFT data and benchmark list
References
- Hu et al. LoRA: Low-Rank Adaptation of Large Language Models (ICLR 2022, originally on arXiv Oct 2021) — the original LoRA paper, low-rank decomposition fine-tuning
- Dettmers et al. QLoRA: Efficient Fine-Tuning of Quantized LLMs (NeurIPS 2023) — 4-bit quantization + LoRA, single-GPU 65B fine-tuning
- Taori et al. Stanford Alpaca: An Instruction-following LLaMA model (2023) — ~52K instruction data open-source project
- Houlsby et al. Parameter-Efficient Transfer Learning for NLP (ICML 2019) — the Adapter method
- Liu et al. P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-Tuning Universally Across Scales and Tasks (2022) — P-Tuning series
- Ouyang et al. Training language models to follow instructions with human feedback (InstructGPT, 2022) — the original source for SFT as step one of RLHF