Skip to content

Fine-Tuning: SFT and Parameter-Efficient Fine-Tuning

At a glance Fine-tuning continues training on a pre-trained foundation with supervised data, transforming "can generate" into "can do things by instruction." This article systematically covers SFT data and objectives, compares full-parameter fine-tuning with LoRA/QLoRA/Adapter/P-Tuning, explains LoRA low-rank decomposition principles, and addresses critical engineering issues like catastrophic forgetting and data quality.

Fine-Tuning: SFT and Parameter-Efficient Fine-Tuning ​

Fine-tuning is updating model parameters with a small amount of supervised data on top of a pre-trained large model, transforming a "can predict the next token" general-purpose language model into a "can follow instructions to complete tasks" usable product. Pretraining determines the model's knowledge ceiling; fine-tuning determines whether it can "connect the knowledge correctly" — pretraining answers "how smart the model is," fine-tuning answers "how usable the model is."

This article covers the post-training theme. See Pretraining: Data and Objectives for pretraining principles; after fine-tuning comes the even deeper Alignment: RLHF and DPO stage; for hands-on practice, see Fine-Tuning Practice: Full LoRA Pipeline.

1. Why Fine-Tuning Is Needed: The Separation of Knowledge and Behavior ​

A pre-trained model learns "statistical patterns of language": predict the next token given preceding context. This capability alone can already generate coherent text, but it has three clear productization flaws:

  1. Doesn't follow instructions: you ask it to "summarize in one sentence," it may output three paragraphs; you tell it to "output only JSON," it produces prose. "Follow instructions" was never in the pretraining objective.
  2. Wrong format: real products need dialogue-style, structured, role-constrained output, while the raw foundation only "continues writing from preceding context."
  3. No behavioral boundaries: the foundation model won't refuse harmful requests or admit it doesn't know.

These flaws aren't fundamentally about "insufficient knowledge" — they're about wrong behavior. Fine-tuning primarily changes behavior and style, not knowledge — the vast majority of a model's learned world knowledge comes from pretraining; "injecting new knowledge" (especially factual knowledge) with a few thousand fine-tuning samples is nearly impossible.

From the technical history perspective, the "pretrain + fine-tune" combination isn't new: GPT-1 (2018) already showed that general pretraining + task-specific fine-tuning beats training from scratch; BERT-era fine-tuning was NLP's standard. The change in the big model era is that with parameter explosion, full-parameter fine-tuning became expensive, making parameter-efficient fine-tuning (PEFT) mainstream, but the paradigm of "adjusting behavior with supervised data" remains continuous.

Knowledge vs Behavior

Think of pretraining as "general education" and fine-tuning as "on-the-job training": general education determines how much a person knows, on-the-job training determines how they perform on the job. To make a model "know more," go back to pretraining with more data (see Scaling Laws); to make it "behave more professionally," that's fine-tuning's home turf. This is also why evaluation distinguishes knowledge benchmarks from behavior benchmarks.

Thus fine-tuning's typical use cases include:

Use CaseWhat to changeTypical data
Instruction followingBehavior: understand and execute instructionsInstruction-response pairs (Alpaca-style)
Chat-styleBehavior + style: multi-turn dialogue, roleplayDialogue trees (multi-turn conversation records)
Domain styleStyle: legal docs, customer service scripts, code comment habitsDomain corpus
Output formatBehavior: JSON/table/summary formattingStructured output samples
Task adaptationCapability: classification/extraction/rewriting downstream tasksTask-annotated data
Inject knowledgeKnowledge (limited, poor effect)Domain Q&A pairs

Fine-tuning can't replace pretraining

"Feeding" new facts to a model via fine-tuning (like internal company docs) usually works poorly: too few samples, too weak signal, and easy to mess up the model's existing knowledge. Such needs should prioritize retrieval-augmented generation (see RAG: Retrieval-Augmented Generation) over fine-tuning.

2. SFT: Supervised Fine-Tuning ​

Supervised Fine-Tuning (SFT) is the most basic, most commonly used type of fine-tuning: given a batch of "input → expected output" samples, use standard cross-entropy loss to teach the model to imitate the expected output. It's the first stop in post-training, and step one of alignment's RLHF (one of InstructGPT's three steps).

1. What SFT Data Looks Like ​

The core form of SFT data is the instruction-response pair:

User: Explain what backpropagation is in one sentence.
Assistant: Backpropagation is an algorithm that computes the gradient of loss with respect to network parameters using the chain rule, used to train neural networks.

Higher quality is the dialogue tree: a multi-turn conversation record where each turn fully expands the historical context. This way the model learns not just "correct answers" but "remembering context across turns, no repetition or contradiction" dialogue behavior.

Famous open-source SFT data includes: Stanford Alpaca (~52K samples, 2023, generated instruction-response pairs using the strongest available API model at the time), ShareGPT (real user conversations with ChatGPT), UltraChat, etc. — full archive at Datasets & Benchmarks.

2. SFT Loss: Still Cross-Entropy ​

SFT's training objective is structurally identical to pretraining — just the data changed — pretraining is "continue writing," SFT is "write according to the standard answer." For each instruction-response pair, only the response tokens' cross-entropy is computed:

text
L_SFT = -Σ log P( y_t | x, y_<t ; θ )
         ↑ sum and average over each token y_t in the response y; x is the full instruction (including system prompt / dialogue history)

In practice, this is usually implemented with "labels masking": the instruction tokens are masked in the loss (label set to -100), so the model only learns the response generation pattern.

3. Data Quality Far Outweighs Data Volume ​

SFT scale is far smaller than pretraining: a few thousand to tens of thousands of high-quality samples often suffice; hundreds of thousands of garbage samples can ruin the model. The main factors determining SFT effect, in order of importance:

  1. Quality and diversity: the instruction types, difficulty, and domain distribution covered matter more than count.
  2. Clean and consistent: errors, duplicates, contradictory samples directly teach bad behavior.
  3. Quantity: only after quality is met does increasing quantity yield positive returns.

The "quality > quantity" litmus test

A practical rule: if as a human you wouldn't consider a sample a "good response," don't put it in SFT data. After generating data with an API model, manually spot-check and deduplicate (Alpaca later discovered a large number of duplicates and meaningless instructions in its data).

4. Pretraining vs SFT Comparison ​

DimensionPretrainingSFT
DataMassive web text (trillions of tokens)Small amount of human/synthetic data (thousands to millions of samples)
ObjectivePredict next token (continuation)Generate expected response to given instruction
LearnsLanguage, knowledge, reasoning foundationInstruction following, style, format
CostTens of thousands of GPU hoursA few GPUs suffice
ChangesKnowledge ceilingBehavior and expression

3. Full-Parameter Fine-Tuning vs Parameter-Efficient Fine-Tuning ​

1. The Cost and Risk of Full-Parameter Fine-Tuning ​

Full-parameter fine-tuning updates all model parameters. For 7B–70B models, this requires a large amount of memory (optimizer states, gradients, activations must all reside), and has two core issues:

  • High cost: an 80GB A100/H100 can barely handle 7B full-parameter training; 70B requires multi-GPU or whole machines.
  • Catastrophic forgetting: to fit fine-tuning data, the model may "overwrite" general capabilities learned in pretraining — typical manifestation: math/code/commonsense reasoning gets worse after fine-tuning.

Catastrophic forgetting is a real risk

Many teams only look at business metrics after fine-tuning (it went up), not general capabilities (they quietly dropped). Before deployment, always run a set of general benchmarks for regression — precisely what Evaluation & Benchmarks repeatedly emphasizes: fine-tuning must accompany an evaluation loop, or you won't know what got "better" alongside what got "worse."

2. Parameter-Efficient Fine-Tuning: Change Only a Small Part ​

Parameter-Efficient Fine-Tuning (PEFT) freezes most model parameters, training only a small number of added/selected parameters, reducing fine-tuning cost by one to two orders of magnitude. Mainstream methods compared:

MethodRepresentative Paper/YearWhat to trainMemory/CostTypical effect
Full fine-tuning—All parametersHighHighest ceiling, but large forgetting risk
LoRAHu et al. 2021Injected low-rank matricesLow<1% trainable params, effect close to full
QLoRADettmers et al. 2023Low-rank matrices + 4-bit baseVery low65B model fine-tunable on single GPU
AdapterHoulsby et al. 2019Small networks between Transformer layersLowAdds params per layer, modular
P-Tuning / PrefixLiu et al. 2022 / Li & Liang 2021Learnable continuous prefix vectorsLowestOnly touches near-embedding layers
BitFitBen Zaken et al. 2021Bias terms onlyLowestExtremely lightweight baseline

3. LoRA Principle: Low-Rank Decomposition ​

LoRA (Low-Rank Adaptation)'s core assumption: the update ΔW that fine-tuning applies to pre-trained weights is low-rank — no need to train a full large matrix, just the product of two small matrices.

Let pre-trained weight be W₀ (shape d×d). LoRA freezes it and introduces two trainable matrices A (d×r) and B (r×d):

text
W = W₀ + ΔW = W₀ + (α/r) · B·A
                ↑      ↑         ↑
             frozen  scaling   low-rank increment, rank r << d

Forward computation becomes h = W₀·x + (α/r)·B·A·x. Since r is small (commonly 8/16/32/64), trainable parameters drop from d×d to 2·d·r, typically less than 1% of full-parameter. The LoRA paper reported: on GPT-3 175B, trainable params reduced by ~10,000× with effect still close to full-parameter fine-tuning.

LoRA's key hyperparameters:

HyperparameterMeaningPractical experience
Rank rDimension of low-rank matrices8–64; simple tasks use 8, complex domains use 64; diminishing returns for r too large
αScaling coefficientOften ~2r; α and r interact, affecting LR perception
target_modulesWhich layers to inject (q_proj/k_proj/v_proj/o_proj, etc.)Common: q/v or all attention projections
Learning rateFine-tuning-specific LRGenerally lower than pretraining (~1e-4 ~ 2e-4)
MergeAdd BA back to W₀ after trainingZero extra overhead at inference

Why LoRA is effective and popular

Three reasons: ① few trainable params, drastically reducing memory and compute needs; ② frozen backbone, catastrophic forgetting significantly mitigated; ③ trained products are just two small matrices, enabling multiple LoRA stacks on one base model for on-demand switching (routing, multi-tenant adaptation).

QLoRA further quantizes the base weights to 4-bit (NF4 format) and freezes them, only dequantizing for LoRA parameter backprop computation, making 65B-scale models fine-tunable on a single high-end GPU (Dettmers et al., 2023). Quantization's system background is in Deployment & Serving.

4. Brief Overview of Other PEFT Methods ​

  • Adapter: insert a small feedforward network within each Transformer layer (down-project then up-project), only updating these Adapters during training. Drawback: multi-layer insertion, slight extra latency at inference.
  • P-Tuning v2 / Prefix Tuning: don't modify weights; only prepend a set of learnable continuous vectors (soft prompts) at the input. Extremely fast to train, but capability improvement on complex tasks is typically weaker than LoRA.
  • BitFit: only train all bias parameters; surprisingly good effect on some tasks, the most lightweight baseline.

5. PEFT vs Full Fine-Tuning: The Effectiveness Gap ​

A question that must be faced honestly: does PEFT's "savings" come at the cost of "weakness"? From public experiments, the gap between PEFT and full is small in most scenarios, but not zero:

EvidenceConclusion
LoRA original paper (GPT-3 175B)Close to full-parameter fine-tuning on most downstream tasks, even better on some
QLoRA paper (Guanaco series)4-bit base + LoRA on multiple benchmarks within acceptable range of full-precision fine-tuning
Community large-scale practiceOn complex generation/reasoning tasks, full or larger-rank LoRA still has ceiling advantage

Comprehensive conclusion: PEFT is fully adequate for simple tasks (classification, extraction, style adaptation); for complex tasks sensitive to ceiling, first use PEFT to validate data and direction, then decide whether to upgrade to full or bigger model. This is far steadier than betting on full-parameter fine-tuning from the start.

4. Four Key Issues in Fine-Tuning Engineering ​

1. Overfitting and Underfitting ​

Fine-tuning data volumes are small, making overfitting easy to trigger. Control methods: lower LR (start at 1e-4 ~ 2e-4), limit epochs (1–3 suffice; SFT gains little from extra rounds), early stopping (watch validation loss), add dropout. Underfitting is the other direction — LR too low or LoRA rank too small, making fine-tuning almost ineffective, performance indistinguishable from the base.

2. Mitigating Catastrophic Forgetting ​

  • Use LoRA/QLoRA to freeze the backbone (most effective);
  • Mix 5%–10% general/domain-preservation data into fine-tuning data (the "replay" idea from continual learning);
  • Run general benchmark regression tests before/after fine-tuning (see Evaluation & Benchmarks);
  • If needed, model merging (e.g., WEMIX-style parameter interpolation) to restore general capabilities.

3. Data Quality and Deduplication ​

Duplicate samples make the model overfit to "memorizing answers" for certain inputs; low-quality samples (irrelevant answers, grammar errors, harmful content) pollute behavior. SFT data preparation flow: format cleaning → deduplication (exact + semantic) → quality filtering → manual spot-check → mix by task type.

4. Evaluation Loop After Fine-Tuning ​

Fine-tuning isn't "train and done" — three questions must be answered: did the target behavior improve? Did general capabilities drop? Did harmful behavior increase? These correspond to business eval sets, general benchmarks (MMLU/GSM8K/HumanEval, etc., see Datasets & Benchmarks), and safety evals (see Safety & Risks). Full methodology in Evaluation in Practice.

5. Common Failure Modes and Troubleshooting ​

SymptomPossible CauseTroubleshooting Direction
Training loss doesn't dropLR too low, LoRA rank too small, data format wrongRaise LR, increase r, check if instruction was masked
Training loss drops but output unchangedOnly learned instruction part, response part didn't participate in lossCheck if labels mask also masked the response
Business metrics up, general capabilities plummetCatastrophic forgettingRoll back to LoRA, mix in preservation data, run general benchmark regression
Output is all the training set's "standard answers"Overfitting / too little/too-uniform dataLower epochs, increase data diversity, add dropout
More hallucinationsNoisy data, or alignment didn't keep upData cleaning, add alignment stage
No difference before/after fine-tuningLR too low / r too small / data overlaps with base capabilityFirst run a small-scale full-parameter test to validate data effectiveness

5. Positioning of Fine-Tuning vs Alignment vs Prompting vs RAG ​

1. Is Alignment Needed After SFT? ​

Yes. SFT solves "can imitate"; alignment solves "knows what to do, what not to do." InstructGPT's practice showed: SFT-only models still output harmful content and still make commonsense errors on simple questions; after SFT, going through alignment: RLHF and DPO (reward model + RL, or direct preference optimization), the model's helpfulness and safety improved significantly.

StageData/signalProblem solved
PretrainingMassive textKnowledge and language capability
SFTInstruction-response pairsInstruction following and behavior format
RLHF/DPOHuman preference rankingValues, safety, helpfulness
EvaluationBenchmarks + red teamVerification and fallback

2. Fine-Tuning vs Prompting vs RAG ​

MethodWhat it changesCostApplicable scenario
Prompt engineeringDoesn't change the modelLowestBehavior tweaking, format constraints, instant effect
RAGDoesn't change model, connects external knowledgeMediumNeed new knowledge, traceable, knowledge updates frequently
Fine-tuningChanges model behaviorHighFixed style/format/task shape, prompting can't handle it

Don't fine-tune right away

For most business problems, start with prompting (see Prompt Engineering); use RAG for new knowledge; only consider fine-tuning when both paths are exhausted — e.g., output style must be 100% stable, no amount of prompting makes it listen. Fine-tuning is the most expensive means and the last resort. Common anti-patterns see Common Pitfalls & Anti-Patterns.

3. The Fine-Tuning Cost Account: When It's Not Worth It ​

Fine-tuning's real cost goes far beyond GPU time — calculate the full account before deciding:

Cost ItemExplanation
Data costLabeling/cleaning/spot-checking labor, often underestimated
Training costGPU rental, trial rounds (3–5 rounds of tuning is common)
Evaluation costBefore/after regression, red team and safety evals
Ops costModel version management, canary rollout, monitoring, rollback

When fine-tuning is clearly not worth it: the need is temporary; data volume insufficient to support stable behavior; business metric changes can't be quantified by evaluation; or the team lacks continuous iteration capability. In these cases, prompting or RAG are more rational choices.

4. How to Iterate After Fine-Tuning ​

A healthy fine-tuning workflow is: small data trial → evaluate → expand data → re-evaluate → canary deploy → online feedback feeds data back. Fine-tuning isn't a one-time action; it's a continuous iteration loop, like any product feature.

Further Reading ​

References ​