Skip to content

Fine-Tuning and PEFT (LoRA)

At a glance Fine-tuning continues training a pretrained model for new objectives while PEFT updates only a small share of parameters; this article explains fine-tuning vs. prompting vs. RAG, the spectrum from full fine-tuning (FFT) to QLoRA, the low-rank decomposition behind LoRA, instruction data, and the complete end-to-end workflow.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

Fine-Tuning and PEFT (LoRA) ​

Fine-tuning (FT) means continuing to train a pretrained model on labeled data from a specific task or domain so it adapts to a new objective; parameter-efficient fine-tuning (PEFT) is a training paradigm that reaches similar results by updating only a small subset of parameters (or even by adding just a few new ones).

Pretrained models (Llama, Qwen, DeepSeek, etc.) picked up general language ability from trillion-scale corpora — they are "generalists who know a little about everything." But the model doesn't know whether you want JSON output or customer-service phrasing, doesn't know your industry jargon, and has no idea where your company's knowledge base lives. Fine-tuning is the step that turns the generalist into a specialist — it doesn't touch the underlying language ability itself, but adjusts the model's "behavioral habits" and "output preferences." Almost every major LLM product you can name today (ChatGPT, Claude, DeepSeek) went through the full path of "pretraining → post-training (including fine-tuning)"; see Large Language Models (LLM) and A Brief History.

1. Why Fine-Tuning: Three Paths from Generalist to Specialist ​

An LLM that already speaks and reasons becomes one that "answers your domain questions in your format" through exactly three mainstream routes:

ApproachWhat it doesChanges model parameters?Typical costBest for
Prompt engineeringSpell out rules, examples, and persona in the inputNoNearly zeroQuick validation, rule-based tasks, things the model can already do
Retrieval-Augmented Generation (RAG)Hook up an external knowledge base and stitch retrieved content into the contextNoBuilding the index + retrieval pipelineKnowledge-intensive work, facts that must be fresh and accurate, content that needs to be traceable
Fine-tuning (FT)Keep training on task data to change model behavior/format/styleYesGPU training costFixed output formats, a distinctive house voice, behavior problems prompts can't fix

Rule of thumb: fine-tuning is for the problems prompting can't fix. If the model knows how but answers wrong → adjust the prompt; if the answer needs external facts → bring in RAG; if the format just won't come out right, the voice never sounds like yours, or everything won't fit into a prompt → then consider fine-tuning. The three are not mutually exclusive; the most common production combination is "fine-tuning for format + RAG for knowledge." For a detailed comparison, see Prompt Engineering and Retrieval-Augmented Generation (RAG).

A Practical Screening Framework

Ask three questions: ① Is the model capable enough (if not → move to a bigger model or add knowledge)? ② Is it a "what to know" problem (→ RAG) or a "how to talk" problem (→ fine-tuning)? ③ Can prompting cover it (if so → don't train yet)? Run a baseline first; talk fine-tuning later.

2. The Fine-Tuning Spectrum: From Full Fine-Tuning to QLoRA ​

"Fine-tuning" is a big umbrella; inside it stretches a full spectrum from "touching every parameter" to "touching only a few matrices." The two key metrics are how many parameters get updated (which drives training memory and storage) and how close the result gets to full fine-tuning.

ApproachFull nameWhat gets updatedRelative memoryTraining speedQuality referenceTypical use
Full fine-tuning (FFT)Full Fine-TuningAll model parametersBaseline (highest)Slowest100% (the upper-bound benchmark)Top teams with ample data and compute
Layer-wise FTLayer-wise FTUnfreeze only the last few layers / certain modulesSlightly lowerSlightly fasterClose to full FT (when the task maps to top-layer semantics)Legacy approach, rarely used today
Adapter tuningAdapter TuningInsert small networks between Transformer layers and train only thoseDown ~1/2FasterClose to full FT on single tasksMulti-task with a shared backbone (one small parameter set per task)
LoRALow-Rank AdaptationFreeze the backbone; train only the injected low-rank matrices (~0.1%–1% of parameters)Down ~2/3Fast≈ full FT on most tasks, slightly lower on a fewToday's mainstream default choice
QLoRAQuantized LoRAQuantize the backbone to 4-bit and freeze it, then train the LoRA matricesAnother ~1/3 down (fine-tunes 70B-class models on a single GPU)FastOn par with LoRAConsumer GPUs / personal research / tight budgets

How to Read This Table

The far right of the spectrum doesn't mean "best quality" — only "lowest cost." Full fine-tuning remains the reference point for the quality ceiling, but fully fine-tuning a 7B model needs roughly 4×14GB of memory (Adam states and gradients all take memory), whereas QLoRA at the same scale runs on a single 16GB consumer GPU. For the underlying principles and engineering trade-offs, see Inference Optimization and Quantization (the quantization idea) and Deployment and Inference Optimization in Practice.

One more word on adapters: each Transformer layer gets a small Bottleneck network inserted (project down, then project back up), which adds one extra computation at inference time. The upside is that many tasks can share the same frozen backbone, each attaching its own small parameter set — this made adapters quite popular for multi-task settings. But LoRA wins across the board on both quality and implementation simplicity, so today the rule is basically "if LoRA works, skip Adapters."

3. How LoRA Works: Splitting One Big Update into Two Small Matrices ​

LoRA (Low-Rank Adaptation) comes from the 2021 paper LoRA: Low-Rank Adaptation of Large Language Models. It rests on a key observation: in full fine-tuning, the weight update itself can be approximated with a very low rank — pretrained weights W encode vast general knowledge, and task adaptation only needs to perturb them along a few small directions.

3.1 Low-Rank Decomposition ​

Suppose a linear layer (say, an attention projection matrix) has original weights W of shape d×d. Full fine-tuning learns the complete update ΔW (also d×d — hundreds of millions of parameters); LoRA instead forces ΔW to decompose into the product of two small matrices:

ΔW ≈ B × A        (A: d×r, B: r×d, with r ≪ d)

Pretrained weights W (frozen)        Delta branch (only this is trained)
┌─────────────────┐        ┌─────────┐ ┌─────────┐
│   W : d×d       │   +    │  B: d×r │×│  A: r×d │
│  (kept as-is)   │        │         │ │         │
└─────────────────┘        └─────────┘ └─────────┘
        ↓ mergeable at forward/inference time
   W' = W + (alpha / r) × B × A

During training only A and B are updated, so the parameter count drops from d×d to d×r + r×d. With d=4096 and r=8, the delta parameters are only about 1/256 of full fine-tuning. A is initialized with Gaussian random values, B with zeros — so at the start of training B×A=0, the model behaves exactly like the pretrained model, and training is more stable.

python
import torch
import torch.nn as nn
import torch.nn.functional as F

class LoRALinear(nn.Module):
    """Inject a low-rank branch into a linear layer: freeze the main weight, train only A and B"""
    def __init__(self, in_dim: int, out_dim: int, r: int = 8, alpha: int = 16):
        super().__init__()
        self.weight = nn.Parameter(torch.randn(out_dim, in_dim) * 0.02)  # simulates pretrained weights
        self.weight.requires_grad = False                                 # freeze
        self.A = nn.Parameter(torch.randn(in_dim, r) * 0.01)              # Gaussian init
        self.B = nn.Parameter(torch.zeros(r, out_dim))                    # zero init
        self.scaling = alpha / r

    def forward(self, x):
        base = F.linear(x, self.weight)              # frozen backbone: plain forward pass
        delta = (x @ self.A) @ self.B * self.scaling # low-rank delta: only this is trained
        return base + delta

The above is a minimal sketch of the idea; in production just use the Hugging Face PEFT library: LoraConfig(r=8, lora_alpha=16, target_modules=["q_proj","v_proj"]) injects LoRA into the specified modules. For the complete engineering workflow, see Fine-Tune Your Own LLM.

3.2 Choosing the Rank r ​

r valueDelta parameter share (7B model reference)Characteristics
r=1–4TinyVery few update directions; suits ultra-light adaptation on very little data; risk of underfitting
r=8~0.1%The safe default for most tasks
r=16–32~0.2%–0.4%More stable for complex tasks with larger datasets; higher quality ceiling
r=64+LargeClearly diminishing returns; the parameter count approaches full fine-tuning, which defeats the purpose

Rule of thumb: start small (r=8), double and compare once the pipeline runs (16/32), and don't jump straight to 64. Doubling r grows memory and storage roughly linearly, while quality gains usually flatten after 8→16.

Zero Overhead at Inference

The A and B matrices from a finished LoRA run can be merged back into the backbone before inference: W_merged = W + scaling × (B @ A). After merging, the model architecture is identical to the original, with zero extra inference latency or memory (for quantized merging, see Inference Optimization and Quantization). You can also skip merging and store A and B separately — one LoRA weight set per task (usually tens to a few hundred MB), mounted on demand. That is exactly why LoRA is so widely used in multi-task and multi-tenant scenarios.

4. Training Data: The Data Science of Instruction Tuning ​

4.1 Anatomy of Instruction Data (SFT Data) ​

The most mainstream form of fine-tuning is instruction tuning, also called supervised fine-tuning (SFT): you feed the model triplets of "instruction + input + expected output" so it learns to follow instructions and act on them. This is the paradigm that kicked off the fine-tuning boom after ChatGPT launched, originating from the InstructGPT paper.

A typical instruction sample looks like this:

json
[
  {
    "instruction": "Rewrite the following sentence in a more formal, professional tone",
    "input": "This plan looks pretty good — let's give it a try.",
    "output": "The plan is feasible; a trial run is recommended."
  },
  {
    "instruction": "Classify the sentiment of the text below. Output only: positive / neutral / negative",
    "input": "The new client app crashed yet again, and support never responds.",
    "output": "negative"
  }
]

4.2 Scale and Quality ​

Data scaleTypical outcomeNotes
A few hundred samplesObservable behavior changeOnly suits ultra-light "format constraints"; don't expect new capabilities
Thousands to tens of thousandsThe quality bar for the vast majority of tasks200–2,000 characters per sample, covering the task's main forms
Hundreds of thousandsClear capability gainsApproaching "rebuilding a general-purpose assistant"; expensive

Three lessons from practice:

  • Quality matters far more than quantity. 1,000 hand-curated samples often beat 100,000 auto-generated dirty ones. The industry has the "LIMA phenomenon": fine-tuning on just 1,000 high-quality instructions can significantly improve alignment and usability.
  • Diversity beats repetition. Covering different forms, difficulty levels, and edge cases of the task is far more useful than copying the same sample ten times.
  • Outputs must be "what the model should answer." SFT is essentially imitating the expected output distribution; if outputs contain wrong formats or hallucinations, the model will absorb them wholesale.

Data Leakage and Contamination

Public instruction datasets (chat data scraped from the web, shared SFT datasets, etc.) often mix in test-set answers, private text, even profanity. Fine-tuning data must be cleaned, deduplicated, and privacy-scrubbed; otherwise the model will memorize answers or pick up bad habits. For dataset tools and catalogs, see Datasets and Tools.

5. The Complete Workflow: From Data to Production ​

The end-to-end process of fine-tuning a model boils down to four steps:

① Data preparation → ② Base model selection → ③ Training → ④ Evaluation & iteration
  Collect/clean/format into instruction JSON   Pick a model for the task   LoRA hyperparameter experiments   Ship/regress/back to ①

① Data preparation: Collect real business samples → clean and deduplicate → write instruction JSON → split into train/val (keep 5%–10% for val). This step eats up over 60% of the project's total time.

② Base model selection: Bigger isn't automatically better; follow three rules of thumb — match the task language to the base model's language (for Chinese tasks, prefer base models with strong Chinese corpora); the base model's knowledge and capability floor must be high enough (fine-tuning changes behavior, not the boundaries of what the model knows — see the risks section below); prefer open-source base models to reduce closed-source dependence. For a quick model reference, see Models and Leaderboards at a Glance.

③ Training: The hyperparameter window for LoRA fine-tuning is narrow; common starting points:

learning_rate = 2e-4      # LoRA typically uses 1e-4 ~ 5e-4, one notch higher than full fine-tuning
num_epochs    = 3         # 1~3 epochs is enough for small data; more will overfit
batch_size    = 4 ~ 16    # bounded by GPU memory; can pair with gradient accumulation
lora_r        = 8         # commonly 8 / 16 / 32
lora_alpha    = 16        # usually 1~2× r
max_seq_len   = 2048      # set to your task's longest input; don't blindly max it out
warmup_steps  = 100       # a short warmup stabilizes training

④ Evaluation and iteration: After training, absolutely do not look only at the training loss — check validation-set performance and business metrics (format compliance rate, task accuracy, style similarity). Regress every historical fine-tuned version against a fixed eval set to prevent "fixing A and breaking B." For evaluation methodology, see LLM Evaluation and Benchmarks; for the tooling, see Building an LLM Evaluation Pipeline.

The One-Line Workflow Mantra

80% of a fine-tuning project's effort goes into the first two stages (data + evaluation design); actual GPU time is usually a small fraction. Build the eval set first; talk about training second.

6. Risks and Pitfalls ​

6.1 Catastrophic Forgetting ​

When a model trains hard on new-task data, it can "wash out" the general abilities learned during pretraining — questions it used to answer correctly start failing, and its common sense degrades. Usual triggers: data that is too small or too narrow, a learning rate that is too high, or too many epochs. Mitigations: keep epochs in check, mix in a small amount of general data (e.g., blend general instruction data at 5%–10%), and note that LoRA itself has a natural buffer against forgetting because it changes so few parameters.

6.2 Overfitting Small Datasets ​

Training for dozens of epochs on a few thousand samples, watching validation loss climb instead of fall, seeing perfect training-set performance while the real task falls apart — this is the most classic fine-tuning failure scene. Countermeasures: with small data, use a larger learning rate + fewer epochs + early stopping (monitor validation loss), and spot-check generation quality with random samples during training.

6.3 The Biggest Misconception: Fine-Tuning ≠ Learning New Knowledge ​

This is the point that bears repeating most. Fine-tuning mainly changes "behavior" (format, style, instruction-following) and almost never injects "knowledge" — during training the model may memorize scattered facts from the training set, but it cannot reliably learn a large body of new domain knowledge; real knowledge comes from the trillion-scale corpora of pretraining and is backed at runtime by Retrieval-Augmented Generation (RAG).

  • Want the model to answer questions about new policies from after May 2026 → use RAG; fine-tuning can't fix recency;
  • Want the model to master your company's 100,000 internal documents → use RAG; fine-tuning only helps it get the "way of answering questions about those documents" right;
  • Want a stable output format and a voice that fits your brand → that's fine-tuning's comfort zone.

Rule of thumb: knowledge problems go to RAG, behavior problems go to fine-tuning, and combining the two is the production norm. For a fuller catalog of traps, see Common Pitfalls and Anti-Patterns.

7. Fine-Tuning and Alignment: Two Sides of Post-Training ​

Fine-tuning isn't an isolated step; it's one member of the "post-training" family. The full lifecycle of an industry LLM looks like this:

Pretraining (learning language from massive corpora) → Post-training (SFT fine-tuning → alignment via RLHF/DPO) → Deployment & inference

Alignment is essentially the second half of fine-tuning: SFT first teaches the model to "follow instructions," then RLHF/DPO tunes it to be "safe, honest, and helpful" — both continue training on the pretrained model; only the objective function changes, from "imitating correct outputs" to "optimizing human preferences." So "fine-tuning" and "alignment" aren't two independent concepts but different stages of the same post-training pipeline; for the deep dive, see Alignment: RLHF and DPO.

Typical domain fine-tuning cases:

DomainRepresentativesKey practices
CodeGitHub Copilot, DeepSeek-Coder, Code LlamaSFT on code-completion data (source code + comments + tests); for a complete breakdown of code-intelligence products, see GitHub Copilot and Code Intelligence
HealthcareMed-PaLM, open-source medical base modelsFine-tune on curated data from medical records, medical QA, and clinical guidelines, while also doing safety alignment
Customer service / financeVarious industry LLMsFine-tune on historical tickets and standard scripts to build a "house voice," then layer RAG on top to supply product knowledge
ReasoningDeepSeek-R1SFT on chain-of-thought reasoning data, then reinforcement learning to elicit deep thinking; see DeepSeek-R1 and Reasoning Models

Beyond that, fine-tuned models are often combined with AI Agents — the stable output format produced by fine-tuning serves as the output constraint for tool calling and task planning. For the overall learning roadmap, see Learning Paths.

Further Reading ​

References ​