Theme
Fine-Tuning in Practice: The Full LoRA Workflow
Fine-tuning isn't "feeding new data to a model"; it's using a controlled training process to modify the model's weight distribution — LoRA's contribution is compressing "updating all weights" into "updating two low-rank matrices," reducing the cost of fine-tuning from "exclusive to large labs" to "doable on a single GPU."
This article provides a complete, copy-paste-ready LoRA workflow from data to evaluation. For the "why LoRA works" theory (low-rank decomposition W = W0 + BA, meaning of r and α), see Fine-Tuning: SFT and Parameter-Efficient Fine-Tuning; for alignment (RLHF/DPO), see Alignment.
First Question: Should You Fine-tune?
Fine-tuning is the most expensive and risky model modification. Before starting, ask yourself using the checklist below:
| Your need | Better approach | When to fine-tune |
|---|---|---|
| Let the model know your private documents/facts | RAG (retrieval-augmented generation) | Fine-tuning is not for injecting knowledge |
| Make output follow a specific format | Prompts + structured output | Prompts hit limits and format must be "inherent" |
| Make output style like someone/some text type | Few-shot examples | Examples don't fit/unstable |
| Let the model master domain-specific tasks and terminology | — | ✅ This is fine-tuning's home turf |
One-sentence criterion: fine-tuning changes "the model's behavioral habits" (style, format, task terminology, persona), not "the model's memory" (knowledge). For knowledge problems, go with RAG in Practice. If your goal is actually style/format, read Prompting in Practice first before deciding.
Three "Fine-tuning is Worse Than Prompting" Signals
- Data < hundreds of examples: prompting/examples may be more cost-effective;
- The need is "a specific fact": RAG suffices, fine-tuning also brings forgetting risk;
- No evaluation method: when you can't quantify before/after fine-tuning, first add Evaluation in Practice.
Data Preparation
1. Format: The Conversation Template
The mainstream format for LLM fine-tuning is the conversation format (list of role/content pairs). This corresponds to the Hugging Face transformers conversation template (apply_chat_template), which automatically adds special tokens like <|im_start|> and <|im_end|> during training. During training, loss is computed on assistant tokens only; user/system parts are typically masked (TRL's SFTTrainer handles this automatically).
json
[
{
"messages": [
{"role": "user", "content": "What is the capital of France?"},
{"role": "assistant", "content": "The capital of France is Paris."}
]
}
]2. Data Quantity and Ratio Guidelines
| Data Dimension | Recommended Range | Notes |
|---|---|---|
| Total count | Hundreds to tens of thousands | LoRA fine-tuning commonly uses thousands; hundreds can work with high quality |
| Repetition | 2~3 epochs, to prevent overfitting | More than 3 epochs easily overfits (see Section V) |
| Length | Match the target scenario | Too-short max_seq_length during training truncates long responses |
| Quality filtering | Deduplicate, remove noise, correct facts | One bad annotation outweighs ten good ones |
Evidence: Data Quality > Data Quantity
Industry repeatedly validates: with the same budget, 500 high-quality, uniformly formatted, deduplicated samples significantly outperform 5000 automatically scraped dirty samples. Data cleaning effort (deduplication, toxicity removal, fact-checking) directly determines the fine-tuning ceiling, echoing the pretraining principle that "data is the model" — except here scale is replaced by precision.
Common Data Augmentation and Denoising Methods
| Method | Practice | Purpose |
|---|---|---|
| Deduplication | Remove by text similarity / embedding | Prevent overfitting and memorization |
| Quality filtering | Remove garbled, ultra-long, meaningless responses | Improve training signal |
| Format standardization | Apply the same conversation template and punctuation rules | Improve output consistency |
| Construct negative samples | Add "this is not how to answer" samples | Clarify boundaries, reduce hallucination |
| Multi-turn expansion | Manually expand single-turn Q&A to multi-turn conversations | Improve multi-turn consistency |
Be restrained with negative samples
Negative samples ("this is a wrong example") in excessive amounts make the model overly cautious and increase refusal rates. Experience shows positive samples should dominate, with negative samples as garnish (about 5~10%), and manually check results.
3. Base Model Selection
| Dimension | Recommendation |
|---|---|
| Capability baseline | The base determines the fine-tuning ceiling; a weak base cannot produce new capabilities through fine-tuning; test prompting first before fine-tuning |
| Parameters vs memory | 7B/8B trainable on single GPU; 70B class needs multi-GPU or QLoRA (see Section VIII) |
| Open-source license | Note model licenses for commercial use (e.g., Llama community license terms) |
| Chinese scenario | Prioritize strong Chinese bases (Qwen series, etc.); Chinese fine-tuning cost is lower |
| Alignment level | Use aligned base (chat version) for fine-tuning tasks; "pure capability" experiments can use base version |
Model profiles and selection at Mainstream Models Archive.
4. LoRA Configuration: Four Key Hyperparameters
LoRA constrains weight updates to low-rank: W' = W0 + (B·A)·α/r. Core hyperparameters:
| Hyperparameter | Meaning | Common Starting Point | Adjustment Direction |
|---|---|---|---|
r (rank) | Rank of low-rank matrix, determines learnable parameter count | 8~16 | Increase for harder tasks/more data (32/64), otherwise overfits |
lora_alpha | Scaling coefficient, controls update intensity | 16~32 (about 2x r) | Increase if weak effect, decrease if overfitting |
target_modules | Modules to inject LoRA | Start with q_proj, v_proj; add k_proj, o_proj, gate_proj, up_proj, down_proj for harder tasks | Wider coverage = stronger but more prone to overfitting |
| Learning rate | Learning rate for LoRA layer | 1e-4 ~ 2e-4 | Reduce to 5e-5 or 3e-5 if unstable |
dropout | Dropout for LoRA layer | 0.05~0.1 | Increase if overfitting |
Complete Training Script (HF PEFT + TRL, Ready to Run)
python
# 1. Base model and tokenizer
model_id = "Qwen/Qwen2.5-7B-Instruct" # Using Chinese 7B as example
# 2. LoRA config
from peft import LoraConfig, get_peft_config
lora_config = LoraConfig(
r=16, # rank
lora_alpha=32, # scaling
target_modules=["q_proj", "v_proj"],
lora_dropout=0.05,
bias="none",
task_type="CAUSAL_LM"
)
# 3. Training parameters (defaults from LLaMA-Factory and similar tools are derived from the same principles)
training_args = dict(
per_device_train_batch_size=1, # gradient accumulation when memory is insufficient
gradient_accumulation_steps=8, # effective batch size = 1 × 8 = 8
learning_rate=2e-4,
num_train_epochs=2,
fp16=True, # use fp16 for older GPUs that don't support bf16
logging_steps=10,
optim="paged_adamw_8bit", # save memory with 8-bit optimizer
warmup_ratio=0.05,
weight_decay=0.01,
max_grad_norm=1.0, # gradient clipping
lr_scheduler_type="cosine", # cosine decay
report_to="tensorboard", # open TensorBoard
save_strategy="steps",
save_steps=100,
output_dir="outputs",
seed=42,
)
# 4. Training (automatically applies conversation template; computes loss on assistant part only)python
# Newer TRL automatically applies conversation template, only computes loss on assistant part;
# Older versions can specify field with: dataset_text_field="messages"How to calculate effective batch size
Effective batch size = per_device_train_batch_size × gradient_accumulation_steps × num_devices. When memory is insufficient, prioritize reducing batch size and use gradient accumulation to compensate — large batch training is more stable; this is the most common approach in LoRA fine-tuning.
5. Training Monitoring: Watch for Overfitting
1. Three Things to Always Do
- Open TensorBoard (
report_to="tensorboard"is already set): watch both the train loss and eval loss curves. - Regular manual verification: every N steps load the latest checkpoint and run 5~10 real questions to check output. Loss cannot replace human judgment.
- Record a baseline: before fine-tuning, first run the same 10 questions on the base model, save outputs for comparison.
2. Loss Curve Interpretation
| Curve Shape | Meaning | Action |
|---|---|---|
| train / eval loss both decrease | Normal learning | Continue |
| train loss continues decreasing, eval loss turns upward | Overfitting | Early stop (use best eval checkpoint), reduce epochs, add data |
| Both high and not decreasing | Learning rate / data issue | Check learning rate, data format, conversation template |
| Loss fluctuates wildly | Learning rate too high / batch too small | Lower learning rate, increase accumulation |
Direct evidence of overfitting: memorizing training data
One manifestation of fine-tuning overfitting is the model starts "reciting" long sentences from the training set — for questions that appeared in the training data, the response is almost word-for-word identical. This usually means too many epochs or r is too large. The rule is: LoRA has few learnable parameters, and overfitting is mainly controlled by epochs and r.
3. Hyperparameter Quick-Reference Table
| Hyperparameter | Common Range | Effect Logic |
|---|---|---|
| Learning rate | 1e-4 ~ 3e-4 | Too high → divergence, too low → won't learn |
| Warmup steps | 1~5% of total steps | Stabilize early training (essential for large models) |
| Weight decay | 0.01~0.1 | Weight regularization, suppress overfitting |
| Gradient clipping | 1.0 (max norm) | Prevent loss spikes |
| Cosine decay | Commonly used | More stable convergence in later stages |
| Batch size (effective) | 8~64 | Larger batch = more stable, faster convergence |
A reproducible starting point
lr=2e-4 + warmup(100 steps) + weight_decay=0.01 + gradient clipping 1.0 + cosine decay + epoch=2. Start with this combination, then fine-tune by Section V curves. Most LoRA tasks need only lr and epoch changes on top of this.
Merging and Export
LoRA weights default to "side-mounted" low-rank matrices (PeftModel), and before deployment you typically need to merge into the main model:
python
from peft import PeftModel
# Load LoRA adapter and merge
base = AutoModelForCausalLM.from_pretrained(model_id)
peft = PeftModel.from_pretrained(base, "outputs/lora-checkpoint")
merged = peft.merge_and_unload() # Merge back into main model weights
merged.save_pretrained("outputs/model-merged") # Export as standard HF modelThe merged model can directly go through Deployment and Serving's vLLM and other inference frameworks. You can also skip merging: deploy using services that support PeftModel (e.g., vLLM supports --enable_lora) to dynamically mount at inference time, saving memory but slightly increasing complexity.
Evaluation Comparison: Before Fine-Tuning vs After
Whether fine-tuning succeeded, must compare using the same evaluation set before and after fine-tuning:
| Comparison Item | Practice |
|---|---|
| Fixed evaluation set | Freeze 50~200 samples before fine-tuning (including a held-out set outside the training set) |
| Metrics | Task accuracy + style/format manual score + general capability regression (prevent "catastrophic forgetting") |
| Control dimensions | Base vs fine-tuned; optionally also compare "prompting only" version |
| General capability regression | Run MMLU subset or other general benchmarks to confirm fine-tuning didn't wash away base capabilities |
python
# General capability regression example: run merged model with lm-evaluation-harnessThe evaluation system (golden set, regression, LLM-as-a-judge) is at Evaluation in Practice.
Multi-Round Iteration: Fine-Tuning Is Not a One-Time Activity
One fine-tuning rarely reaches the target directly; the standard flow is "fine-tune → evaluate → find gaps → add data → fine-tune again" cycle:
| Round | Goal | Common Actions |
|---|---|---|
| Round 1 | Verify data format and pipeline | Run with minimal data (e.g., 200 samples) |
| Round 2 | Approach target effect | Expand data, tune lr/epochs, add negative samples |
| Round 3 | Fix systematic errors found by evaluation | Add targeted data for error types (see error analysis) |
| Final | Prevent forgetting regression | Run general benchmarks, confirm capabilities unchanged |
Don't endlessly tune hyperparameters on the "same data"
If scores stop improving after Round 2, the problem is usually data (not enough, imbalanced, inconsistent annotations) rather than hyperparameters. At this point, add data or change annotation criteria instead of continuing to search for the learning rate.
8. QLoRA Memory Estimation
QLoRA freezes the main model at 4-bit quantization and trains only LoRA parameters, dramatically reducing memory requirements. Rule of thumb:
Inference memory ≈ weight bytes (2 bytes/parameter, fp16)
QLoRA training memory ≈ weights at 4-bit (0.5 bytes/parameter) + LoRA gradients/optimizer + activations + headroom| Model | Full-parameter (fp16) | LoRA (fp16) | QLoRA (4-bit) | What one 24GB card can do |
|---|---|---|---|---|
| 0.5B | ~1GB | ~2~3GB | ~2~3GB | Easy |
| 7B/8B | ~14GB (just enough) | ~16~18GB | ~6~8GB | ✅ Common combination |
| 13B | ~26GB (exceeds) | ~30GB+ (exceeds) | ~10~12GB | ✅ QLoRA can train |
| 70B | ~140GB (needs multi-GPU) | ~160GB (needs multi-GPU) | ~45~50GB | Single-GPU limit; multi-GPU more stable |
Note on numbers
Above are estimated ranges (training memory also affected by batch size, sequence length, activation overhead); actual values depend on testing. The QLoRA paper's benchmark conclusion: can fine-tune a 65B model on a single 48GB card (References). Doubling sequence length significantly increases activation memory; this is the most often overlooked variable on small cards.
9. Common Tools Comparison
| Tool | Purpose | Strengths | Suitable For |
|---|---|---|---|
HF peft + transformers | Low-level library | Flexible, controllable, broadest ecosystem | Custom pipelines, learning principles |
HF trl (SFTTrainer, etc.) | High-level trainer | Conversation templates, DPO/PPO family | SFT/alignment integrated |
| LLaMA-Factory | One-click training platform | Chinese-friendly, WebUI, multiple methods out-of-the-box | Quick experiments, parameter tuning |
| Axolotl | Config-driven trainer | YAML config, supports multi-GPU and advanced tricks | Reproducing research configs, batch experiments |
Advice for beginners
Round one: use LLaMA-Factory or TRL to run through the full pipeline (half a day); Round two: drop to peft and look at each parameter's role. Tools are means; understanding "training objective + data + hyperparameters" is the core.
10. Common Failures and Troubleshooting
| Failure Symptom | Root Cause | Fix |
|---|---|---|
| Out-of-memory OOM | batch/sequence too long, activation explosion | Reduce per_device_train_batch_size, max_seq_length; use QLoRA |
| Training looks fine but output unchanged | LoRA not activated (wrong target_modules), LR too low, epochs too few | Print trainable parameter count to verify; check target_modules vs actual layer names |
| Loss decreases but effect worsens | Data quality issue / inconsistent data format | Manually spot-check 20 samples; standardize template and annotations |
| Irrelevant/nonsensical answers | Over-fine-tuning or data contamination | Roll back to best eval checkpoint; reduce r and epochs |
| Chinese gets worse | Base model weak in Chinese, training data lacks Chinese | Switch to Chinese-capable base; check data language ratio |
| Conversation template not working | Tokenizer lacks apply_chat_template | Confirm model has conversation template; use newer transformers version |
Further Reading
- Fine-Tuning: SFT and Parameter-Efficient Fine-Tuning — LoRA principle (low-rank decomposition, α/r), full-parameter vs LoRA comparison
- Alignment: RLHF and DPO — Post-SFT alignment steps, when needed
- Evaluation in Practice — Before/after comparison and regression testing methods
- Prompting in Practice — "Zero-cost" methods to try before fine-tuning
- RAG in Practice — Alternative for knowledge-based needs
- Datasets and Benchmarks Archive — Instruction data (Alpaca/ShareGPT, etc.) and evaluation benchmarks
- Deployment and Serving — Production path for fine-tuned models
References
- LoRA: Low-Rank Adaptation of Large Language Models (arXiv:2106.09685) — Original LoRA paper
- QLoRA: Efficient Finetuning of Quantized LLMs (arXiv:2305.14314) — Original QLoRA paper, includes memory benchmarks
- Hugging Face PEFT (GitHub) — Official library for LoRA and other parameter-efficient fine-tuning methods
- Hugging Face TRL (GitHub) — Official library for SFTTrainer / DPOTrainer
- LLaMA-Factory (GitHub) — One-click fine-tuning platform
- Axolotl (GitHub) — Config-driven fine-tuning framework
- Hugging Face Transformers Documentation — Model loading, conversation templates, trainer official docs
- lm-evaluation-harness (GitHub) — General benchmark regression testing tool for before/after fine-tuning comparison