Skip to content

Fine-Tuning in Practice: The Full LoRA Workflow

At a glance Walk through the complete fine-tuning pipeline with LoRA — data preparation, base model selection, LoRA configuration, training monitoring, merging and export, evaluation comparison — including QLoRA VRAM estimation, common tool comparison, failure troubleshooting, and criteria for "fine-tuning vs prompting."

Fine-Tuning in Practice: The Full LoRA Workflow ​

Fine-tuning isn't "feeding new data to a model"; it's using a controlled training process to modify the model's weight distribution — LoRA's contribution is compressing "updating all weights" into "updating two low-rank matrices," reducing the cost of fine-tuning from "exclusive to large labs" to "doable on a single GPU."

This article provides a complete, copy-paste-ready LoRA workflow from data to evaluation. For the "why LoRA works" theory (low-rank decomposition W = W0 + BA, meaning of r and α), see Fine-Tuning: SFT and Parameter-Efficient Fine-Tuning; for alignment (RLHF/DPO), see Alignment.

First Question: Should You Fine-tune? ​

Fine-tuning is the most expensive and risky model modification. Before starting, ask yourself using the checklist below:

Your needBetter approachWhen to fine-tune
Let the model know your private documents/factsRAG (retrieval-augmented generation)Fine-tuning is not for injecting knowledge
Make output follow a specific formatPrompts + structured outputPrompts hit limits and format must be "inherent"
Make output style like someone/some text typeFew-shot examplesExamples don't fit/unstable
Let the model master domain-specific tasks and terminology—✅ This is fine-tuning's home turf

One-sentence criterion: fine-tuning changes "the model's behavioral habits" (style, format, task terminology, persona), not "the model's memory" (knowledge). For knowledge problems, go with RAG in Practice. If your goal is actually style/format, read Prompting in Practice first before deciding.

Three "Fine-tuning is Worse Than Prompting" Signals

  1. Data < hundreds of examples: prompting/examples may be more cost-effective;
  2. The need is "a specific fact": RAG suffices, fine-tuning also brings forgetting risk;
  3. No evaluation method: when you can't quantify before/after fine-tuning, first add Evaluation in Practice.

Data Preparation ​

1. Format: The Conversation Template ​

The mainstream format for LLM fine-tuning is the conversation format (list of role/content pairs). This corresponds to the Hugging Face transformers conversation template (apply_chat_template), which automatically adds special tokens like <|im_start|> and <|im_end|> during training. During training, loss is computed on assistant tokens only; user/system parts are typically masked (TRL's SFTTrainer handles this automatically).

json
[
  {
    "messages": [
      {"role": "user", "content": "What is the capital of France?"},
      {"role": "assistant", "content": "The capital of France is Paris."}
    ]
  }
]

2. Data Quantity and Ratio Guidelines ​

Data DimensionRecommended RangeNotes
Total countHundreds to tens of thousandsLoRA fine-tuning commonly uses thousands; hundreds can work with high quality
Repetition2~3 epochs, to prevent overfittingMore than 3 epochs easily overfits (see Section V)
LengthMatch the target scenarioToo-short max_seq_length during training truncates long responses
Quality filteringDeduplicate, remove noise, correct factsOne bad annotation outweighs ten good ones

Evidence: Data Quality > Data Quantity

Industry repeatedly validates: with the same budget, 500 high-quality, uniformly formatted, deduplicated samples significantly outperform 5000 automatically scraped dirty samples. Data cleaning effort (deduplication, toxicity removal, fact-checking) directly determines the fine-tuning ceiling, echoing the pretraining principle that "data is the model" — except here scale is replaced by precision.

Common Data Augmentation and Denoising Methods ​

MethodPracticePurpose
DeduplicationRemove by text similarity / embeddingPrevent overfitting and memorization
Quality filteringRemove garbled, ultra-long, meaningless responsesImprove training signal
Format standardizationApply the same conversation template and punctuation rulesImprove output consistency
Construct negative samplesAdd "this is not how to answer" samplesClarify boundaries, reduce hallucination
Multi-turn expansionManually expand single-turn Q&A to multi-turn conversationsImprove multi-turn consistency

Be restrained with negative samples

Negative samples ("this is a wrong example") in excessive amounts make the model overly cautious and increase refusal rates. Experience shows positive samples should dominate, with negative samples as garnish (about 5~10%), and manually check results.

3. Base Model Selection ​

DimensionRecommendation
Capability baselineThe base determines the fine-tuning ceiling; a weak base cannot produce new capabilities through fine-tuning; test prompting first before fine-tuning
Parameters vs memory7B/8B trainable on single GPU; 70B class needs multi-GPU or QLoRA (see Section VIII)
Open-source licenseNote model licenses for commercial use (e.g., Llama community license terms)
Chinese scenarioPrioritize strong Chinese bases (Qwen series, etc.); Chinese fine-tuning cost is lower
Alignment levelUse aligned base (chat version) for fine-tuning tasks; "pure capability" experiments can use base version

Model profiles and selection at Mainstream Models Archive.

4. LoRA Configuration: Four Key Hyperparameters ​

LoRA constrains weight updates to low-rank: W' = W0 + (B·A)·α/r. Core hyperparameters:

HyperparameterMeaningCommon Starting PointAdjustment Direction
r (rank)Rank of low-rank matrix, determines learnable parameter count8~16Increase for harder tasks/more data (32/64), otherwise overfits
lora_alphaScaling coefficient, controls update intensity16~32 (about 2x r)Increase if weak effect, decrease if overfitting
target_modulesModules to inject LoRAStart with q_proj, v_proj; add k_proj, o_proj, gate_proj, up_proj, down_proj for harder tasksWider coverage = stronger but more prone to overfitting
Learning rateLearning rate for LoRA layer1e-4 ~ 2e-4Reduce to 5e-5 or 3e-5 if unstable
dropoutDropout for LoRA layer0.05~0.1Increase if overfitting

Complete Training Script (HF PEFT + TRL, Ready to Run) ​

python
# 1. Base model and tokenizer
model_id = "Qwen/Qwen2.5-7B-Instruct"        # Using Chinese 7B as example

# 2. LoRA config
from peft import LoraConfig, get_peft_config
lora_config = LoraConfig(
    r=16,                              # rank
    lora_alpha=32,                     # scaling
    target_modules=["q_proj", "v_proj"],
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM"
)

# 3. Training parameters (defaults from LLaMA-Factory and similar tools are derived from the same principles)
training_args = dict(
    per_device_train_batch_size=1,     # gradient accumulation when memory is insufficient
    gradient_accumulation_steps=8,     # effective batch size = 1 × 8 = 8
    learning_rate=2e-4,
    num_train_epochs=2,
    fp16=True,                          # use fp16 for older GPUs that don't support bf16
    logging_steps=10,
    optim="paged_adamw_8bit",           # save memory with 8-bit optimizer
    warmup_ratio=0.05,
    weight_decay=0.01,
    max_grad_norm=1.0,                  # gradient clipping
    lr_scheduler_type="cosine",         # cosine decay
    report_to="tensorboard",            # open TensorBoard
    save_strategy="steps",
    save_steps=100,
    output_dir="outputs",
    seed=42,
)

# 4. Training (automatically applies conversation template; computes loss on assistant part only)
python
# Newer TRL automatically applies conversation template, only computes loss on assistant part;
# Older versions can specify field with: dataset_text_field="messages"

How to calculate effective batch size

Effective batch size = per_device_train_batch_size × gradient_accumulation_steps × num_devices. When memory is insufficient, prioritize reducing batch size and use gradient accumulation to compensate — large batch training is more stable; this is the most common approach in LoRA fine-tuning.

5. Training Monitoring: Watch for Overfitting ​

1. Three Things to Always Do ​

  • Open TensorBoard (report_to="tensorboard" is already set): watch both the train loss and eval loss curves.
  • Regular manual verification: every N steps load the latest checkpoint and run 5~10 real questions to check output. Loss cannot replace human judgment.
  • Record a baseline: before fine-tuning, first run the same 10 questions on the base model, save outputs for comparison.

2. Loss Curve Interpretation ​

Curve ShapeMeaningAction
train / eval loss both decreaseNormal learningContinue
train loss continues decreasing, eval loss turns upwardOverfittingEarly stop (use best eval checkpoint), reduce epochs, add data
Both high and not decreasingLearning rate / data issueCheck learning rate, data format, conversation template
Loss fluctuates wildlyLearning rate too high / batch too smallLower learning rate, increase accumulation

Direct evidence of overfitting: memorizing training data

One manifestation of fine-tuning overfitting is the model starts "reciting" long sentences from the training set — for questions that appeared in the training data, the response is almost word-for-word identical. This usually means too many epochs or r is too large. The rule is: LoRA has few learnable parameters, and overfitting is mainly controlled by epochs and r.

3. Hyperparameter Quick-Reference Table ​

HyperparameterCommon RangeEffect Logic
Learning rate1e-4 ~ 3e-4Too high → divergence, too low → won't learn
Warmup steps1~5% of total stepsStabilize early training (essential for large models)
Weight decay0.01~0.1Weight regularization, suppress overfitting
Gradient clipping1.0 (max norm)Prevent loss spikes
Cosine decayCommonly usedMore stable convergence in later stages
Batch size (effective)8~64Larger batch = more stable, faster convergence

A reproducible starting point

lr=2e-4 + warmup(100 steps) + weight_decay=0.01 + gradient clipping 1.0 + cosine decay + epoch=2. Start with this combination, then fine-tune by Section V curves. Most LoRA tasks need only lr and epoch changes on top of this.

Merging and Export ​

LoRA weights default to "side-mounted" low-rank matrices (PeftModel), and before deployment you typically need to merge into the main model:

python
from peft import PeftModel

# Load LoRA adapter and merge
base = AutoModelForCausalLM.from_pretrained(model_id)
peft = PeftModel.from_pretrained(base, "outputs/lora-checkpoint")
merged = peft.merge_and_unload()              # Merge back into main model weights
merged.save_pretrained("outputs/model-merged") # Export as standard HF model

The merged model can directly go through Deployment and Serving's vLLM and other inference frameworks. You can also skip merging: deploy using services that support PeftModel (e.g., vLLM supports --enable_lora) to dynamically mount at inference time, saving memory but slightly increasing complexity.

Evaluation Comparison: Before Fine-Tuning vs After ​

Whether fine-tuning succeeded, must compare using the same evaluation set before and after fine-tuning:

Comparison ItemPractice
Fixed evaluation setFreeze 50~200 samples before fine-tuning (including a held-out set outside the training set)
MetricsTask accuracy + style/format manual score + general capability regression (prevent "catastrophic forgetting")
Control dimensionsBase vs fine-tuned; optionally also compare "prompting only" version
General capability regressionRun MMLU subset or other general benchmarks to confirm fine-tuning didn't wash away base capabilities
python
# General capability regression example: run merged model with lm-evaluation-harness

The evaluation system (golden set, regression, LLM-as-a-judge) is at Evaluation in Practice.

Multi-Round Iteration: Fine-Tuning Is Not a One-Time Activity ​

One fine-tuning rarely reaches the target directly; the standard flow is "fine-tune → evaluate → find gaps → add data → fine-tune again" cycle:

RoundGoalCommon Actions
Round 1Verify data format and pipelineRun with minimal data (e.g., 200 samples)
Round 2Approach target effectExpand data, tune lr/epochs, add negative samples
Round 3Fix systematic errors found by evaluationAdd targeted data for error types (see error analysis)
FinalPrevent forgetting regressionRun general benchmarks, confirm capabilities unchanged

Don't endlessly tune hyperparameters on the "same data"

If scores stop improving after Round 2, the problem is usually data (not enough, imbalanced, inconsistent annotations) rather than hyperparameters. At this point, add data or change annotation criteria instead of continuing to search for the learning rate.

8. QLoRA Memory Estimation ​

QLoRA freezes the main model at 4-bit quantization and trains only LoRA parameters, dramatically reducing memory requirements. Rule of thumb:

Inference memory ≈ weight bytes (2 bytes/parameter, fp16)
QLoRA training memory ≈ weights at 4-bit (0.5 bytes/parameter) + LoRA gradients/optimizer + activations + headroom
ModelFull-parameter (fp16)LoRA (fp16)QLoRA (4-bit)What one 24GB card can do
0.5B~1GB~2~3GB~2~3GBEasy
7B/8B~14GB (just enough)~16~18GB~6~8GB✅ Common combination
13B~26GB (exceeds)~30GB+ (exceeds)~10~12GB✅ QLoRA can train
70B~140GB (needs multi-GPU)~160GB (needs multi-GPU)~45~50GBSingle-GPU limit; multi-GPU more stable

Note on numbers

Above are estimated ranges (training memory also affected by batch size, sequence length, activation overhead); actual values depend on testing. The QLoRA paper's benchmark conclusion: can fine-tune a 65B model on a single 48GB card (References). Doubling sequence length significantly increases activation memory; this is the most often overlooked variable on small cards.

9. Common Tools Comparison ​

ToolPurposeStrengthsSuitable For
HF peft + transformersLow-level libraryFlexible, controllable, broadest ecosystemCustom pipelines, learning principles
HF trl (SFTTrainer, etc.)High-level trainerConversation templates, DPO/PPO familySFT/alignment integrated
LLaMA-FactoryOne-click training platformChinese-friendly, WebUI, multiple methods out-of-the-boxQuick experiments, parameter tuning
AxolotlConfig-driven trainerYAML config, supports multi-GPU and advanced tricksReproducing research configs, batch experiments

Advice for beginners

Round one: use LLaMA-Factory or TRL to run through the full pipeline (half a day); Round two: drop to peft and look at each parameter's role. Tools are means; understanding "training objective + data + hyperparameters" is the core.

10. Common Failures and Troubleshooting ​

Failure SymptomRoot CauseFix
Out-of-memory OOMbatch/sequence too long, activation explosionReduce per_device_train_batch_size, max_seq_length; use QLoRA
Training looks fine but output unchangedLoRA not activated (wrong target_modules), LR too low, epochs too fewPrint trainable parameter count to verify; check target_modules vs actual layer names
Loss decreases but effect worsensData quality issue / inconsistent data formatManually spot-check 20 samples; standardize template and annotations
Irrelevant/nonsensical answersOver-fine-tuning or data contaminationRoll back to best eval checkpoint; reduce r and epochs
Chinese gets worseBase model weak in Chinese, training data lacks ChineseSwitch to Chinese-capable base; check data language ratio
Conversation template not workingTokenizer lacks apply_chat_templateConfirm model has conversation template; use newer transformers version

Further Reading ​

References ​