Appearance
Fine-Tune Your Own LLM
Fine-tuning means taking a pretrained large model and continuing to train it on your own data, so the model "learns" your business language, output formats, and behavioral preferences. Thanks to parameter-efficient fine-tuning techniques like LoRA (Low-Rank Adaptation), this is no longer the exclusive domain of a few big labs: with one consumer GPU with 8 GB of VRAM, a few thousand to tens of thousands of high-quality instruction samples, and an afternoon, you can turn Qwen or Llama into a little assistant that understands your business.
This article doesn't rehash the theory of fine-tuning. Instead it answers three practical questions: Should your use case be fine-tuned at all? How do you prepare VRAM and the software environment? What does the complete pipeline from data to deployment look like? What follows is an operations manual you can copy verbatim — data preparation, LoRA training, merge and export, and inference validation, with runnable code at every step. For the theory behind fine-tuning (why LoRA works, what low-rank approximation actually is), see Fine-Tuning and PEFT (LoRA); this article focuses on how to get it running.
Prerequisites: you should already understand the basics of large language models (LLMs) and have hands-on experience running a prompt engineering or RAG application. Every stage in this article has a corresponding in-depth article on this site; links are provided throughout.
1. Decide First: Is Your Use Case Worth Fine-Tuning
1. The One-Sentence Test
Prompting first, RAG second, and only then consider fine-tuning. Fine-tuning isn't "more advanced" — it's "more expensive": it demands data, compute, and iteration time, and every change requires retraining. The vast majority of business problems are already solved at the prompting and RAG stages.
| Dimension | Prompt Engineering | RAG | Fine-Tuning (LoRA) |
|---|---|---|---|
| Cost of change | A few minutes; just edit the prompt | Update the retrieval store and index | Hours to days; retrain the model |
| Data required | A few examples (few-shot) | A document corpus | Thousands of high-quality labeled samples |
| Knowledge freshness | Good | Best (changes take effect immediately) | Poor (frozen at training time) |
| Deep domain capability | Weak | Moderate | Strong |
| Private knowledge | Relies on prompt injection | Best fit | Can be baked in, but goes stale easily |
| Output format/style control | Moderate | Weak | Strong |
| Deployment cost | Lowest | Requires a retrieval stack | Larger model and VRAM overhead |
| Typical use cases | General assistants, one-off needs | Support knowledge bases, enterprise Q&A | Domain language, fixed-format output |
Four signals that fine-tuning is worth it:
- The model clearly misunderstands domain terminology, jargon, and proper nouns, and documents retrieved via RAG can't "teach" it (think medicine, law, or industry-specific slang);
- You need a stable, fixed output format (a specific JSON schema, a specific reply style) and prompts keep drifting off;
- You're deploying offline / locally / inside an intranet and can't carry a huge context with every request;
- Inference cost matters: rather than stuffing 20k tokens of context into every request, bake the high-frequency knowledge directly into the weights.
When to hit the brakes:
- Knowledge changes frequently → use RAG instead of fine-tuning;
- Data is scarce or of dubious quality (you can't produce 500+ high-quality samples) → polish your prompts first;
- You have no validation set or evaluation method → read Building an LLM Evaluation first; otherwise you won't be able to tell good results from bad after fine-tuning.
2. How to Choose a Base Model
Pick the wrong base model and everything downstream is wasted effort. Evaluate three dimensions: Chinese vs. English, model size vs. VRAM, instruct model vs. base model. For parameter counts and leaderboard data on each model, see Model and Leaderboard Quick Reference.
| Model | Size | Characteristics | Best for |
|---|---|---|---|
| Qwen2.5-7B-Instruct | 7B | Strong Chinese, good instruction following, mature ecosystem | Chinese-language business, general fine-tuning on-ramp |
| Llama-3.1-8B-Instruct | 8B | Strong English, richest community ecosystem | English tasks, tool calling |
| Mistral-7B-Instruct | 7B | Efficient inference, VRAM-friendly | English, tight resources |
| Qwen2.5-14B / Llama-3.1-13B | 14B / 13B | Stronger but hungrier for VRAM | Large datasets, quality first |
| Qwen2.5-3B / Llama-3.2-3B | 3B | Runs on minimal VRAM | Teaching, edge devices |
Three rules of thumb:
- For Chinese-language business, prefer Qwen; for general English work, prefer Llama. This is experience the community has validated repeatedly — see Model and Leaderboard Quick Reference for detailed comparison data;
- Start at 7B / 8B: on consumer GPUs, 7B is the value-for-money sweet spot; at 13B and above, VRAM and training time climb steeply;
- Choose an "instruct" model, not a base model: the
-Instruct/-Chatvariants have already been aligned (RLHF/DPO) and behave better after fine-tuning; fine-tuning a base model directly means you have to build its conversational ability yourself — generally not recommended.
2. Environment Setup: The VRAM Budget and Dependencies
1. VRAM Requirements Table (Rules of Thumb)
Fine-tuning VRAM consists of four parts: model weights + gradients + optimizer states + activations. Full fine-tuning (FFT) has to maintain all of those complete states at once, so its VRAM demand is typically 5-10x that of QLoRA. The table below gives rule-of-thumb figures (in GB); actual usage varies with max_seq_len and batch_size:
| Model size | Full fine-tuning (BF16) | QLoRA (4-bit + LoRA) | Suggested GPU |
|---|---|---|---|
| 2B / 3B | ~18 GB | ~3-4 GB | 6-8 GB is enough |
| 7B / 8B | ~60-70 GB (needs A100/H100 class) | ~7-9 GB | RTX 4070 / 3090 |
| 13B / 14B | ~120-140 GB (multi-GPU) | ~12-15 GB | RTX 4090 |
| 70B | Multi-node, multi-GPU | ~35-45 GB | A100 80G / multi-GPU |
The conclusion is clear: on consumer GPUs, QLoRA is the only realistic option. QLoRA combines 4-bit quantization (NF4) with LoRA — the base weights are compressed to a quarter of their size and stay frozen throughout, while trainable parameters account for just 0.1%-1% of all parameters. For how quantization works, see Inference Optimization and Quantization.
Not enough VRAM?
Turn three dials in order: drop max_seq_length from 2048 to 1024 → set per_device_train_batch_size to 1 and recover the effective batch size with gradient_accumulation_steps → lower LoRA's r from 16 to 8. Of these, "batch=1 + gradient accumulation" is nearly lossless in quality and should be your first lever.
2. Installing Dependencies
bash
# Python 3.10+ recommended; create an isolated environment first
conda create -n ft python=3.10 -y && conda activate ft
pip install "transformers>=4.43" "peft>=0.11" "trl>=0.9" \
"datasets>=2.19" "bitsandbytes>=0.43" "accelerate>=0.32"| Package | Role |
|---|---|
| transformers | Model and tokenizer loading, core training API |
| peft | LoRA configuration and model wrapping (LoraConfig / get_peft_model) |
| trl | High-level trainers such as SFTTrainer and DPOTrainer |
| datasets | Dataset loading and preprocessing |
| bitsandbytes | Low-level 4-bit / 8-bit quantization library (required for QLoRA) |
| accelerate | Infrastructure for distributed and mixed-precision training |
Linux plus NVIDIA drivers is the combination with the best official support. On Windows, bitsandbytes has offered official support since 0.43, but WSL2 is recommended. Mac (MPS) doesn't support 4-bit training; switch to 8-bit or a smaller model.
3. Data Preparation: 80% of Fine-Tuning Success
A saying in the industry: the bottleneck of LoRA fine-tuning is never compute — it's data. One high-quality instruction dataset beats three low-quality ones.
1. Instruction Dataset Format
Fine-tuning is fundamentally about teaching the model an input → output mapping. The most universal format is the Alpaca-style triple:
json
[
{
"instruction": "Explain what a vector database is in one sentence",
"input": "",
"output": "A vector database is built around vector indexes and specializes in storing and retrieving high-dimensional vectors, typically paired with an embedding model for semantic search and RAG scenarios."
},
{
"instruction": "Rewrite the product description below into ad copy of 20 characters or fewer for an e-commerce platform",
"input": "This is a smartwatch with wireless charging, IP68 water resistance, and a 7-day battery life",
"output": "Seven-day battery, wireless fast charging, ready for anywhere."
}
]instruction: the task description;input: the task input (may be empty);output: the desired model response.
Before training, each record must be assembled into full conversation text with role markers (ChatML format — Qwen's official training format):
python
from datasets import load_dataset
def format_example(example):
user = example["instruction"]
if example.get("input"):
user += "\n" + example["input"]
return {
"text": (
f"<|im_start|>user\n{user}<|im_end|>\n"
f"<|im_start|>assistant\n{example['output']}<|im_end|>\n"
)
}
ds = load_dataset("json", data_files="data.jsonl", split="train")
ds = ds.map(format_example, remove_columns=ds.column_names)
print(ds[0]["text"]) # Be sure to print one example before training and check the template character by character2. Starting from Public Datasets
Don't want to build data from scratch? Reuse community datasets — the full list is in Dataset and Tool Directory:
| Dataset | Size | Notes | Good for |
|---|---|---|---|
| Alpaca (stanford_alpaca) | 52k | The granddaddy of instruction tuning, general English | General capability boost |
| alpaca-cleaned / Chinese edition | 52k | Cleaned version, Chinese translation | Chinese general fine-tuning |
| ShareGPT | 90k+ | Real multi-turn conversations | Conversation style learning |
| OpenOrca / SlimOrca | 500k+ | Higher-quality English instructions | Large-scale general fine-tuning |
| Custom business data | Hundreds to tens of thousands | Your domain | Domain fine-tuning |
Key reminder: mix public data with business data, usually at a general : business ratio between 3:1 and 1:1. Fine-tuning on pure business data quickly overfits and erodes general capabilities (see the common pitfalls section below).
3. Six Rules of Data Cleaning
- Deduplicate: repeated samples make the model memorize high-frequency content instead of learning to generalize; deduplicate by hash or embedding similarity;
- Length filtering: drop samples that are too short (< 10 tokens) or too long (beyond 80% of
max_seq_length); - Format consistency: field semantics must be uniform across all samples (e.g. is
inputan empty string or null — pick one), otherwise the model learns behavior that only works sometimes; - Answer quality review: the most common defect in instruction data is wrong answers — manually review a sample of 100; if the error rate exceeds 5%, go back and fix the data;
- Privacy and sensitive data: training data ends up "living in" the model weights, so PII (phone numbers, ID numbers) must be masked or removed;
- Distribution coverage: check that all business scenarios are covered, especially edge cases (ambiguous inputs, extra-long inputs, special characters).
4. Step-by-Step Implementation: Five Steps from Loading to Deployment
The following end-to-end example ties the whole pipeline together: use Qwen2.5-7B-Instruct + LoRA on a consumer GPU (about 8 GB VRAM) to fine-tune an assistant that "knows vector databases".
Step 1: Load the Model and Tokenizer (4-bit QLoRA Configuration)
python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
model_id = "Qwen/Qwen2.5-7B-Instruct"
# 4-bit NF4 quantization config (the core of QLoRA)
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4", # 4-bit NormalFloat
bnb_4bit_use_double_quant=True, # Double quantization, saving another ~0.4 bits/param
bnb_4bit_compute_dtype=torch.bfloat16,
)
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
# Some Chinese models require the pad token to be set explicitly
if tokenizer.pad_token is None:
tokenizer.pad_token = tokenizer.eos_token
model = AutoModelForCausalLM.from_pretrained(
model_id,
quantization_config=bnb_config,
device_map="auto", # Distribute layers across VRAM/CPU automatically
trust_remote_code=True,
)
model.config.use_cache = False # Disable the KV cache during training to save VRAMKey points:
bnb_4bit_compute_dtype=torch.bfloat16keeps compute in bf16, saving VRAM while preserving precision — the QLoRA default recommendation;device_map="auto"places layers on the GPU, with any that don't fit falling back to CPU (very slow — shrink the model or sequence length if you can). For the full discussion of quantization and inference optimization, see Inference Optimization and Quantization.
Step 2: LoRA Configuration and PEFT Wrapping
LoRA's core idea: freeze the original weights, train only the injected low-rank matrices, then merge the deltas back in after training. For the theory, see Fine-Tuning and PEFT (LoRA).
python
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training
# QLoRA requirement: convert the quantized model weights to a trainable layout first
model = prepare_model_for_kbit_training(model)
lora_config = LoraConfig(
r=16, # Rank of the low-rank matrices: larger = more expressive, more VRAM
lora_alpha=32, # Scaling factor, usually 1-2x r
lora_dropout=0.05,
bias="none",
task_type="CAUSAL_LM",
target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj"], # Inject into all attention + FFN layers
)
model = get_peft_model(model, lora_config)
model.print_trainable_parameters()
# Output looks like: trainable params: 21,626,880 || all params: 7,683,094,528 || trainable%: 0.2815Parameter quick reference:
| Parameter | Meaning | Recommendation |
|---|---|---|
r | Low rank, the width of the LoRA-injected matrices | 8 (small data / saving VRAM) to 64 (large data) |
lora_alpha | Scaling factor; the effective scale is alpha / r | Usually 2 × r; no need to tune it precisely |
lora_dropout | Dropout on the injected layers | 0.05-0.1 |
target_modules | Which linear layers get LoRA injected | Default is q/k/v/o; adding the FFN layers works better |
Step 3: Training (trl SFTTrainer)
python
from trl import SFTTrainer, SFTConfig
training_args = SFTConfig(
output_dir="./qwen2.5-7b-lora",
per_device_train_batch_size=1,
gradient_accumulation_steps=16, # Effective batch size = 1 × 16 = 16
learning_rate=2e-4, # 1e-4 ~ 5e-4 is typical for LoRA
num_train_epochs=3,
lr_scheduler_type="cosine",
warmup_ratio=0.03,
optim="paged_adamw_8bit", # 8-bit optimizer, saves even more VRAM
logging_steps=10,
save_steps=200,
save_total_limit=3,
max_seq_length=1024, # Maximum training sequence length
bf16=True, # Use bf16 on Ampere or newer GPUs
report_to="none",
)
trainer = SFTTrainer(
model=model,
args=training_args,
train_dataset=ds,
processing_class=tokenizer, # Syntax for trl >= 0.12; older versions use tokenizer=tokenizer
)
trainer.train()
trainer.save_model("./qwen2.5-7b-lora/final") # Saves only the LoRA adapterHyperparameter quick reference (rules of thumb):
| Hyperparameter | Meaning | Typical range | Notes |
|---|---|---|---|
learning_rate | Learning rate | 1e-4 ~ 5e-4 (LoRA) | Full fine-tuning uses ~1e-5; LoRA can be more aggressive |
num_train_epochs | Number of epochs | 1-3 | Small data (< 5k): 3 epochs; large data: 1-2 |
per_device_train_batch_size | Per-GPU batch size | 1-4 | Limited by VRAM; usually 1 |
gradient_accumulation_steps | Gradient accumulation steps | 8-32 | Effective batch = batch_size × accumulation steps |
max_seq_length | Sequence length | 512-2048 | Drives activation VRAM; just make it long enough |
warmup_ratio | Warmup ratio | 0.03-0.1 | Stabilizes the early phase of training |
Three Training Disciplines
- Watch the training loss curve: ideally it descends smoothly and stabilizes within 1-3 epochs; if loss isn't dropping, check the learning rate first (too high = oscillation, too low = no movement).
- Don't chase a train loss of zero: zero means overfitting — the model has memorized the training data.
- Save checkpoints regularly: every
save_stepssave means a training crash won't force you to start over.
Step 4: Merge and Export
The output of LoRA training is a set of adapter files totaling a few dozen MB. At inference time you either merge them back into the base model or load the adapter directly through peft. To export as a standard model:
python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base_model_id = "Qwen/Qwen2.5-7B-Instruct"
adapter_path = "./qwen2.5-7b-lora/final"
export_path = "./qwen2.5-7b-merged"
# Reload the base in BF16 (no quantization), then attach the LoRA adapter
base = AutoModelForCausalLM.from_pretrained(
base_model_id, torch_dtype=torch.bfloat16, device_map="auto",
)
model = PeftModel.from_pretrained(base, adapter_path)
model = model.merge_and_unload() # Merge the LoRA deltas into the weights
model.save_pretrained(export_path)
AutoTokenizer.from_pretrained(base_model_id).save_pretrained(export_path)
print("merged model saved to", export_path)The merged model can be loaded directly with AutoModelForCausalLM.from_pretrained(export_path). For production deployment and inference optimization (quantization, vLLM, concurrency control), see Deploying LLMs and Inference Optimization in Practice.
Step 5: Inference Validation — Before vs. After Fine-Tuning
The first thing to do once training finishes: take samples the model never saw during training and compare its answers before and after fine-tuning.
python
def chat(model, tokenizer, prompt, max_new_tokens=128):
messages = [{"role": "user", "content": prompt}]
text = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=max_new_tokens, do_sample=False)
return tokenizer.decode(out[0][inputs["input_ids"].shape[1]:],
skip_special_tokens=True)
# The fine-tuned model
finetuned = AutoModelForCausalLM.from_pretrained(
export_path, torch_dtype=torch.bfloat16, device_map="auto"
)
# The base model before fine-tuning (for comparison)
base = AutoModelForCausalLM.from_pretrained(
"Qwen/Qwen2.5-7B-Instruct", torch_dtype=torch.bfloat16, device_map="auto"
)
test_prompt = "Explain the HNSW index to a product manager in one sentence"
print("Before fine-tuning:", chat(base, tokenizer, test_prompt))
print("After fine-tuning:", chat(finetuned, tokenizer, test_prompt))A typical side-by-side output:
Before fine-tuning: The HNSW index is a graph-based approximate nearest neighbor search algorithm that builds a layered graph structure in high-dimensional space and finds nearest neighbors through greedy search.
After fine-tuning: HNSW is a "skip list + graph" hierarchical small-world index built specifically for nearest neighbor retrieval over high-dimensional vectors. As a product manager, remember three things: fast queries (millisecond-level), controllable memory, and support for incremental inserts.Notice the post-fine-tuning changes: shorter sentences, analogies that match your internal messaging, a consistent wrap-up phrasing — exactly what "the model has learned your data distribution" looks like. If the two answers barely differ, the problem is almost certainly in the data, not the training.
5. Evaluate and Iterate: Never Accept Fine-Tuning on a "Feeling"
1. Build a Validation Set
Hold out 10%-20% of samples that never participate in training as a validation set (or build 30-100 dedicated "exam questions"). Every few epochs during training, run inference on the validation set and score it with the metrics from Building an LLM Evaluation. If training metrics rise while validation metrics stall or even fall, that's overfitting — at that point, stop training, cut the epochs, and lower the learning rate.
2. Manual Spot-Checks Are the Bare Minimum
Automatic metrics (loss, BLEU) perceive little about open-ended responses. Manually review at least 20-50 real outputs every iteration, checking three problem types:
- Hallucination: did the model fabricate "facts" that aren't in the training data;
- Style drift: has the internal tone gone off-track (e.g. the training data is mostly English while production requires Chinese output);
- Edge behavior: does it lose control on inputs outside the training distribution (profanity, extremely long questions, empty inputs).
3. Plugging Into Your Evaluation Framework
Fine-tuning isn't "train and done" — it's one link in the evaluation loop. Hook the fine-tuned model into your evaluation framework, A/B it against the base model or the previous version on a golden dataset, and decide whether to ship based on data, not impressions. For choosing evaluation metrics and using LLM-as-judge, see LLM Evaluation and Benchmarks.
6. Common Pitfalls: Troubleshooting from One Table
| Pitfall | Symptom | Cause | Diagnosis / Fix |
|---|---|---|---|
| Wrong data format | Training runs fine but generation is garbled, with repeated "assistant" markers | ChatML template doesn't match the tokenizer, or `< | im_end |
| Overfitting on small data | Low training loss but poor validation results; answers sound "recited" | Little data (< 1k), too many epochs, r too large | Cut epochs, mix in general data, augment the data |
| Catastrophic forgetting | General capabilities degrade (math and code get worse) | Fine-tuning on pure business data, prior knowledge overwritten | Mix in general data at 3:1; do multi-task mixed training |
| Chinese tokenization issues | Chinese replies break sentences oddly or drop characters | Tokenizer lacks a pad token, sequences truncated | Set pad_token = eos_token explicitly; check max_seq_length |
| VRAM OOM | Crash mid-training | Sequences too long / batch too large | Lower max_seq_length, batch=1 + gradient accumulation, drop r to 8 |
| Quantized loading fails | bitsandbytes throws errors | Driver / CUDA version mismatch | pip install -U bitsandbytes, check nvidia-smi |
| Training loss won't drop | Oscillating from the very start | Learning rate too high | Retry with lr down at 5e-5 |
The one-line rule: when anything goes wrong with fine-tuning, suspect the data first, the training config second, and the code last. For a fuller list of training and deployment traps, see Common Pitfalls and Anti-Patterns.
7. Going Further: DPO and Multi-Turn Conversations
1. DPO: Using Preference Data to Make the Model More Likable
SFT can only teach the model to "answer correctly"; it can't learn "which answer is better". DPO (Direct Preference Optimization) needs no reward model — it optimizes the policy directly on paired "good answer vs. bad answer" data, making it a frontline alignment method. For the theory, see Alignment: RLHF and DPO.
python
from trl import DPOConfig, DPOTrainer
dpo_config = DPOConfig(
output_dir="./qwen2.5-7b-dpo",
per_device_train_batch_size=1,
gradient_accumulation_steps=8,
learning_rate=1e-6, # DPO typically uses a much smaller learning rate
num_train_epochs=1,
bf16=True,
beta=0.1, # DPO temperature coefficient; controls how strongly the model follows preferences
)
dpo_trainer = DPOTrainer(
model=model, # Can reuse the LoRA model from SFT
args=dpo_config,
train_dataset=dpo_ds, # Three fields per sample: {prompt, chosen, rejected}
processing_class=tokenizer,
)
dpo_trainer.train()Each sample in dpo_ds contains three fields: prompt (the user question), chosen (the better answer), and rejected (the worse answer). Preference data can be accumulated from human labeling, user feedback (thumbs up/down), or side-by-side outputs from two models — for the latter's collection pipeline, see the production evaluation section of Building an LLM Evaluation.
SFT → DPO is the standard one-two punch: first use SFT to pull the model into your domain, then use DPO to fine-tune style and preferences. If budget or time is limited, do SFT only first.
2. Multi-Turn Conversation Fine-Tuning
A model trained on single-turn instruction data will "forget" history. Multi-turn conversation fine-tuning uses the ShareGPT format, where each sample is a complete conversation history:
json
{
"conversations": [
{"from": "human", "value": "Recommend earbuds suitable for running a marathon"},
{"from": "gpt", "value": "Considering a secure fit and battery life, I'd recommend the XX bone conduction earbuds, with 10 hours of battery."},
{"from": "human", "value": "What about a budget under 500 yuan?"},
{"from": "gpt", "value": "Then go with the YY — also bone conduction, 8 hours of battery, 399 yuan."}
]
}Key points for constructing multi-turn samples: keep the full history so the model learns to "resolve references in context"; also keep some samples "truncated mid-conversation" to simulate real conversations whose opening got cut off. For the underlying mechanics of chat formats and templates, see Large Language Models (LLM).
Further Reading
- Fine-Tuning and PEFT (LoRA) — the theory companion to this article: how low-rank approximation works, and how different PEFT methods compare
- Prompt Engineering — what to try before fine-tuning; the two are often used together
- Retrieval-Augmented Generation (RAG) — the other half of the "fine-tune or RAG" decision
- Inference Optimization and Quantization — the quantization principles behind QLoRA and inference optimization at deployment
- Building an LLM Evaluation — the acceptance-testing toolkit for after fine-tuning, referenced in Section 5
- LLM Evaluation and Benchmarks — the complete framework of evaluation metrics and benchmarks
- Alignment: RLHF and DPO — the full theory behind the DPO covered in Section 7
- Deploying LLMs and Inference Optimization in Practice — how to go live after fine-tuning
- Common Pitfalls and Anti-Patterns — the complete catalog of training and deployment traps
- Model and Leaderboard Quick Reference — the basis for choosing a base model
- Dataset and Tool Directory — where to find fine-tuning datasets
- Large Language Models (LLM) — the model literacy to build up before fine-tuning
References
- Hu et al. LoRA: Low-Rank Adaptation of Large Language Models (ICLR 2022) — the original LoRA paper; the source of all LoRA configurations in this article
- Dettmers et al. QLoRA: Efficient Finetuning of Quantized LLMs (NeurIPS 2023) — the original QLoRA paper: NF4 quantization, double quantization, paged optimizer
- Hugging Face PEFT documentation — official API reference for LoraConfig, get_peft_model, and merge_and_unload
- Hugging Face TRL documentation — official docs for SFTTrainer / DPOTrainer
- bitsandbytes official repository — the low-level 4-bit quantization library
- Stanford Alpaca project — the source of the 52k instruction dataset
- ShareGPT dataset (Hugging Face) — commonly used data for multi-turn conversation fine-tuning
- Rafailov et al. Direct Preference Optimization (NeurIPS 2023) — the original DPO paper
- OpenAI Fine-tuning guide — the managed-API alternative for fine-tuning (when you don't want to manage GPUs yourself)
- Qwen official fine-tuning docs — the official training guide for Qwen, the example model in this article
Suggested order: get the minimal example in Steps 1-5 running first (even with only 100 samples) → swap in real business data → iterate 2-3 rounds of manual spot-checks on the validation set → sign off with the evaluation framework → only then consider DPO fine-tuning.