Skip to content

Fine-Tune Your Own LLM

At a glance Fine-tune your own model with Qwen/Llama and LoRA on a consumer GPU: decide whether fine-tuning is warranted, prepare VRAM and data, then follow complete runnable code from QLoRA training and merge/export through inference validation and evaluation to launch.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

Fine-Tune Your Own LLM ​

Fine-tuning means taking a pretrained large model and continuing to train it on your own data, so the model "learns" your business language, output formats, and behavioral preferences. Thanks to parameter-efficient fine-tuning techniques like LoRA (Low-Rank Adaptation), this is no longer the exclusive domain of a few big labs: with one consumer GPU with 8 GB of VRAM, a few thousand to tens of thousands of high-quality instruction samples, and an afternoon, you can turn Qwen or Llama into a little assistant that understands your business.

This article doesn't rehash the theory of fine-tuning. Instead it answers three practical questions: Should your use case be fine-tuned at all? How do you prepare VRAM and the software environment? What does the complete pipeline from data to deployment look like? What follows is an operations manual you can copy verbatim — data preparation, LoRA training, merge and export, and inference validation, with runnable code at every step. For the theory behind fine-tuning (why LoRA works, what low-rank approximation actually is), see Fine-Tuning and PEFT (LoRA); this article focuses on how to get it running.

Prerequisites: you should already understand the basics of large language models (LLMs) and have hands-on experience running a prompt engineering or RAG application. Every stage in this article has a corresponding in-depth article on this site; links are provided throughout.

1. Decide First: Is Your Use Case Worth Fine-Tuning ​

1. The One-Sentence Test ​

Prompting first, RAG second, and only then consider fine-tuning. Fine-tuning isn't "more advanced" — it's "more expensive": it demands data, compute, and iteration time, and every change requires retraining. The vast majority of business problems are already solved at the prompting and RAG stages.

DimensionPrompt EngineeringRAGFine-Tuning (LoRA)
Cost of changeA few minutes; just edit the promptUpdate the retrieval store and indexHours to days; retrain the model
Data requiredA few examples (few-shot)A document corpusThousands of high-quality labeled samples
Knowledge freshnessGoodBest (changes take effect immediately)Poor (frozen at training time)
Deep domain capabilityWeakModerateStrong
Private knowledgeRelies on prompt injectionBest fitCan be baked in, but goes stale easily
Output format/style controlModerateWeakStrong
Deployment costLowestRequires a retrieval stackLarger model and VRAM overhead
Typical use casesGeneral assistants, one-off needsSupport knowledge bases, enterprise Q&ADomain language, fixed-format output

Four signals that fine-tuning is worth it:

  • The model clearly misunderstands domain terminology, jargon, and proper nouns, and documents retrieved via RAG can't "teach" it (think medicine, law, or industry-specific slang);
  • You need a stable, fixed output format (a specific JSON schema, a specific reply style) and prompts keep drifting off;
  • You're deploying offline / locally / inside an intranet and can't carry a huge context with every request;
  • Inference cost matters: rather than stuffing 20k tokens of context into every request, bake the high-frequency knowledge directly into the weights.

When to hit the brakes:

  • Knowledge changes frequently → use RAG instead of fine-tuning;
  • Data is scarce or of dubious quality (you can't produce 500+ high-quality samples) → polish your prompts first;
  • You have no validation set or evaluation method → read Building an LLM Evaluation first; otherwise you won't be able to tell good results from bad after fine-tuning.

2. How to Choose a Base Model ​

Pick the wrong base model and everything downstream is wasted effort. Evaluate three dimensions: Chinese vs. English, model size vs. VRAM, instruct model vs. base model. For parameter counts and leaderboard data on each model, see Model and Leaderboard Quick Reference.

ModelSizeCharacteristicsBest for
Qwen2.5-7B-Instruct7BStrong Chinese, good instruction following, mature ecosystemChinese-language business, general fine-tuning on-ramp
Llama-3.1-8B-Instruct8BStrong English, richest community ecosystemEnglish tasks, tool calling
Mistral-7B-Instruct7BEfficient inference, VRAM-friendlyEnglish, tight resources
Qwen2.5-14B / Llama-3.1-13B14B / 13BStronger but hungrier for VRAMLarge datasets, quality first
Qwen2.5-3B / Llama-3.2-3B3BRuns on minimal VRAMTeaching, edge devices

Three rules of thumb:

  • For Chinese-language business, prefer Qwen; for general English work, prefer Llama. This is experience the community has validated repeatedly — see Model and Leaderboard Quick Reference for detailed comparison data;
  • Start at 7B / 8B: on consumer GPUs, 7B is the value-for-money sweet spot; at 13B and above, VRAM and training time climb steeply;
  • Choose an "instruct" model, not a base model: the -Instruct / -Chat variants have already been aligned (RLHF/DPO) and behave better after fine-tuning; fine-tuning a base model directly means you have to build its conversational ability yourself — generally not recommended.

2. Environment Setup: The VRAM Budget and Dependencies ​

1. VRAM Requirements Table (Rules of Thumb) ​

Fine-tuning VRAM consists of four parts: model weights + gradients + optimizer states + activations. Full fine-tuning (FFT) has to maintain all of those complete states at once, so its VRAM demand is typically 5-10x that of QLoRA. The table below gives rule-of-thumb figures (in GB); actual usage varies with max_seq_len and batch_size:

Model sizeFull fine-tuning (BF16)QLoRA (4-bit + LoRA)Suggested GPU
2B / 3B~18 GB~3-4 GB6-8 GB is enough
7B / 8B~60-70 GB (needs A100/H100 class)~7-9 GBRTX 4070 / 3090
13B / 14B~120-140 GB (multi-GPU)~12-15 GBRTX 4090
70BMulti-node, multi-GPU~35-45 GBA100 80G / multi-GPU

The conclusion is clear: on consumer GPUs, QLoRA is the only realistic option. QLoRA combines 4-bit quantization (NF4) with LoRA — the base weights are compressed to a quarter of their size and stay frozen throughout, while trainable parameters account for just 0.1%-1% of all parameters. For how quantization works, see Inference Optimization and Quantization.

Not enough VRAM?

Turn three dials in order: drop max_seq_length from 2048 to 1024 → set per_device_train_batch_size to 1 and recover the effective batch size with gradient_accumulation_steps → lower LoRA's r from 16 to 8. Of these, "batch=1 + gradient accumulation" is nearly lossless in quality and should be your first lever.

2. Installing Dependencies ​

bash
# Python 3.10+ recommended; create an isolated environment first
conda create -n ft python=3.10 -y && conda activate ft

pip install "transformers>=4.43" "peft>=0.11" "trl>=0.9" \
            "datasets>=2.19" "bitsandbytes>=0.43" "accelerate>=0.32"
PackageRole
transformersModel and tokenizer loading, core training API
peftLoRA configuration and model wrapping (LoraConfig / get_peft_model)
trlHigh-level trainers such as SFTTrainer and DPOTrainer
datasetsDataset loading and preprocessing
bitsandbytesLow-level 4-bit / 8-bit quantization library (required for QLoRA)
accelerateInfrastructure for distributed and mixed-precision training

Linux plus NVIDIA drivers is the combination with the best official support. On Windows, bitsandbytes has offered official support since 0.43, but WSL2 is recommended. Mac (MPS) doesn't support 4-bit training; switch to 8-bit or a smaller model.

3. Data Preparation: 80% of Fine-Tuning Success ​

A saying in the industry: the bottleneck of LoRA fine-tuning is never compute — it's data. One high-quality instruction dataset beats three low-quality ones.

1. Instruction Dataset Format ​

Fine-tuning is fundamentally about teaching the model an input → output mapping. The most universal format is the Alpaca-style triple:

json
[
  {
    "instruction": "Explain what a vector database is in one sentence",
    "input": "",
    "output": "A vector database is built around vector indexes and specializes in storing and retrieving high-dimensional vectors, typically paired with an embedding model for semantic search and RAG scenarios."
  },
  {
    "instruction": "Rewrite the product description below into ad copy of 20 characters or fewer for an e-commerce platform",
    "input": "This is a smartwatch with wireless charging, IP68 water resistance, and a 7-day battery life",
    "output": "Seven-day battery, wireless fast charging, ready for anywhere."
  }
]
  • instruction: the task description;
  • input: the task input (may be empty);
  • output: the desired model response.

Before training, each record must be assembled into full conversation text with role markers (ChatML format — Qwen's official training format):

python
from datasets import load_dataset

def format_example(example):
    user = example["instruction"]
    if example.get("input"):
        user += "\n" + example["input"]
    return {
        "text": (
            f"<|im_start|>user\n{user}<|im_end|>\n"
            f"<|im_start|>assistant\n{example['output']}<|im_end|>\n"
        )
    }

ds = load_dataset("json", data_files="data.jsonl", split="train")
ds = ds.map(format_example, remove_columns=ds.column_names)
print(ds[0]["text"])   # Be sure to print one example before training and check the template character by character

2. Starting from Public Datasets ​

Don't want to build data from scratch? Reuse community datasets — the full list is in Dataset and Tool Directory:

DatasetSizeNotesGood for
Alpaca (stanford_alpaca)52kThe granddaddy of instruction tuning, general EnglishGeneral capability boost
alpaca-cleaned / Chinese edition52kCleaned version, Chinese translationChinese general fine-tuning
ShareGPT90k+Real multi-turn conversationsConversation style learning
OpenOrca / SlimOrca500k+Higher-quality English instructionsLarge-scale general fine-tuning
Custom business dataHundreds to tens of thousandsYour domainDomain fine-tuning

Key reminder: mix public data with business data, usually at a general : business ratio between 3:1 and 1:1. Fine-tuning on pure business data quickly overfits and erodes general capabilities (see the common pitfalls section below).

3. Six Rules of Data Cleaning ​

  1. Deduplicate: repeated samples make the model memorize high-frequency content instead of learning to generalize; deduplicate by hash or embedding similarity;
  2. Length filtering: drop samples that are too short (< 10 tokens) or too long (beyond 80% of max_seq_length);
  3. Format consistency: field semantics must be uniform across all samples (e.g. is input an empty string or null — pick one), otherwise the model learns behavior that only works sometimes;
  4. Answer quality review: the most common defect in instruction data is wrong answers — manually review a sample of 100; if the error rate exceeds 5%, go back and fix the data;
  5. Privacy and sensitive data: training data ends up "living in" the model weights, so PII (phone numbers, ID numbers) must be masked or removed;
  6. Distribution coverage: check that all business scenarios are covered, especially edge cases (ambiguous inputs, extra-long inputs, special characters).

4. Step-by-Step Implementation: Five Steps from Loading to Deployment ​

The following end-to-end example ties the whole pipeline together: use Qwen2.5-7B-Instruct + LoRA on a consumer GPU (about 8 GB VRAM) to fine-tune an assistant that "knows vector databases".

Step 1: Load the Model and Tokenizer (4-bit QLoRA Configuration) ​

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig

model_id = "Qwen/Qwen2.5-7B-Instruct"

# 4-bit NF4 quantization config (the core of QLoRA)
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",        # 4-bit NormalFloat
    bnb_4bit_use_double_quant=True,   # Double quantization, saving another ~0.4 bits/param
    bnb_4bit_compute_dtype=torch.bfloat16,
)

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
# Some Chinese models require the pad token to be set explicitly
if tokenizer.pad_token is None:
    tokenizer.pad_token = tokenizer.eos_token

model = AutoModelForCausalLM.from_pretrained(
    model_id,
    quantization_config=bnb_config,
    device_map="auto",                # Distribute layers across VRAM/CPU automatically
    trust_remote_code=True,
)
model.config.use_cache = False        # Disable the KV cache during training to save VRAM

Key points:

  • bnb_4bit_compute_dtype=torch.bfloat16 keeps compute in bf16, saving VRAM while preserving precision — the QLoRA default recommendation;
  • device_map="auto" places layers on the GPU, with any that don't fit falling back to CPU (very slow — shrink the model or sequence length if you can). For the full discussion of quantization and inference optimization, see Inference Optimization and Quantization.

Step 2: LoRA Configuration and PEFT Wrapping ​

LoRA's core idea: freeze the original weights, train only the injected low-rank matrices, then merge the deltas back in after training. For the theory, see Fine-Tuning and PEFT (LoRA).

python
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training

# QLoRA requirement: convert the quantized model weights to a trainable layout first
model = prepare_model_for_kbit_training(model)

lora_config = LoraConfig(
    r=16,                # Rank of the low-rank matrices: larger = more expressive, more VRAM
    lora_alpha=32,       # Scaling factor, usually 1-2x r
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM",
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
                    "gate_proj", "up_proj", "down_proj"],  # Inject into all attention + FFN layers
)

model = get_peft_model(model, lora_config)
model.print_trainable_parameters()
# Output looks like: trainable params: 21,626,880 || all params: 7,683,094,528 || trainable%: 0.2815

Parameter quick reference:

ParameterMeaningRecommendation
rLow rank, the width of the LoRA-injected matrices8 (small data / saving VRAM) to 64 (large data)
lora_alphaScaling factor; the effective scale is alpha / rUsually 2 × r; no need to tune it precisely
lora_dropoutDropout on the injected layers0.05-0.1
target_modulesWhich linear layers get LoRA injectedDefault is q/k/v/o; adding the FFN layers works better

Step 3: Training (trl SFTTrainer) ​

python
from trl import SFTTrainer, SFTConfig

training_args = SFTConfig(
    output_dir="./qwen2.5-7b-lora",
    per_device_train_batch_size=1,
    gradient_accumulation_steps=16,     # Effective batch size = 1 × 16 = 16
    learning_rate=2e-4,                 # 1e-4 ~ 5e-4 is typical for LoRA
    num_train_epochs=3,
    lr_scheduler_type="cosine",
    warmup_ratio=0.03,
    optim="paged_adamw_8bit",           # 8-bit optimizer, saves even more VRAM
    logging_steps=10,
    save_steps=200,
    save_total_limit=3,
    max_seq_length=1024,                # Maximum training sequence length
    bf16=True,                          # Use bf16 on Ampere or newer GPUs
    report_to="none",
)

trainer = SFTTrainer(
    model=model,
    args=training_args,
    train_dataset=ds,
    processing_class=tokenizer,         # Syntax for trl >= 0.12; older versions use tokenizer=tokenizer
)

trainer.train()
trainer.save_model("./qwen2.5-7b-lora/final")   # Saves only the LoRA adapter

Hyperparameter quick reference (rules of thumb):

HyperparameterMeaningTypical rangeNotes
learning_rateLearning rate1e-4 ~ 5e-4 (LoRA)Full fine-tuning uses ~1e-5; LoRA can be more aggressive
num_train_epochsNumber of epochs1-3Small data (< 5k): 3 epochs; large data: 1-2
per_device_train_batch_sizePer-GPU batch size1-4Limited by VRAM; usually 1
gradient_accumulation_stepsGradient accumulation steps8-32Effective batch = batch_size × accumulation steps
max_seq_lengthSequence length512-2048Drives activation VRAM; just make it long enough
warmup_ratioWarmup ratio0.03-0.1Stabilizes the early phase of training

Three Training Disciplines

  • Watch the training loss curve: ideally it descends smoothly and stabilizes within 1-3 epochs; if loss isn't dropping, check the learning rate first (too high = oscillation, too low = no movement).
  • Don't chase a train loss of zero: zero means overfitting — the model has memorized the training data.
  • Save checkpoints regularly: every save_steps save means a training crash won't force you to start over.

Step 4: Merge and Export ​

The output of LoRA training is a set of adapter files totaling a few dozen MB. At inference time you either merge them back into the base model or load the adapter directly through peft. To export as a standard model:

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base_model_id = "Qwen/Qwen2.5-7B-Instruct"
adapter_path = "./qwen2.5-7b-lora/final"
export_path  = "./qwen2.5-7b-merged"

# Reload the base in BF16 (no quantization), then attach the LoRA adapter
base = AutoModelForCausalLM.from_pretrained(
    base_model_id, torch_dtype=torch.bfloat16, device_map="auto",
)
model = PeftModel.from_pretrained(base, adapter_path)
model = model.merge_and_unload()        # Merge the LoRA deltas into the weights
model.save_pretrained(export_path)
AutoTokenizer.from_pretrained(base_model_id).save_pretrained(export_path)
print("merged model saved to", export_path)

The merged model can be loaded directly with AutoModelForCausalLM.from_pretrained(export_path). For production deployment and inference optimization (quantization, vLLM, concurrency control), see Deploying LLMs and Inference Optimization in Practice.

Step 5: Inference Validation — Before vs. After Fine-Tuning ​

The first thing to do once training finishes: take samples the model never saw during training and compare its answers before and after fine-tuning.

python
def chat(model, tokenizer, prompt, max_new_tokens=128):
    messages = [{"role": "user", "content": prompt}]
    text = tokenizer.apply_chat_template(
        messages, tokenize=False, add_generation_prompt=True
    )
    inputs = tokenizer(text, return_tensors="pt").to(model.device)
    out = model.generate(**inputs, max_new_tokens=max_new_tokens, do_sample=False)
    return tokenizer.decode(out[0][inputs["input_ids"].shape[1]:],
                            skip_special_tokens=True)

# The fine-tuned model
finetuned = AutoModelForCausalLM.from_pretrained(
    export_path, torch_dtype=torch.bfloat16, device_map="auto"
)
# The base model before fine-tuning (for comparison)
base = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen2.5-7B-Instruct", torch_dtype=torch.bfloat16, device_map="auto"
)

test_prompt = "Explain the HNSW index to a product manager in one sentence"
print("Before fine-tuning:", chat(base, tokenizer, test_prompt))
print("After fine-tuning:", chat(finetuned, tokenizer, test_prompt))

A typical side-by-side output:

Before fine-tuning: The HNSW index is a graph-based approximate nearest neighbor search algorithm that builds a layered graph structure in high-dimensional space and finds nearest neighbors through greedy search.
After fine-tuning: HNSW is a "skip list + graph" hierarchical small-world index built specifically for nearest neighbor retrieval over high-dimensional vectors. As a product manager, remember three things: fast queries (millisecond-level), controllable memory, and support for incremental inserts.

Notice the post-fine-tuning changes: shorter sentences, analogies that match your internal messaging, a consistent wrap-up phrasing — exactly what "the model has learned your data distribution" looks like. If the two answers barely differ, the problem is almost certainly in the data, not the training.

5. Evaluate and Iterate: Never Accept Fine-Tuning on a "Feeling" ​

1. Build a Validation Set ​

Hold out 10%-20% of samples that never participate in training as a validation set (or build 30-100 dedicated "exam questions"). Every few epochs during training, run inference on the validation set and score it with the metrics from Building an LLM Evaluation. If training metrics rise while validation metrics stall or even fall, that's overfitting — at that point, stop training, cut the epochs, and lower the learning rate.

2. Manual Spot-Checks Are the Bare Minimum ​

Automatic metrics (loss, BLEU) perceive little about open-ended responses. Manually review at least 20-50 real outputs every iteration, checking three problem types:

  • Hallucination: did the model fabricate "facts" that aren't in the training data;
  • Style drift: has the internal tone gone off-track (e.g. the training data is mostly English while production requires Chinese output);
  • Edge behavior: does it lose control on inputs outside the training distribution (profanity, extremely long questions, empty inputs).

3. Plugging Into Your Evaluation Framework ​

Fine-tuning isn't "train and done" — it's one link in the evaluation loop. Hook the fine-tuned model into your evaluation framework, A/B it against the base model or the previous version on a golden dataset, and decide whether to ship based on data, not impressions. For choosing evaluation metrics and using LLM-as-judge, see LLM Evaluation and Benchmarks.

6. Common Pitfalls: Troubleshooting from One Table ​

PitfallSymptomCauseDiagnosis / Fix
Wrong data formatTraining runs fine but generation is garbled, with repeated "assistant" markersChatML template doesn't match the tokenizer, or `<im_end
Overfitting on small dataLow training loss but poor validation results; answers sound "recited"Little data (< 1k), too many epochs, r too largeCut epochs, mix in general data, augment the data
Catastrophic forgettingGeneral capabilities degrade (math and code get worse)Fine-tuning on pure business data, prior knowledge overwrittenMix in general data at 3:1; do multi-task mixed training
Chinese tokenization issuesChinese replies break sentences oddly or drop charactersTokenizer lacks a pad token, sequences truncatedSet pad_token = eos_token explicitly; check max_seq_length
VRAM OOMCrash mid-trainingSequences too long / batch too largeLower max_seq_length, batch=1 + gradient accumulation, drop r to 8
Quantized loading failsbitsandbytes throws errorsDriver / CUDA version mismatchpip install -U bitsandbytes, check nvidia-smi
Training loss won't dropOscillating from the very startLearning rate too highRetry with lr down at 5e-5

The one-line rule: when anything goes wrong with fine-tuning, suspect the data first, the training config second, and the code last. For a fuller list of training and deployment traps, see Common Pitfalls and Anti-Patterns.

7. Going Further: DPO and Multi-Turn Conversations ​

1. DPO: Using Preference Data to Make the Model More Likable ​

SFT can only teach the model to "answer correctly"; it can't learn "which answer is better". DPO (Direct Preference Optimization) needs no reward model — it optimizes the policy directly on paired "good answer vs. bad answer" data, making it a frontline alignment method. For the theory, see Alignment: RLHF and DPO.

python
from trl import DPOConfig, DPOTrainer

dpo_config = DPOConfig(
    output_dir="./qwen2.5-7b-dpo",
    per_device_train_batch_size=1,
    gradient_accumulation_steps=8,
    learning_rate=1e-6,          # DPO typically uses a much smaller learning rate
    num_train_epochs=1,
    bf16=True,
    beta=0.1,                    # DPO temperature coefficient; controls how strongly the model follows preferences
)

dpo_trainer = DPOTrainer(
    model=model,                 # Can reuse the LoRA model from SFT
    args=dpo_config,
    train_dataset=dpo_ds,        # Three fields per sample: {prompt, chosen, rejected}
    processing_class=tokenizer,
)
dpo_trainer.train()

Each sample in dpo_ds contains three fields: prompt (the user question), chosen (the better answer), and rejected (the worse answer). Preference data can be accumulated from human labeling, user feedback (thumbs up/down), or side-by-side outputs from two models — for the latter's collection pipeline, see the production evaluation section of Building an LLM Evaluation.

SFT → DPO is the standard one-two punch: first use SFT to pull the model into your domain, then use DPO to fine-tune style and preferences. If budget or time is limited, do SFT only first.

2. Multi-Turn Conversation Fine-Tuning ​

A model trained on single-turn instruction data will "forget" history. Multi-turn conversation fine-tuning uses the ShareGPT format, where each sample is a complete conversation history:

json
{
  "conversations": [
    {"from": "human", "value": "Recommend earbuds suitable for running a marathon"},
    {"from": "gpt", "value": "Considering a secure fit and battery life, I'd recommend the XX bone conduction earbuds, with 10 hours of battery."},
    {"from": "human", "value": "What about a budget under 500 yuan?"},
    {"from": "gpt", "value": "Then go with the YY — also bone conduction, 8 hours of battery, 399 yuan."}
  ]
}

Key points for constructing multi-turn samples: keep the full history so the model learns to "resolve references in context"; also keep some samples "truncated mid-conversation" to simulate real conversations whose opening got cut off. For the underlying mechanics of chat formats and templates, see Large Language Models (LLM).

Further Reading ​

References ​

Suggested order: get the minimal example in Steps 1-5 running first (even with only 100 samples) → swap in real business data → iterate 2-3 rounds of manual spot-checks on the validation set → sign off with the evaluation framework → only then consider DPO fine-tuning.