Theme
Frameworks and Tooling: A Selection Guide
The essence of tool selection is matching complexity to your scenario: frameworks solve 80% of common problems at the cost of 20% flexibility. Picking the wrong framework doesn't just mean "installing one extra package" — it locks the entire team into the framework's design.
This article breaks the LLM toolchain into six layers, each with a scenario → recommendation decision table. After reading, you should have a "selection map that gives you confidence" rather than "pick whatever's trending." The full open-source ecosystem landscape is at Llama and the Open-Source Ecosystem.
I. Layered Tool Map
┌─────────────────────────────────────────────┐
│ Application layer Dify / custom apps │
├─────────────────────────────────────────────┤
│ Orchestration layer LangChain / LlamaIndex │
├─────────────────────────────────────────────┤
│ Evaluation layer lm-eval-harness / OpenCompass / RAGAS │
├─────────────────────────────────────────────┤
│ Inference layer vLLM / SGLang / TensorRT-LLM / llama.cpp │
├─────────────────────────────────────────────┤
│ Training layer transformers / PEFT / TRL / DeepSpeed │
│ Megatron-LM / LLaMA-Factory / Axolotl │
├─────────────────────────────────────────────┤
│ Model & data layer Transformers / Datasets / Tokenizers │
└─────────────────────────────────────────────┘General selection principles (applied in every layer below):
- Default to the thinnest: If a standard library or API can handle it, don't introduce a framework.
- Get it working before abstracting: Implement the first version the most direct way; introduce abstractions only when repeated needs emerge.
- Maintain replaceability: Wrap core modules (e.g., vector DBs, inference engines) behind interfaces for easy swapping.
II. Model and Data Layer: Nearly Mandatory
| Tool | Purpose | When to Pick |
|---|---|---|
Hugging Face transformers | Standard library for model loading/inference/training | Almost all scenarios (de facto standard) |
datasets | Dataset loading and preprocessing | Mandatory for training/fine-tuning |
tokenizers | High-performance tokenization | When building custom tokenizers |
transformers's single-line from_pretrained loads models with unified AutoModel/AutoTokenizer interfaces, turning "swap models" from code changes into parameter changes — this is the foundation of the entire ecosystem. Without special reason, skip selection at the model layer and just use transformers.
III. Training Layer: Pick by Scale and Method
| Tool | Purpose | Applicable Scenarios | Not Applicable |
|---|---|---|---|
transformers Trainer | Single-GPU / small-scale training | Small model fine-tuning, teaching | Large-scale distributed training |
PEFT | LoRA/QLoRA and other parameter-efficient fine-tuning | Most fine-tuning tasks | Full-parameter fine-tuning |
TRL | SFT/DPO/PPO trainer | Instruction fine-tuning and alignment | Custom training loops |
LLaMA-Factory | One-click platform | Quick experiments, switching methods | Deep customization |
Axolotl | Configuration-driven training | Reproducing configs, batch experiments | Dynamic logic needs |
DeepSpeed | Distributed training engine (ZeRO) | Large-model multi-GPU training | Small-model single-GPU |
Megatron-LM | NVIDIA large-scale training (tensor parallelism, etc.) | 10B+ parameter training | Conventional fine-tuning |
Training selection decision table:
What are you doing?
├─ Fine-tuning under 7B, fits on one GPU → PEFT/TRL or LLaMA-Factory
├─ Fine-tuning 13B+, need quantization for memory → QLoRA (bitsandbytes) + PEFT
├─ Pretraining / full-parameter 70B+ multi-GPU → DeepSpeed (ZeRO-3) or Megatron-LM
└─ Alignment (RLHF/DPO) → TRL (DPOTrainer) or custom + trlx, etc.Three Disciplines of Training Engineering
Beyond frameworks, the most often overlooked aspect of training is engineering discipline:
| Discipline | Practice | Problem Solved |
|---|---|---|
| Reproducible experiments | Fix random seeds + record full hyperparameters (including data version, framework version) | Two experiments give different conclusions |
| Checkpoint management | Save regularly, keep best by eval score | Overfitting / rollback after incidents |
| Training logs | Record loss / gradient norm / learning rate per step | Locate divergence, judge convergence |
Training logs' value lies in offline postmortems: debugging a "fine-tuned but dumber" model relies on eval loss inflection points in the training logs.
A realistic note on training selection
Don't introduce DeepSpeed for problems solvable by PEFT/TRL. The complexity of distributed training (node communication, gradient synchronization, mixed precision) is an order-of-magnitude increase. Ask yourself first: is single-GPU + QLoRA enough? Full workflow at Fine-Tuning in Practice: Full LoRA Workflow.
IV. Inference Layer: The Engineering Battlefield of Throughput and Latency
Inference engines directly determine deployment costs and user experience. Metric definitions (TTFT/TPOT/throughput) are at Deployment and Serving.
| Engine | Characteristics | Applicable Scenarios |
|---|---|---|
| vLLM | PagedAttention, continuous batching, OpenAI-compatible API | Production default pick |
| SGLang | RadixAttention, structured output optimization | High concurrency, complex prompting, tool calling |
| TensorRT-LLM | NVIDIA deep optimization, TensorRT compilation | NVIDIA clusters, extreme throughput |
| llama.cpp (GGUF) | Runs on CPU / low memory, fully local | Local / edge, personal use, offline |
| Ollama | Wrapper around llama.cpp, one-click run | Local experience, teaching demos |
transformers native inference | Simple but inefficient (no batching optimization) | Prototyping, teaching |
Selection decision table (inference):
| Scenario | Recommendation | Why |
|---|---|---|
| Production API service (GPU) | vLLM or SGLang | High continuous-batching throughput, API compatible |
| Need function calling / strong JSON constraints | SGLang / vLLM (structured output) | Native constrained decoding support |
| Local PC / no discrete GPU | llama.cpp / Ollama (GGUF quantization) | Runs on low resources |
| NVIDIA large-scale cluster | TensorRT-LLM | Extreme performance (but complex build) |
| Teaching / prototyping | transformers or Ollama | Simple and intuitive |
Trade-offs of Inference Engines: No Silver Bullet
| Need | vLLM | SGLang | TensorRT-LLM | llama.cpp |
|---|---|---|---|---|
| Out-of-the-box (OpenAI-compatible API) | ★★★ | ★★★ | ★★ | ★ (needs separate service) |
| High-concurrency throughput | ★★★ | ★★★ | ★★★ | ★ |
| Structured output / tool calling | ★★ | ★★★ | ★★ | ★ |
| Low resources (CPU / low memory) | ★ | ★ | ★ | ★★★ |
| Community activity / update speed | ★★★ | ★★ | ★★ | ★★★ |
Selection conclusion: Default to vLLM; try SGLang for high-concurrency + complex prompt scenarios; consider TensorRT-LLM only for NVIDIA clusters pursuing extreme throughput; use llama.cpp for personal/edge scenarios. Don't run two engines simultaneously — the inference layer is where operational costs are highest; keeping a single engine is a massive simplification.
Don't use raw transformers for production inference
It lacks continuous batching and KV cache optimization; for the same throughput, the memory/GPU requirements are several times those of vLLM. "Testing the model through in transformers" and "launching a service" must pass through an inference engine conversion.
V. Application Orchestration Layer: LangChain / LlamaIndex / Dify
| Tool | Purpose | Strengths | Weaknesses |
|---|---|---|---|
| LangChain | General orchestration framework (chains, Agents, tools) | Large ecosystem, full components | Thick abstraction layer, frequent updates |
| LlamaIndex | Data ingestion and deep RAG optimization | Most mature RAG components | Data-focused, weaker Agent support |
| Dify | Low-code application platform (visual) | Non-developers can build too; includes RAG/workflow | Limited deep customization |
| Custom | Write calling logic yourself | Fully controllable, no version lock-in | Build from scratch, maintenance cost |
Selection decision table (orchestration):
Your team and tech profile?
├─ Pure backend engineering team, long-term maintenance → Custom lightweight wrapper (most stable)
├─ Quick RAG / multi-Agent prototyping → LlamaIndex or LangChain
├─ Business/product people building apps → Dify
└─ Already locked by framework code → Evaluate "minimal rewrite path", don't keep stackingAgents and Workflows: The Next Stage of Orchestration
When requirements evolve from "single-turn Q&A" to "multi-tool, multi-step," the orchestration layer shifts from RAG to Agents and workflows. Tool calling mechanisms are at LLM-Powered Agents. Key selection points at this layer:
| Solution | Form | Suitable For |
|---|---|---|
| LangGraph | Graph-based workflows (official LangChain) | Complex flows needing fine-grained state control |
| AutoGen / CrewAI | Multi-Agent collaboration | Multi-role division experiments |
| Custom event loop | Manage "call LLM → call tool → fill back" loop yourself | Production long-term maintenance, fully controllable |
| Platforms (Dify/Coze, etc.) | Visual orchestration | Non-engineering teams, quick launch |
Agents are complexity amplifiers
Agent loops (multi-step reasoning + tool calling) have a much larger failure surface than single-turn RAG: tool-calling format errors, runaway loops, cost overruns. Ask yourself first: do you really need an Agent? Many "Agent requirements" can be solved by "workflows + tool functions," an order of magnitude lower in complexity.
"Don't let frameworks hijack you" — the most expensive hidden cost
Two hidden costs of frameworks: API drift (frequent API changes, hard to upgrade once locked) and abstraction leakage (when something breaks, you must read the framework source to debug). Criterion: is the debugging time you spend on the framework exceeding the dev time the framework saves you? The closer the code is to your application core, the more worth it is to build custom.
VI. Vector Database Layer: See RAG in Practice
For a full decision table on selection dimensions, see RAG in Practice Section 3. Quick view:
| Vector DB | Form | Scenario |
|---|---|---|
| FAISS | Library | Prototyping, offline, single-machine |
| pgvector | PG extension | Already using PG, need transactional consistency |
| Qdrant / Milvus | Standalone service | Production-grade, billion-scale vectors, hybrid retrieval |
| Chroma | Embedded | Local teaching, lightweight |
VII. Evaluation Layer: Evaluation Tools
| Tool | Purpose | Characteristics |
|---|---|---|
| lm-evaluation-harness | General benchmark evaluation (MMLU/GSM8K, etc.) | Community standard, comprehensive tasks |
| OpenCompass | Chinese-friendly, multi-model comparison | Good report visualization |
| RAGAS | RAG-specific evaluation | Faithfulness / relevance / answer relevance |
| Custom pipeline | Business golden set | Irreplaceable, see Evaluation in Practice |
Evaluation tools ≠ evaluation systems
lm-eval is just "the script that runs benchmarks." A real evaluation system = your golden set + judge + regression gate + online feedback. Tools execute; systems interpret. See Evaluation in Practice.
VIII. Master Selection Quick-Reference Table
| Layer | Your Scenario | Recommendation (Primary) |
|---|---|---|
| Model library | Anything | transformers |
| Fine-tuning | Single-GPU, under 7B | PEFT/TRL or LLaMA-Factory |
| Distributed training | 70B+, multi-GPU | DeepSpeed / Megatron-LM |
| Production inference | GPU service | vLLM |
| Local inference | Personal / edge | llama.cpp / Ollama |
| App orchestration | Quick prototyping | LangChain / LlamaIndex |
| App orchestration | Long-term maintenance | Custom lightweight wrapper |
| Vector DB | Prototyping | FAISS |
| Vector DB | Production | Qdrant / Milvus / pgvector |
| Evaluation | General benchmarks | lm-eval-harness / OpenCompass |
| Evaluation | Business regression | Custom golden set pipeline |
A Complete Selection Walkthrough (Example)
A team builds a Chinese Q&A system from scratch. Selection walkthrough:
| Decision Point | Scenario | Conclusion |
|---|---|---|
| Model | Chinese, limited budget | Open-source 7B (Qwen series) |
| Fine-tuning | Try prompting first | Skip fine-tuning for now, PEFT as backup |
| Inference | Production API service | vLLM |
| Orchestration | Team is backend engineers | Custom lightweight wrapper |
| Vector DB | Small data volume, already have PG | pgvector |
| Evaluation | Need to build a system | Custom golden set + lm-eval regression |
Selection is "a combination of trade-offs" — "the best" at any single point can be offset by other links (e.g., choosing vLLM but having the orchestration layer locked by a framework, and the whole thing still feels painful).
IX. Experiment Management and Observability
The "last layer" of the toolchain is experiment management and observability: without it, the output of the previous eight layers cannot accumulate.
| Tool | Purpose | Selection Criteria |
|---|---|---|
| Weights & Biases | Training experiment tracking | Training curves, hyperparameter comparison, team sharing |
| MLflow | Experiments + model registry + deployment tracking | Model version management, integration with engineering pipeline |
| Logs / tracing (OpenTelemetry, etc.) | Online request tracing | Request-level latency, token usage, error attribution |
Selection conclusion: Use W&B or MLflow for training experiments; connect online services with standard logging and tracing. Experiment management is not an "optional optimization" but infrastructure for team collaboration and regression debugging — all numbers (Evaluation in Practice) need to be recorded to be trusted.
X. Principles Summary: Three Questions Before Every Selection
Before every selection, ask yourself three questions:
- Do I really need this framework? If a standard library or ten lines of code can solve it, don't introduce one.
- Is it solving my problem or its own? The more the framework's problem domain overlaps with yours, the more worth it is.
- Can I leave it in three years? The harder to leave (the more locked in), the more you should evaluate building custom or wrapping.
Frameworks are liabilities, understanding is assets
Tools go out of date (LangChain's API changes three times a year, vLLM adds new features quarterly), but your understanding of principles like "retrieval → generation," "batching → throughput" never goes out of date. Prefer tools that "teach you principles" over those that "think for you." This perspective also applies to the more complete resource list at Curated Resources.
Further Reading
- Deployment and Serving — Metrics and memory formulas behind inference engines
- RAG in Practice — Practical usage of orchestration frameworks and vector DBs
- Evaluation in Practice — Evaluation tools and custom evaluation systems
- Fine-Tuning in Practice: Full LoRA Workflow — Concrete usage of PEFT/TRL/LLaMA-Factory
- Llama and the Open-Source Ecosystem — Full landscape of open-source models and tool ecosystems
- Mainstream Models: Reference Archive — Model selection (pick the right model before frameworks)
- Curated Resources — More complete tool and learning resource index
References
- Hugging Face Transformers (GitHub) — De facto standard at the model layer
- Microsoft DeepSpeed (GitHub) — Distributed training engine
- NVIDIA Megatron-LM (GitHub) — Large-scale training framework
- Hugging Face PEFT (GitHub) — Parameter-efficient fine-tuning
- Hugging Face TRL (GitHub) — SFT/DPO/PPO trainer
- LLaMA-Factory (GitHub) — One-click fine-tuning platform
- vLLM (GitHub) — Production inference engine
- SGLang (GitHub) — High-performance inference framework
- TensorRT-LLM (GitHub) — NVIDIA inference engine
- llama.cpp (GitHub) — Local inference that runs on CPU
- Ollama — One-click local model runner
- LangChain — Orchestration framework
- LlamaIndex — Data and RAG framework
- Dify — Low-code LLM application platform
- OpenCompass (GitHub) — Evaluation platform