Skip to content

Frameworks and Tooling: A Selection Guide

At a glance A six-layer breakdown of the LLM toolchain — model libraries → training → inference → orchestration → vector DBs → evaluation — with scenario-driven selection decision tables and the "don't let frameworks hijack you" principles plus three self-build criteria.

Frameworks and Tooling: A Selection Guide ​

The essence of tool selection is matching complexity to your scenario: frameworks solve 80% of common problems at the cost of 20% flexibility. Picking the wrong framework doesn't just mean "installing one extra package" — it locks the entire team into the framework's design.

This article breaks the LLM toolchain into six layers, each with a scenario → recommendation decision table. After reading, you should have a "selection map that gives you confidence" rather than "pick whatever's trending." The full open-source ecosystem landscape is at Llama and the Open-Source Ecosystem.

I. Layered Tool Map ​

┌─────────────────────────────────────────────┐
│ Application layer  Dify / custom apps        │
├─────────────────────────────────────────────┤
│ Orchestration layer  LangChain / LlamaIndex   │
├─────────────────────────────────────────────┤
│ Evaluation layer  lm-eval-harness / OpenCompass / RAGAS │
├─────────────────────────────────────────────┤
│ Inference layer  vLLM / SGLang / TensorRT-LLM / llama.cpp │
├─────────────────────────────────────────────┤
│ Training layer  transformers / PEFT / TRL / DeepSpeed  │
│                 Megatron-LM / LLaMA-Factory / Axolotl    │
├─────────────────────────────────────────────┤
│ Model & data layer  Transformers / Datasets / Tokenizers │
└─────────────────────────────────────────────┘

General selection principles (applied in every layer below):

  1. Default to the thinnest: If a standard library or API can handle it, don't introduce a framework.
  2. Get it working before abstracting: Implement the first version the most direct way; introduce abstractions only when repeated needs emerge.
  3. Maintain replaceability: Wrap core modules (e.g., vector DBs, inference engines) behind interfaces for easy swapping.

II. Model and Data Layer: Nearly Mandatory ​

ToolPurposeWhen to Pick
Hugging Face transformersStandard library for model loading/inference/trainingAlmost all scenarios (de facto standard)
datasetsDataset loading and preprocessingMandatory for training/fine-tuning
tokenizersHigh-performance tokenizationWhen building custom tokenizers

transformers's single-line from_pretrained loads models with unified AutoModel/AutoTokenizer interfaces, turning "swap models" from code changes into parameter changes — this is the foundation of the entire ecosystem. Without special reason, skip selection at the model layer and just use transformers.

III. Training Layer: Pick by Scale and Method ​

ToolPurposeApplicable ScenariosNot Applicable
transformers TrainerSingle-GPU / small-scale trainingSmall model fine-tuning, teachingLarge-scale distributed training
PEFTLoRA/QLoRA and other parameter-efficient fine-tuningMost fine-tuning tasksFull-parameter fine-tuning
TRLSFT/DPO/PPO trainerInstruction fine-tuning and alignmentCustom training loops
LLaMA-FactoryOne-click platformQuick experiments, switching methodsDeep customization
AxolotlConfiguration-driven trainingReproducing configs, batch experimentsDynamic logic needs
DeepSpeedDistributed training engine (ZeRO)Large-model multi-GPU trainingSmall-model single-GPU
Megatron-LMNVIDIA large-scale training (tensor parallelism, etc.)10B+ parameter trainingConventional fine-tuning

Training selection decision table:

What are you doing?
├─ Fine-tuning under 7B, fits on one GPU → PEFT/TRL or LLaMA-Factory
├─ Fine-tuning 13B+, need quantization for memory → QLoRA (bitsandbytes) + PEFT
├─ Pretraining / full-parameter 70B+ multi-GPU → DeepSpeed (ZeRO-3) or Megatron-LM
└─ Alignment (RLHF/DPO) → TRL (DPOTrainer) or custom + trlx, etc.

Three Disciplines of Training Engineering ​

Beyond frameworks, the most often overlooked aspect of training is engineering discipline:

DisciplinePracticeProblem Solved
Reproducible experimentsFix random seeds + record full hyperparameters (including data version, framework version)Two experiments give different conclusions
Checkpoint managementSave regularly, keep best by eval scoreOverfitting / rollback after incidents
Training logsRecord loss / gradient norm / learning rate per stepLocate divergence, judge convergence

Training logs' value lies in offline postmortems: debugging a "fine-tuned but dumber" model relies on eval loss inflection points in the training logs.

A realistic note on training selection

Don't introduce DeepSpeed for problems solvable by PEFT/TRL. The complexity of distributed training (node communication, gradient synchronization, mixed precision) is an order-of-magnitude increase. Ask yourself first: is single-GPU + QLoRA enough? Full workflow at Fine-Tuning in Practice: Full LoRA Workflow.

IV. Inference Layer: The Engineering Battlefield of Throughput and Latency ​

Inference engines directly determine deployment costs and user experience. Metric definitions (TTFT/TPOT/throughput) are at Deployment and Serving.

EngineCharacteristicsApplicable Scenarios
vLLMPagedAttention, continuous batching, OpenAI-compatible APIProduction default pick
SGLangRadixAttention, structured output optimizationHigh concurrency, complex prompting, tool calling
TensorRT-LLMNVIDIA deep optimization, TensorRT compilationNVIDIA clusters, extreme throughput
llama.cpp (GGUF)Runs on CPU / low memory, fully localLocal / edge, personal use, offline
OllamaWrapper around llama.cpp, one-click runLocal experience, teaching demos
transformers native inferenceSimple but inefficient (no batching optimization)Prototyping, teaching

Selection decision table (inference):

ScenarioRecommendationWhy
Production API service (GPU)vLLM or SGLangHigh continuous-batching throughput, API compatible
Need function calling / strong JSON constraintsSGLang / vLLM (structured output)Native constrained decoding support
Local PC / no discrete GPUllama.cpp / Ollama (GGUF quantization)Runs on low resources
NVIDIA large-scale clusterTensorRT-LLMExtreme performance (but complex build)
Teaching / prototypingtransformers or OllamaSimple and intuitive

Trade-offs of Inference Engines: No Silver Bullet ​

NeedvLLMSGLangTensorRT-LLMllama.cpp
Out-of-the-box (OpenAI-compatible API)★★★★★★★★★ (needs separate service)
High-concurrency throughput★★★★★★★★★★
Structured output / tool calling★★★★★★★★
Low resources (CPU / low memory)★★★★★★
Community activity / update speed★★★★★★★★★★

Selection conclusion: Default to vLLM; try SGLang for high-concurrency + complex prompt scenarios; consider TensorRT-LLM only for NVIDIA clusters pursuing extreme throughput; use llama.cpp for personal/edge scenarios. Don't run two engines simultaneously — the inference layer is where operational costs are highest; keeping a single engine is a massive simplification.

Don't use raw transformers for production inference

It lacks continuous batching and KV cache optimization; for the same throughput, the memory/GPU requirements are several times those of vLLM. "Testing the model through in transformers" and "launching a service" must pass through an inference engine conversion.

V. Application Orchestration Layer: LangChain / LlamaIndex / Dify ​

ToolPurposeStrengthsWeaknesses
LangChainGeneral orchestration framework (chains, Agents, tools)Large ecosystem, full componentsThick abstraction layer, frequent updates
LlamaIndexData ingestion and deep RAG optimizationMost mature RAG componentsData-focused, weaker Agent support
DifyLow-code application platform (visual)Non-developers can build too; includes RAG/workflowLimited deep customization
CustomWrite calling logic yourselfFully controllable, no version lock-inBuild from scratch, maintenance cost

Selection decision table (orchestration):

Your team and tech profile?
├─ Pure backend engineering team, long-term maintenance → Custom lightweight wrapper (most stable)
├─ Quick RAG / multi-Agent prototyping → LlamaIndex or LangChain
├─ Business/product people building apps → Dify
└─ Already locked by framework code → Evaluate "minimal rewrite path", don't keep stacking

Agents and Workflows: The Next Stage of Orchestration ​

When requirements evolve from "single-turn Q&A" to "multi-tool, multi-step," the orchestration layer shifts from RAG to Agents and workflows. Tool calling mechanisms are at LLM-Powered Agents. Key selection points at this layer:

SolutionFormSuitable For
LangGraphGraph-based workflows (official LangChain)Complex flows needing fine-grained state control
AutoGen / CrewAIMulti-Agent collaborationMulti-role division experiments
Custom event loopManage "call LLM → call tool → fill back" loop yourselfProduction long-term maintenance, fully controllable
Platforms (Dify/Coze, etc.)Visual orchestrationNon-engineering teams, quick launch

Agents are complexity amplifiers

Agent loops (multi-step reasoning + tool calling) have a much larger failure surface than single-turn RAG: tool-calling format errors, runaway loops, cost overruns. Ask yourself first: do you really need an Agent? Many "Agent requirements" can be solved by "workflows + tool functions," an order of magnitude lower in complexity.

"Don't let frameworks hijack you" — the most expensive hidden cost

Two hidden costs of frameworks: API drift (frequent API changes, hard to upgrade once locked) and abstraction leakage (when something breaks, you must read the framework source to debug). Criterion: is the debugging time you spend on the framework exceeding the dev time the framework saves you? The closer the code is to your application core, the more worth it is to build custom.

VI. Vector Database Layer: See RAG in Practice ​

For a full decision table on selection dimensions, see RAG in Practice Section 3. Quick view:

Vector DBFormScenario
FAISSLibraryPrototyping, offline, single-machine
pgvectorPG extensionAlready using PG, need transactional consistency
Qdrant / MilvusStandalone serviceProduction-grade, billion-scale vectors, hybrid retrieval
ChromaEmbeddedLocal teaching, lightweight

VII. Evaluation Layer: Evaluation Tools ​

ToolPurposeCharacteristics
lm-evaluation-harnessGeneral benchmark evaluation (MMLU/GSM8K, etc.)Community standard, comprehensive tasks
OpenCompassChinese-friendly, multi-model comparisonGood report visualization
RAGASRAG-specific evaluationFaithfulness / relevance / answer relevance
Custom pipelineBusiness golden setIrreplaceable, see Evaluation in Practice

Evaluation tools ≠ evaluation systems

lm-eval is just "the script that runs benchmarks." A real evaluation system = your golden set + judge + regression gate + online feedback. Tools execute; systems interpret. See Evaluation in Practice.

VIII. Master Selection Quick-Reference Table ​

LayerYour ScenarioRecommendation (Primary)
Model libraryAnythingtransformers
Fine-tuningSingle-GPU, under 7BPEFT/TRL or LLaMA-Factory
Distributed training70B+, multi-GPUDeepSpeed / Megatron-LM
Production inferenceGPU servicevLLM
Local inferencePersonal / edgellama.cpp / Ollama
App orchestrationQuick prototypingLangChain / LlamaIndex
App orchestrationLong-term maintenanceCustom lightweight wrapper
Vector DBPrototypingFAISS
Vector DBProductionQdrant / Milvus / pgvector
EvaluationGeneral benchmarkslm-eval-harness / OpenCompass
EvaluationBusiness regressionCustom golden set pipeline

A Complete Selection Walkthrough (Example) ​

A team builds a Chinese Q&A system from scratch. Selection walkthrough:

Decision PointScenarioConclusion
ModelChinese, limited budgetOpen-source 7B (Qwen series)
Fine-tuningTry prompting firstSkip fine-tuning for now, PEFT as backup
InferenceProduction API servicevLLM
OrchestrationTeam is backend engineersCustom lightweight wrapper
Vector DBSmall data volume, already have PGpgvector
EvaluationNeed to build a systemCustom golden set + lm-eval regression

Selection is "a combination of trade-offs" — "the best" at any single point can be offset by other links (e.g., choosing vLLM but having the orchestration layer locked by a framework, and the whole thing still feels painful).

IX. Experiment Management and Observability ​

The "last layer" of the toolchain is experiment management and observability: without it, the output of the previous eight layers cannot accumulate.

ToolPurposeSelection Criteria
Weights & BiasesTraining experiment trackingTraining curves, hyperparameter comparison, team sharing
MLflowExperiments + model registry + deployment trackingModel version management, integration with engineering pipeline
Logs / tracing (OpenTelemetry, etc.)Online request tracingRequest-level latency, token usage, error attribution

Selection conclusion: Use W&B or MLflow for training experiments; connect online services with standard logging and tracing. Experiment management is not an "optional optimization" but infrastructure for team collaboration and regression debugging — all numbers (Evaluation in Practice) need to be recorded to be trusted.

X. Principles Summary: Three Questions Before Every Selection ​

Before every selection, ask yourself three questions:

  1. Do I really need this framework? If a standard library or ten lines of code can solve it, don't introduce one.
  2. Is it solving my problem or its own? The more the framework's problem domain overlaps with yours, the more worth it is.
  3. Can I leave it in three years? The harder to leave (the more locked in), the more you should evaluate building custom or wrapping.

Frameworks are liabilities, understanding is assets

Tools go out of date (LangChain's API changes three times a year, vLLM adds new features quarterly), but your understanding of principles like "retrieval → generation," "batching → throughput" never goes out of date. Prefer tools that "teach you principles" over those that "think for you." This perspective also applies to the more complete resource list at Curated Resources.

Further Reading ​

References ​