Skip to content

How to Choose Frameworks and Tools

Quick overview scikit-learn, PyTorch, vLLM, or cloud API? Starting from data shape, team capability, and deployment environment, this article gives framework selection comparison tables for three scenarios — tabular data, deep learning, and large models — plus a checkbox decision checklist.

How to Choose Frameworks and Tools ​

One-sentence conclusion: framework selection isn't about choosing the "most popular," but the "best fit for the current problem, team, and deployment environment." The ML tool ecosystem has long passed the era of "everyone using one toolkit" — today's choices are abundant enough to cause decision paralysis, and most "wrong framework" failures aren't because the tool itself is bad, but because the answer was decided first, then the problem was reverse-engineered.

This article splits selection into three decision layers: data shape determines the tool family, team capability determines the upper bound of tool complexity, and deployment environment determines the "last mile" of the tool. Walk through in this order, and most hesitation disappears naturally.

I. Selection Principles: Ask the Question First, Then Choose the Tool ​

1. Three Real Decision Dimensions ​

Popularity, GitHub stars, job posting demand — these are result metrics, not decision bases. Only three dimensions truly determine selection:

                ┌─────────────────────────────────────┐
                │         The problem you're solving    │
                └─────────────────────────────────────┘
                          │            │            │
              ┌───────────▼─┐   ┌───────▼───────┐   ┌▼──────────┐
              │  ① Data shape │   │ ② Team capability  │   │ ③ Deployment env │
              │  Tabular/text/ │   │ Algorithm background │   │ Offline/online/  │
              │  image/time    │   │ Engineering skill/   │   │ cloud/edge/      │
              │                │   │ time budget          │   │ resources/compliance│
              └─────────────┘   └───────────────┘   └───────────┘
                          │            │            │
                          └─────►  Tool family  ◄──────┘
  • Data shape: determines which "tool family" to use. Tabular data → sklearn/GBDT family; image/speech/text → deep learning frameworks; unstructured large-scale text → large model ecosystem. Forcing deep learning frameworks on tabular data, or using tree models on images, is fighting the data shape.
  • Team capability: determines the upper bound of tool complexity. A three-person, two-week-delivery team shouldn't directly deploy distributed training frameworks; a 20-person algorithm team absolutely has the reason to maintain a self-built training platform. Tools are a function of the team, not the other way around.
  • Deployment environment: determines the "last mile." A customer site running only on CPU, an offline-inference compliance scenario, an online service with peak fluctuations — these constraints directly eliminate a batch of "looks great in theory" frameworks.

2. Why "Choosing by Popularity" Is a Trap ​

ChatGPT went viral so everyone wants to use Transformers — not because you solved a text problem, but because it's trendy; the most GitHub-starred framework usually only proves it is easiest to write a hello world for, not that it best fits your production constraints. Historically, TensorFlow once dominated with first-mover advantage, while PyTorch has since surpassed it as the de facto standard (detailed in Section III); ecosystem status flips every three to five years, while your model needs to go live for three to five years. People who treat "trends" as "standards" find themselves always rewriting code.

The right posture: score each of the three dimensions, then let the tool family emerge:

Problem You FacePreferred Tool FamilyAlternative / Mixed Use
Tabular data, structured features, medium scalescikit-learn + GBDT family (XGBoost/LightGBM/CatBoost)AutoML tools for baseline
Unstructured data (images/audio/text)PyTorch familyTensorFlow/Keras, JAX
LLM inference and fine-tuningHugging Face + vLLM / cloud APIOllama (local single machine)
Feature engineering, pipeline orchestrationscikit-learn Pipeline / Polars / feature platformDirectly write business scripts

Minimum actionable step for selection

Don't make a "framework research PPT." Spend half a day writing a minimum viable prototype in each candidate framework, run on your real data, validate in your real deployment environment, and choose the one where "the prototype ran fastest and was least painful engineering-wise." One 8-hour hands-on experiment beats three weeks of trend reports.

II. Tabular Data Scenario: sklearn and the GBDT Three Giants ​

Tabular data (structured data, feature tables) is still the largest share of ML problems in industry — risk control, marketing, supply chain, ops alerts are all tabular problems. Selection here is very mature, with the least controversy.

1. scikit-learn: Ecosystem and Pipeline, Not Precision Ceiling ​

scikit-learn's value isn't in single-model precision, but in how it turns the "full modeling pipeline" into a unified API:

python
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.compose import ColumnTransformer
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import cross_val_score

pipe = Pipeline(steps=[
    ("preprocess", ColumnTransformer([
        ("num", StandardScaler(), ["age", "income"]),
        ("cat", OneHotEncoder(), ["city", "channel"]),
    ])),
    ("model", RandomForestClassifier(n_estimators=200)),
])

scores = cross_val_score(pipe, X_train, y_train, cv=5, scoring="roc_auc")

The fit / predict / transform trio + Pipeline + cross-validation + GridSearchCV form an almost impossible to get wrong evaluation flow. It's also the protocol layer of the entire Python ML ecosystem — XGBoost, LightGBM all offer sklearn-compatible interfaces, meaning you can seamlessly drop GBDT into your Pipeline for model comparison.

sklearn's shortcomings are also clear: natively slower, struggles with big data (some algorithms support n_jobs parallelism but still eat memory), no GPU support, weak deep models (MLP). So its correct usage is pipeline skeleton + baseline model + mixed use with GBDT.

2. The GBDT Three Giants: The Triangle of Precision, Speed, and Ease of Use ​

Gradient Boosted Decision Trees (mechanics in Tree Models and Ensemble Learning) are the ceiling for tabular data. XGBoost, LightGBM, and CatBoost share the same math (all are boosting trees), but have vastly different engineering implementations:

DimensionXGBoostLightGBMCatBoost
Creator / YearTianqi Chen et al., 2014Microsoft, 2017Yandex, 2017
Core accelerationPre-sorting + approximate histogramHistogram + leaf-wise growthSymmetric trees + Ordered Boosting
Training speedFast (faster in hist mode)Fastest, lowest memory usageMedium (very fast on GPU)
PrecisionExtremely strong with good tuning, stableSimilar to XGBoost, slightly worse on small dataWins when many categorical features, anti-overfitting
Categorical featuresNeed manual encodingNative support (category dtype)Best native support, no encoding needed
Missing valuesAutomatically learns split directionAutomatically handlesAutomatically handles
Ease of useMedium (many hyperparameters)MediumMost hassle-free (defaults are competitive)
Ecosystem breadthVery wide (R/Julia/Scala/Spark)Wide (Python/R/C++)Moderately wide (Python/R/Java)
Representative usersEarly Kaggle dominant playerAlibaba, Meituan, large traffic scenariosRecommendation, search, high-cardinality categorical

Three impressive engineering differences:

LightGBM's leaf-wise growth is the source of its speed and its trap. It splits the leaf with the largest gain each time, trains faster, and can approach the optimum more closely, but without a tree-depth limit, it easily overfits (num_leaves defaults to 31, needs to be paired with max_depth and min_data_in_leaf). XGBoost's level-wise layer-by-layer growth is more conservative, hence often more stable on small datasets.

CatBoost's symmetric trees (oblivious trees) have every node at the same depth splitting on the same feature; the structure is simple, naturally anti-overfitting, and inference speed is extremely fast; combined with Ordered Boosting to mitigate prediction shift, it's almost a crushing favorite for "high-cardinality categorical features + moderate sample size" — saves feature engineering, using raw categorical columns directly:

python
from catboost import CatBoostClassifier

model = CatBoostClassifier(iterations=500, learning_rate=0.05, verbose=0)
model.fit(X_train, y_train, cat_features=["city", "device_type"])  # pass categorical column names directly

XGBoost's historical status: it's what turned GBDT into an industrial standard, with regularized objectives, column sampling, GPU support, and all the bells and whistles; in the Spark ecosystem, XGBoostClassifier remains the mainstay for big-data training. If your team needs to integrate with Spark or run multiple languages, XGBoost has the widest ecosystem radius.

3. Conclusion: How to Land in the Tabular Scenario ​

python
# A recommended minimum workflow: sklearn baseline first, then GBDT advanced, AutoML fallback
from sklearn.linear_model import LogisticRegression
from lightgbm import LGBMClassifier

baseline = LogisticRegression(max_iter=1000)          # ① Interpretable baseline
advanced = LGBMClassifier(n_estimators=1000,          # ② Precision workhorse
                          learning_rate=0.05,
                          num_leaves=64,
                          n_jobs=-1)

Practical guidelines

  • Sample size < 10k, seeking interpretability: sklearn (logistic regression / random forest) is enough, don't over-engineer;
  • 10k ~ million-scale tabular data, need precision: LightGBM (default) or XGBoost (when Spark / deep tuning is needed);
  • Extremely many categorical features, don't want to do feature engineering: CatBoost;
  • Don't know how to choose: put all three in an sklearn Pipeline, cross-validate once, whoever wins is who you use — you'll get a conclusion in a day. Tuning details are in Hyperparameter Tuning Practice.

III. Deep Learning Scenario: PyTorch vs TensorFlow vs JAX ​

Entering unstructured data (images, speech, text, video), it's the domain of deep learning frameworks. The landscape in this track today is the result of a ten-year "research-driven vs industrial-driven" tug-of-war.

1. Three Positionings ​

DimensionPyTorchTensorFlowJAX
CreatorMeta (formerly Facebook AI)GoogleGoogle Research
First open source20172015 (TF2 from 2019)2018
Computation graphDynamic graph (define-by-run)Dynamic graph (TF2 eager) + static graph exportFunctional + JIT (jax.jit)
Programming paradigmImperative, close to PythonHigh-level Keras API is simple, low-level complexPurely functional, immutable, composable
Debugging difficultyLowest (native Python errors)Medium (many API layers, version fragmentation)High (XLA compilation errors are opaque)
EcosystemDe facto standard: HF, Lightning, paper reproductionsProduction deployment, mobile, TPUResearch, RL, large-scale scientific computing
Production deploymentTorchServe / ONNX / service-ificationTF Serving / TFLite matureNo official mature serving solution
Learning curveMedium (first learn autograd + DataLoader)Low (Keras one-liner models)High (functional mindset)
Primary usersVast majority of academia, mainstream in industryLegacy systems, mobile/embedded, TPU usersDeepMind-family, RL/frontier research

A minimal comparison, feeling the difference in writing "linear layer + training" across the three:

python
# PyTorch: imperative, print anytime, breakpoint anywhere
import torch
import torch.nn as nn

model = nn.Linear(10, 2)
opt = torch.optim.SGD(model.parameters(), lr=0.01)
for xb, yb in dataloader:
    loss = nn.functional.cross_entropy(model(xb), yb)
    loss.backward()
    opt.step()
    opt.zero_grad()
    print(loss.item())  # debugging is just this natural
python
# JAX: functional + explicit transforms, pip install jax
import jax, jax.numpy as jnp
from jax import grad, jit

def loss_fn(params, x, y):
    pred = x @ params["w"] + params["b"]
    return jnp.mean((pred - y) ** 2)

params = {"w": jnp.zeros((10, 2)), "b": jnp.zeros(2)}
grads = jit(grad(loss_fn))(params, xb, yb)   # jit compiles, grad differentiates, vmap batch-processes

2. Why PyTorch Became the De Facto Standard ​

The answer to this question is the most important lesson in deep learning engineering history:

① Dynamic graphs match research iteration. When PyTorch appeared in 2017, TensorFlow 1.x was still the painful "build static graph first, then execute in a session" model — debugging a graph error required going through a compilation layer. PyTorch's "define-by-run" made the code execution order equal the computation graph construction order; print sees intermediate tensors, pdb can enter the loss internals. For researchers who experiment ten times a day, this is a dimensional strike.

② Paper-reproduction ecosystem snowballs. Official code from top conferences shifted massively to PyTorch from 2018; Hugging Face's Transformers chose PyTorch as its preferred backend, essentially moving the entire NLP domain over. Reproducing someone else's model "pip install and run" became PyTorch's ecosystem moat, and this moat in turn forced new papers to continue using PyTorch — a self-reinforcing loop, once formed, is hard to shake even with later technology.

③ Industry feedback completes the loop. Around 2020, the skepticism that "PyTorch only works for research, not production" was dismantled one by one by torch.compile (compilation acceleration), TorchServe, torch.onnx.export, TensorRT integration, and major self-built inference engines (e.g., vLLM uses the PyTorch ecosystem internally). PyTorch 2.x+'s performance is fundamentally no different from static graph solutions, and the huge convenience of "research code goes directly to production" brought industry to PyTorch too.

Don't step on TensorFlow because of this

PyTorch's mainstream status doesn't mean TensorFlow has no value: legacy systems, mobile/embedded (TFLite), TPU cloud, and a large number of online Keras services won't be migrated in the short term. If your team maintains these systems, continuing with TensorFlow is perfectly reasonable — technical selection follows existing assets and team skills, not community sentiment. JAX continues to shine in reinforcement learning, physics simulation, and large-scale scientific computing; teams working in these frontiers are worth investing in learning.

3. "Signal" for When to Switch Frameworks ​

  • Your training runs on TPU → seriously consider JAX (natively optimal) or TensorFlow;
  • Your model needs to go into iOS/Android/embedded → seriously consider TensorFlow Lite (or PyTorch's ExecuTorch, still maturing);
  • You need large-scale distributed training for 100B+ parameter models → JAX/DeepSpeed family (PyTorch also supports, but JAX is smoother in HPC);
  • None of the above → PyTorch is the default answer, just start; see Deep Learning Fundamentals and CNN Case Studies.

IV. Large Model Scenario: Hugging Face, vLLM, Ollama, and Cloud APIs ​

After 2023, a heavyweight scenario was added to the selection question: large language models (LLM). This isn't a "framework" debate but an "self-hosted vs cloud API" architectural decision, and the supporting toolchain has already solidified.

1. Toolchain Map ​

ToolPositioningTypical UsageTarget Audience
Hugging Face TransformersModel library + unified fine-tuning/inference APIpipeline() runs models in one line; Trainer fine-tunesAlmost everyone using open-source models
Hugging Face HubModel weight hosting and downloadPull Llama, Qwen, DeepSeek, etc. weightsEveryone
vLLMHigh-throughput LLM inference engine (PagedAttention)Deploy open-source models as OpenAI-compatible servicesProduction environments for self-hosted services
OllamaOne-click local model runningollama run qwen2.5:7b, works out of the boxPersonal / prototype / offline single machine
Cloud APICommercial model-as-a-serviceOpenAI, Anthropic, Google Gemini, DeepSeek, Qwen, etc.Quick deployment, teams without GPU

2. Three Tools, Three Postures ​

Hugging Face Transformers is the "unified API" of the open-source model world. What it does is analogous to what sklearn did for traditional ML: converging thousands of architectures (BERT, Llama, Qwen, Mistral...) under the same interface:

python
from transformers import pipeline

# One-line inference: auto-downloads model, loads tokenizer, runs generation
summarizer = pipeline("summarization", model="facebook/bart-large-cnn")
print(summarizer("your long text...", max_length=100)[0]["summary_text"])

On the training side, Trainer wraps the training loop, gradient accumulation, mixed precision, and checkpoints; fine-tuning an open-source model (e.g., SFT on domain data, mechanics in Large Language Models) is typically within a few hundred lines of code. As long as you take the "open-source model route," Hugging Face is the default entry point, no exceptions.

vLLM bridges the gap between "model can run" and "model can handle scale." It optimizes KV cache memory with PagedAttention, supports continuous batching, and can achieve several times the throughput of naive implementations on a single card; it comes with an OpenAI-compatible interface — meaning business code can be written against OpenAI first, then seamlessly switched to self-hosted:

python
# After vLLM deployment, the client is fully isomorphic with calling cloud APIs
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.chat.completions.create(
    model="Qwen/Qwen2.5-7B-Instruct",
    messages=[{"role": "user", "content": "hello"}],
)

Ollama is the entry point for local experience: download, ollama run, done. It reduces the complexity of "running a large model" to the level of "running a Docker container," supports Metal acceleration on macOS and CPU inference, and is a great tool for personal learning, offline demos, and prototype validation — but it doesn't handle high concurrency or multi-machine scheduling; for production scale, switch to vLLM or cloud services.

3. Self-Hosted vs Cloud API: A Real Architectural Decision ​

This is the most important decision table for LLM selection:

Decision DimensionLean Toward Cloud APILean Toward Self-Hosted
Data sensitivityData can cross borders, no compliance constraintsData stays within boundaries, privacy/compliance red lines (healthcare, finance, internal docs)
Call volumeLow traffic, highly variable, validation phaseStable high traffic (TCO breaks even point, self-hosted becomes cheaper)
Capability requirementsNeed the newest and strongest modelsOpen-source models are already sufficient
Latency requirementsCan accept tens to hundreds of milliseconds to secondsUltra-low latency, offline usable
Team capabilityNo GPU, no deployment manpowerHave GPU and engineering team
Cost structureZero fixed cost, pay-per-useOne-time hardware investment + ops

A common decision framework: start with API, end with self-hosted. Use cloud APIs during product validation — fastest and cheapest; when traffic stabilizes, model capability needs solidify, and the TCO calculation shows self-hosting pays back in two years, then migrate to vLLM self-hosting. This "cloud first, self-hosted finally" route avoids two pitfalls: buying GPUs too early and locking suppliers too early.

Three hidden costs of cloud APIs

① Cost out of control: long context, Agent multi-turn calls burn money quickly; must monitor token usage. ② Vendor lock-in: prompt engineering and fine-tuning assets bind to the API, migration costs are high. ③ Data compliance: every piece of text sent leaves your boundary; the legal/security team must approve first. When choosing self-hosted, don't forget that open-source models also need alignment and safety evaluation — "open-source means safe" is not true.

V. Automation and High-Level Tools: AutoML / PyCaret / H2O, When Worth Using ​

There's another category of "tools that select for you" on the selection checklist. AutoML's philosophy: let the tool automatically complete "feature engineering → model selection → hyperparameter search → ensembling," reducing the cognitive cost of the modeling step to a minimum.

1. Main Options ​

ToolOriginCharacteristicsBest Use Case
PyCaretOpen-source low-codeRun full modeling flow + auto-compare 20+ models in a few linesAnalysts, quick validation, teaching
H2O AutoMLH2O.ai (Java core)Veteran enterprise-grade, auto-does Stacked Ensemble, supports R/PythonEnterprise-scale tabular modeling
AutoGluonAWS open-sourceAuto-ensemble king for tabular data, multi-round stacking, extremely strong precisionTabular data chasing upper-limit precision
FLAMLMicrosoft open-sourceLightweight, embeddable in existing code, cost-aware searchAuto-tuning within existing pipelines

A "three-line full modeling" example in PyCaret, feeling the high-level tool's hand-feel:

python
from pycaret.classification import setup, compare_models, finalize_model

setup(data, target="churn", session_id=42)   # ① auto data cleaning/splitting
best = compare_models()                       # ② auto-compare dozens of models + tuning
final_model = finalize_model(best)            # ③ retrain on full data

2. When Worth Using, When Not ​

Worth using:

  • Baseline phase: AutoML's results from one hour usually beat a novice's week of manual tuning — let it tell you "what score this problem theoretically reaches," then decide whether to go deep manually;
  • No dedicated algorithm engineer on the team: business analysts can deliver 80-point models with PyCaret/AutoGluon;
  • Tabular data precision sprints: AutoGluon's stacking ensembles can approach or even exceed manual Tuning ceilings in many Kaggle tabular competitions.

Not worth using:

  • Problem definition unclear: AutoML can only optimize the metric you give it; if you haven't defined "what to predict, how to measure," the tool only amplifies errors;
  • Poor data quality: dirty data, target leakage, sample bias — AutoML won't help you find them, it will only "elegously" swallow them;
  • Need interpretability and controllability: AutoML's Stacked Ensemble is black box inside black box; risk control / healthcare scenarios become harder to explain;
  • Treating AutoML as a religion: it solves only the small ring of "algorithm selection"; it can't manage feature acquisition, data pipelines, deployment monitoring — these account for 80% of project workload.

Positioning AutoML

AutoML doesn't "replace data scientists"; it compresses the "algorithm trial-and-error" step. Mature teams' usage: AutoML quickly produces a baseline → humans do feature engineering and domain modeling on the baseline → manually take over tuning when necessary. Treat it as "a very fast second colleague" not "the boss."

VI. Experiment Management: MLflow and W&B ​

One of the most regret-not-doing-sooner parts of the selection checklist is experiment management. The typical symptoms of not doing experiment management: model files called model_final_v2_really_final.pt, metrics logged in chat history, reproducing a result requires flipping through half of history.

1. Two Main Solutions ​

DimensionMLflowWeights & Biases (W&B)
FormatOpen-source, deployable privatelyCommercial SaaS (free tier), also supports private deployment
Core capabilitiesTracking (metrics/params/artifacts) + Models Registry + ProjectsExperiment tracking + visualization + Sweep hyperparameter search + team collaboration
Deployment costSelf-built (SQLite/Postgres + storage)Zero ops, register and go
StrengthsEngineering: model registry, version management, integration with serving/CIResearch experience: interactive curves, shareable dashboards, Sweep automated search
Best forProduction teams, need private deployment, need to enter MLOps pipelineResearch teams, heavy on visualization and collaboration

A minimal MLflow usage — it changes "the way you write experiments":

python
import mlflow

with mlflow.start_run():
    mlflow.log_param("learning_rate", 0.01)
    mlflow.log_param("num_leaves", 64)
    mlflow.log_metric("val_auc", 0.921)
    mlflow.log_metric("val_logloss", 0.31)
    mlflow.log_artifact("model.pkl")          # artifacts auto-archived, no more v2_final

One command mlflow ui starts a local dashboard; all experiment params, metrics, and model files are traceable on the timeline. Start using it from the very first serious experiment — it's an order of magnitude cheaper than adding it after running 50 experiments.

2. Three Levels of Experiment Management ​

  1. Personal level: MLflow Tracking or W&B free tier,accumulating "params + metrics + artifacts";
  2. Team level: model registry center (MLflow Models / W&B Registry), unify "which model is a production candidate," paired with code review and regression comparison;
  3. Organizational level: integrated into a full MLOps pipeline — training, evaluation, deployment, monitoring, rollback closed loop. This level's content is in MLOps and Model Lifecycle, the natural extension of "experiment management."

Don't over-engineer

If your project is still in "exploration phase, model not yet solid," a local MLflow service + team shared dashboard is enough. Don't deploy K8s scheduling + feature platform + full-link monitoring in week one — experiment management tools, like any others, should be upgraded when the scale demands it; premature infrastructure is another kind of waste.

VII. Notebooks vs Scripts vs Project Engineering: When to Graduate from Notebooks ​

The last "tool" is the organizational form of code. Many teams didn't get the framework wrong; they died because "the code stays in notebooks forever."

1. Three Form Factor Positionings ​

DimensionJupyter NotebookPython ScriptsProject Engineering (packages/pipelines)
PositioningExploration, visualization, storytellingReproduction, batch processing, one-off tasksMaintainable, testable, deployable systems
State managementImplicit global state, cell-order sensitiveFile-level, re-run = reproducibleExplicit I/O, parameterizable
TestabilityAlmost impossibleModerateUnit tests + integration tests
Version controlVery poor (ipynb diff is a nightmare)GoodGood
Suitable stageProblem exploration, EQA analysis, writing docsData pipelines, training scriptsProduction services, complex products
Not suitableProduction, long pipelines, multi-person collaborationInteractive analysisQuick validation

2. Four Signals for Graduating from Notebooks ​

  • Notebook exceeds 500–1000 lines, or "must execute cells in order or it errors" — you've written an untestable, unmaintainable script, just disguised as a notebook;
  • Code goes to production: needs monitoring, scheduled runs, maintenance by others — notebooks hide implicit state; if any step's df is secretly modified by the previous one, it's a ticking bomb;
  • Need unit tests and code review: model code correctness (data split leakage, feature alignment, scaler parameter consistency) can only be guaranteed by testing;
  • Multi-person collaboration: merging ipynb in git is painful; use scripts + configurability (config files / CLI arguments) instead.

3. A Pragmatic Migration Path ​

notebook (exploration)
   │ Extract "determined logic" into functions
   ▼
utils.py / train.py (scripts, with --config params, rerunnable)
   │ Add data/model abstractions, add tests
   ▼
src/ package structure + pipeline orchestration (pipeline script or MLflow project)
   │ Connect to deployment and monitoring
   ▼
Production system (serving / batch processing platform / MLOps pipeline)

Realistic advice

No need for all-in engineering. Staying in notebooks during the exploration phase is completely correct — it's the most efficient analysis tool. The key is the timing of "graduation": when a piece of code is about to be reused a second time, or needs to run into production, extract it into a function and put it in a script immediately. Every time you copy-paste from a notebook to production code, you're creating debt for your future self. For a complete project setup workflow, see Building an ML Project from Scratch.

VIII. Decision Checklist: A Checklist You Can Actually Check Off ​

Converging the entire article into a decision checklist you can follow line by line. From "the most expensive question" to "the cheapest tool," each line corresponds to a section above:

Step 1: Freeze the problem first (prerequisites)

  • [ ] I've clarified "what to predict, for whom, and what the success metric is"
  • [ ] I've confirmed data is accessible, data quality is preliminarily reviewed (no target leakage, no major missing data)
  • [ ] I'm clear on constraints: can data cross borders, latency requirements, available compute

Step 2: Determine tool family by data shape

  • [ ] Tabular / structured data → select [sklearn baseline] + [GBDT: LightGBM/XGBoost/CatBoost]
  • [ ] Images / audio / text / unstructured → select [PyTorch] (unless TPU/mobile/existing TensorFlow constraints)
  • [ ] Large language model tasks → go to Step 4

Step 3: Refine by team and deployment constraints

  • [ ] Team is mostly business analysts, time-pressed → introduce AutoML (PyCaret/AutoGluon) for baseline
  • [ ] CPU only / customer site offline → confirm GBDT and sklearn as primary, avoid heavy deep learning
  • [ ] Need to go on Spark / multi-language → XGBoost ecosystem first
  • [ ] Need interpretability (risk control / healthcare) → linear models + tree models primary, use AutoML black-box ensembles with caution

Step 4: LLM-specific decision

  • [ ] Data-sensitive / compliance red lines → self-hosted (vLLM + open-source models)
  • [ ] Quick validation / low traffic / need latest capability → cloud API
  • [ ] Local single machine / offline demo → Ollama
  • [ ] Taking open-source models → default Hugging Face entry
  • [ ] Stable traffic and self-hosted TCO is favorable → migrate from API to vLLM

Step 5: Engineering infrastructure, done right the first time

  • [ ] Connected experiment management (MLflow or W&B) from the first experiment
  • [ ] Code evolves at the "notebook exploration → script solidification → engineering" pace, no debt carried
  • [ ] Confirmed deployment path (Serving/batch processing/edge) and selection compatibility, no "framework runs but can't deploy"
  • [ ] Finally: spend a day running a minimal prototype on real data with each candidate framework, use the result to make the final call

After checking through this line, your selection rationale can be explained to both the tech team and business stakeholders — the most important thing in selection is "having evidence to support it," not "everyone was using it at the time."

Further Reading ​

References ​