Theme
Building an ML Project from Scratch
Reading ten tutorials is worth less than running a complete project end-to-end. This article takes you through every stage of a real project using the classic, minimal "Titanic survival prediction" — not "calling APIs to get results," but making clear why every step is done the way it is.
Many beginners, having learned many algorithms, still don't know what an ML project actually looks like: how are files organized? what's the order of steps? which parts can be skipped, which absolutely can't? This article walks the complete project workflow from start to finish using one thread (Titanic survival prediction). What you get isn't "0.81 score on this dataset," but a process skeleton that can be applied to any tabular classification problem.
A complete project has seven steps; first, the overall map:
Raw data
│
▼
① Data loading and exploration (EDA) ── Answer: "what does the data look like, what problems?"
▼
② Data cleaning and feature engineering ── Answer: "which columns enter the model, how to turn them into numbers?"
▼
③ Split train / val / test ── Answer: "how to guarantee the evaluation is honest?"
▼
④ Baseline model ────────── Answer: "what's the best a simple solution can do?"
▼
⑤ Iterative improvement ─── Answer: "how much more can a more complex model add?"
▼
⑥ Evaluation and error analysis ── Answer: "where does the model err, and why?"
▼
⑦ Result presentation and visualization ─ Answer: "how to explain conclusions clearly?"Throughout the process, we deliberately follow one discipline: the test set is only touched once, at the end. This discipline runs through the entire article, and it's the defense against #1 in Common Pitfalls: "data leakage."
Prerequisites
This article assumes you understand the basic concepts: what supervised learning is (Supervised Learning), how models are evaluated (Model Evaluation and Validation), what feature engineering is (Feature Engineering). If you need to fill basics, read What is Machine Learning and Overall Architecture Anatomy first, then come back.
I. Project Selection: A Simple but Real Problem
The first question: what to choose for a practice project? Many people's first project is "predicting stock prices with LSTM," which almost certainly fails — not because the technology is hard, but because the problem itself is unsolvable. There are three criteria for choosing a project, all required.
| Criterion | Why | Consequence of Violating |
|---|---|---|
| Has labels (has ground truth) | You need to know "the correct answer" to evaluate whether the model learned | Without labels, there's nothing to evaluate; you're just feeling good about yourself |
| Small scale (hundreds to thousands of samples) | Each iteration runs in minutes, so you can focus on process and concepts | Big data projects spend half the time tuning Spark; you don't learn modeling |
| Interpretable (human-understandable feature meanings) | You can use common sense to judge whether features are reasonable; error analysis works | "Why it's wrong" for images/text needs additional tools; beginners can't get started |
By these criteria, here's a comparison of three classic practice projects:
| Project | Labels | Scale | Interpretability | Suitability |
|---|---|---|---|---|
| Titanic survival prediction | Yes (Survived) | 891 training samples | ★★★ All features are plain language (age, cabin class, gender) | ★★★ Beginner's first choice |
| Boston / California housing price prediction | Yes (continuous values) | ~20k | ★★☆ Features have real-world meaning | ★★★ Best for regression practice |
| Spam email classification | Yes (spam/ham) | 5k+ emails | ★☆☆ Features are word vectors; need text processing | ★★☆ Slightly advanced |
This article chooses Titanic survival prediction, for three reasons:
- Clear labels: whether each passenger survived is a determined historical fact, no "label subjectivity" issue.
- Rich and interpretable features: age, gender, cabin class (Pclass), fare, boarding port — every column can be explained by common sense, suitable for error analysis.
- Small scale: 891 training samples; any model runs in milliseconds, so you can iterate freely.
What the Dataset Looks Like
The Titanic dataset (from the Kaggle same-named entry competition) in its raw form:
| Column | Meaning | Type | Has Missing? |
|---|---|---|---|
| PassengerId | Passenger ID | int | No |
| Survived | Whether survived (1=yes, 0=no) | int (label) | No |
| Pclass | Cabin class (1/2/3) | int | No |
| Name | Name | str | No |
| Sex | Gender (male/female) | str | No |
| Age | Age | float | Yes (177 missing) |
| SibSp | Number of siblings/spouses aboard | int | No |
| Parch | Number of parents/children aboard | int | No |
| Ticket | Ticket number | str | No |
| Fare | Fare | float | No |
| Cabin | Cabin number | str | Yes (687 missing!) |
| Embarked | Boarding port (C/Q/S) | str | Yes (2 missing) |
Note: these 891 rows are the training set for the Kaggle competition. In real projects, data is often scattered across different systems and formats, and loading and cleaning take up half the work — precisely because of this, the "clean dataset" training style is suitable for learning the workflow. For dataset search and evaluation methods, see Datasets and Tools Archive.
Define "success" before writing code
When translating a problem into a machine learning task, you must answer three questions:
- Task type: survival prediction is binary classification (survived / died).
- Evaluation metric: in this dataset, ~38% survived — this is class imbalanced. Accuracy is misleading (see Misused Metrics), so we use ROC-AUC as the primary metric, supplemented by precision / recall.
- Baseline: a model that "predicts all died" has ~62% accuracy. Any model of yours must significantly exceed this number, otherwise it means nothing was learned.
Project success is defined before code is written — this is the basic skill of Model Evaluation and Validation.
II. Environment Setup
1. Version Requirements
This article's code is based on the following versions (2024 stable versions, backwards compatible with older ones):
| Software | Version | Purpose |
|---|---|---|
| Python | 3.10+ | Language itself |
| pandas | 2.x | Data loading and processing |
| scikit-learn | 1.3+ | Modeling, splitting, evaluation, tuning |
| matplotlib | 3.x | Visualization |
| jupyterlab | 4.x | Interactive exploration |
2. Create a Virtual Environment (venv)
Never install dependencies into system Python. A virtual environment is an isolation container for project dependencies, ensuring no cross-contamination when switching machines / projects:
bash
# Enter project root directory
mkdir titanic && cd titanic
# Create virtual environment (activation commands differ between Windows and macOS/Linux)
python -m venv .venv
source .venv/bin/activate # macOS / Linux
# .venv\Scripts\activate # Windows (PowerShell)
# .venv\Scripts\activate.bat # Windows (CMD)
# Install dependencies (one-line install)
pip install pandas scikit-learn matplotlib jupyterlab
# Freeze the dependency list, ensuring others can reproduce with one click
pip freeze > requirements.txtWhy use venv instead of global install
- Different projects may depend on different versions of numpy/sklearn; global install forces you to "upgrade/downgrade back and forth."
requirements.txtgenerated bypip freezeis the first step to reproducible experiments — others get the exact same environment withpip install -r requirements.txt.- In team collaboration, reproducible environment = reproducible experiment results.
3. Start Jupyter and Verify
bash
jupyter labCreate a new notebook, first run the "environment smoke test":
python
import pandas as pd, numpy as np, sklearn, matplotlib
print("pandas:", pd.__version__)
print("sklearn:", sklearn.__version__)
print("numpy:", np.__version__)
# One-line smoke test: confirm sklearn's core components all work
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
print("Environment OK")When version numbers output, the environment is ready. Recommend putting EDA in notebooks, putting reusable logic in .py scripts (directory structure in Section IV) — notebooks suit exploration, scripts suit reuse and testing.
III. Complete Workflow Step by Step
Step 1: Data Loading and Exploration (EDA)
EDA (Exploratory Data Analysis) isn't about "looking at data"; it's about answering three questions: what problems does the data have (missing, types), what do individual features look like (distributions), and what's the relationship between features and labels (where is the signal)?
python
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
# Load data (downloaded train.csv from Kaggle, also available from this site's dataset archive)
df = pd.read_csv("data/train.csv")
# ── ① Check shape and head ────────────────────────────────
print("Shape:", df.shape) # (891, 12) → 891 passengers, 12 columns
df.head()
# ── ② Check per-column types and missing ──────────────────
df.info()
# ── ③ Check numeric column stats (count/mean/std/min/25%/50%/75%/max)──
df.describe()
# ── ④ Check missing values at a glance ────────────────────
df.isnull().sum()df.info() immediately exposes three problems: Age is missing 177, Cabin is missing 687, Embarked is missing 2. df.describe() exposes a fourth problem: Fare's standard deviation is 49.7, mean 32 while median 14 — indicating severe right skew (a few first-class tickets with extremely high prices pulled the mean up).
Next, explore the target distribution:
python
# Survival ratio: 38.4%, meaning the majority class is "died" (62%)
print(df["Survived"].value_counts(normalize=True))
# 0 0.616162
# 1 0.383838Then the relationship between features and labels — this is where EDA is most information-dense:
python
# Gender and survival: female survival rate 74.2%, male 18.9% — the strongest signal
print(df.groupby("Sex")["Survived"].mean())
# female 0.742038
# male 0.188908
# Cabin class and survival: first class 63.0%, second 47.3%, third 24.2%
print(df.groupby("Pclass")["Survived"].mean())
# 1 0.629630
# 2 0.472826
# 3 0.242363
# Age distribution (look at non-missing part for now)
df["Age"].hist(bins=30)
plt.title("Age Distribution")
plt.show()Write down the EDA conclusions (not just glance at them); these are the basis for subsequent feature engineering:
| Finding | Meaning | Subsequent Action |
|---|---|---|
Survived ~38% survival | Class imbalance | Use AUC/recall, don't look at raw accuracy |
Sex strongly correlated with survival (female 74% vs male 19%) | Core signal | Must encode into the model |
Pclass strongly correlated with survival (decreasing 1→3) | Core signal | Keep, consider making it a categorical feature |
Age missing 177, Cabin missing 687 | Severe missingness | Needs imputation or exclusion (next step) |
Fare severely right-skewed | Skewed distribution | Log transform or binning |
Name / Ticket / PassengerId | High cardinality / identifiers | Little direct predictive power, but can extract features |
Signs of a good EDA
A good EDA produces a findings list, where each finding leads to a subsequent decision. If you do EDA just by "printing a dozen tables and moving on," you haven't entered the right state. EDA is essentially hypothesis testing: you come with hypotheses about "what factors affect survival," and use data to validate or refute them.
Step 2: Data Cleaning and Feature Engineering
EDA findings need to be implemented in code. General principle: the output of cleaning and feature engineering is a unified X (feature matrix) + y (label), and all our models consume this pair.
① Missing value handling
| Column | Missing rate | Handling method | Reason |
|---|---|---|---|
Age | 20% | Fill with median (or group-fill by sex/cabin class) | Age is right-skewed; median is more robust than mean |
Embarked | 0.2% | Fill with mode (most frequent S) | Only 2 rows, negligible impact |
Cabin | 77% | Don't fill directly; convert to "has cabin number" binary feature | Missingness itself may contain info (low survival rate for those without cabin records); filling with fake values introduces noise |
python
# Age: group-fill is more fine-grained than global-fill (median by sex/cabin differs)
df["Age"] = df.groupby(["Sex", "Pclass"])["Age"].transform(
lambda s: s.fillna(s.median())
)
# Embarked: mode fill
df["Embarked"] = df["Embarked"].fillna(df["Embarked"].mode()[0])
# Cabin: convert to binary feature "has cabin record" (higher information density than the cabin itself)
df["HasCabin"] = df["Cabin"].notna().astype(int)
df = df.drop(columns=["Cabin"])The "leakage" red line for missing value handling
All quantities learned from data (median, mode, mean, min/max, regression coefficients) must only be computed from the training set, then applied to validation/test sets. Doing fillna(df["Age"].median()) on the full dataset is fine for practice, but in real projects, this is a form of data leakage — test-set information flows into the training set. The rigorous way is to put it in a Pipeline (see below), letting sklearn automatically "fit on the training set, transform on other sets."
② Categorical encoding
Sex and Embarked are strings; models only accept numbers:
python
# One-Hot encoding: Embarked has 3 classes → 3 0/1 columns (sklearn auto-drops the first to avoid collinearity)
df = pd.get_dummies(df, columns=["Sex", "Embarked"], drop_first=True)Sex_male becomes 1/0, Embarked_Q, Embarked_S each 1/0. Note: Pclass has numeric form 1/2/3, but it's a rank (ordinal category) not a continuous quantity; treating it as a categorical feature is more stable — we use pd.get_dummies to expand Pclass into 3 columns, or directly keep the original value and let the tree model handle it (trees don't assume linear relationships of numbers).
③ Manual feature engineering: letting domain knowledge enter the model
The title (Mr/Mrs/Miss/Master) in Name is a classic feature engineering example — it strongly correlates with age and gender, especially useful for distinguishing age groups with missing values:
python
df["Title"] = df["Name"].str.extract(r" ([A-Za-z]+)\.", expand=False)
# Merge rare titles, reduce categories
df["Title"] = df["Title"].replace(
["Lady","Countess","Capt","Col","Don","Dr","Major","Rev","Sir","Jonkheer","Dona"],
"Rare"
)
df["Title"] = df["Title"].replace(["Mlle","Ms"], "Miss")
df["Title"] = df["Title"].replace(["Mme"], "Mrs")
df = pd.get_dummies(df, columns=["Title"], drop_first=True)Then construct a family size feature — "whole families together" is a significant factor in survival:
python
df["FamilySize"] = df["SibSp"] + df["Parch"] + 1 # +1 counts oneself
# Solo vs accompanied: solo survival rate significantly lower
df["IsAlone"] = (df["FamilySize"] == 1).astype(int)④ Numerical scaling
Logistic regression (and any model optimized by gradient descent) is sensitive to feature scale: Fare is tens to hundreds, Age is teens — a 10× scale difference distorts the optimizer and regularization. Tree models don't care about scaling, but uniform scaling does no harm and makes the pipeline more general:
python
from sklearn.preprocessing import StandardScaler
# ⚠️ For demonstration only. The correct approach: put in a Pipeline, let scaler only fit on the training set
scaler = StandardScaler()
df["Age_scaled"] = scaler.fit_transform(df[["Age"]])
df["Fare_scaled"] = scaler.fit_transform(df[["Fare"]])More techniques in feature engineering
This article only used imputation, encoding, scaling, and two manual features. A more systematic framework — binning, log transforms, time features, target encoding — is in Feature Engineering. The flip side of feature engineering is feature selection (removing redundancy, preventing overfitting), see Overfitting and Regularization.
Finally, assemble the modeling matrix:
python
# Feature column list (drop labels, IDs, and original columns that were already processed)
drop_cols = ["PassengerId", "Name", "Ticket", "Survived",
"Age", "Fare", "SibSp", "Parch"]
X = df.drop(columns=drop_cols)
y = df["Survived"]
print("Feature matrix:", X.shape) # (891, 20)
print(X.columns.tolist())Step 3: Splitting Train / Val / Test
This step is the core anti-cheat measure. Three-layer splitting has different responsibilities:
Full data (891)
├── Training set train (70%) ──── model learns parameters here
├── Validation set val (15%) ───── here you select models / tune hyperparams (can look repeatedly)
└── Test set test (15%) ────────── only touched once at the end (simulates "unseen future data")python
from sklearn.model_selection import train_test_split
# Step 1: split off test set (only allowed to touch once at the end)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.15, random_state=42, stratify=y
)
# Step 2: split off validation set from the remainder
X_train, X_val, y_train, y_val = train_test_split(
X_train, y_train, test_size=0.15 / 0.85, random_state=42, stratify=y_train
)
print(f"train {X_train.shape[0]} / val {X_val.shape[0]} / test {X_test.shape[0]}")Two key points:
stratify=y(stratified sampling): ensures the survival ratio in all three sets matches the original data (~38%). Without stratification, random splitting might make the training set's "female passenger ratio skewed," distorting evaluation. Stratification is mandatory under class imbalance.- Split test first, then val: two-layer splitting must be done serially, rather than cutting the data into three parts and "arbitrarily designating one as test." The
0.15 / 0.85above splits 15% from the remaining 85% — the ratio happens to be 15% of the total.
Why keep a separate test set? Isn't validation enough?
Because you repeatedly tune on the validation set, the validation set score gets "contaminated" by the tuning behavior itself — tune enough, and the good performance on the validation set might just be lucky overfitting. The test set is the only "judge" completely immune to your will. Even with small data, insist on this discipline; it's the defense against data leakage and evaluation traps.
Step 4: Baseline Models
Baselines have two levels: weak baseline (proving "the model learned something") and strong baseline (proving "the new model is worth deploying").
① Majority-class baseline: learn nothing
python
from sklearn.dummy import DummyClassifier
from sklearn.metrics import accuracy_score, roc_auc_score
dummy = DummyClassifier(strategy="most_frequent")
dummy.fit(X_train, y_train)
y_pred = dummy.predict(X_val)
print("Accuracy:", accuracy_score(y_val, y_pred)) # ≈ 0.616 (all predict died)
# Note: majority class has no "positive probability," AUC is meaningless (constantly 0.5)All predict died, accuracy 61.6% — this is the lower bound every model must exceed.
② Logistic regression: the first real model
python
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
# Pipeline: scaler only fits on the training set; test/validation sets auto-use the training set's parameters
pipe_lr = Pipeline([
("scaler", StandardScaler()),
("lr", LogisticRegression(max_iter=1000, random_state=42)),
])
pipe_lr.fit(X_train, y_train)
print("Val accuracy:", pipe_lr.score(X_val, y_val))Pipeline isn't a bonus — it forces the discipline of "fit on training, transform on others" into the structure, preventing the most common leakage mistake from manual handling.
③ See the model's "thoughts": logistic regression is interpretable
Logistic regression's value isn't just the score; it directly tells you the direction and intensity of each feature:
python
import numpy as np
coef = pd.Series(pipe_lr.named_steps["lr"].coef_[0], index=X_train.columns)
print(coef.sort_values(ascending=False).to_string())
# Typical output (illustrative):
# Title_Miss 2.74 ← Title "Miss" (young female) strongly points to survival
# Sex_male -2.58 ← Male strongly points to death
# Title_Mr -1.96 ← Title "Mr" (adult male) points to death
# Pclass_3 -0.98 ← Third class points to death
# ...A positive coefficient means "increases survival probability when this feature increases," negative means the opposite. What the model learned aligns completely with EDA observations (higher survival rate for females / first class), which itself is a validation: the model didn't learn the wrong things. More on interpretability is in the trade-offs section of What is Machine Learning.
Step 5: Iterative Improvement
After the baseline "passes," it's time for more complex models. The order is always: simple baseline → medium model → tuning; jumping straight to the most complex model and then tuning is the beginner's most common waste.
① Random forest: tree models enter the field
python
from sklearn.ensemble import RandomForestClassifier
pipe_rf = Pipeline([
("rf", RandomForestClassifier(n_estimators=200, random_state=42)),
])
pipe_rf.fit(X_train, y_train)
print("Random Forest val AUC:",
roc_auc_score(y_val, pipe_rf.predict_proba(X_val)[:, 1]))Random forest doesn't need scaling, captures non-linearity (e.g., "age × cabin" interactions), and is usually a strong baseline for tabular data. See Tree Models and Ensemble Learning for tree model family mechanics.
② Feature importance: what the tree model tells you it used
python
importances = pd.Series(
pipe_rf.named_steps["rf"].feature_importances_,
index=X_train.columns,
).sort_values(ascending=False)
importances.plot.barh(figsize=(8, 8))
plt.title("Random Forest Feature Importance")
plt.show()Age, Fare, Sex_male, FamilySize usually rank at the top. If a feature you consider important has importance near 0, it's worth going back and thinking: did the feature not get done right, or was the signal never there?
③ Hyperparameter tuning: search on the validation set
Random forest's tunable hyperparams are mainly: tree count n_estimators, max tree depth max_depth, minimum samples for splitting min_samples_split. Use GridSearchCV for cross-validation search:
python
from sklearn.model_selection import GridSearchCV
param_grid = {
"rf__n_estimators": [100, 200, 400],
"rf__max_depth": [3, 5, None],
"rf__min_samples_split": [2, 5, 10],
}
gs = GridSearchCV(
pipe_rf,
param_grid,
cv=5, # 5-fold cross-validation
scoring="roc_auc", # search by primary metric
n_jobs=-1,
verbose=1,
)
gs.fit(X_train, y_train)
print("Best params:", gs.best_params_)
print("Best CV AUC:", gs.best_score_)
print("Val AUC:", roc_auc_score(y_val, gs.best_estimator_.predict_proba(X_val)[:, 1]))What cross-validation and grid search are
GridSearchCV does K-fold CV internally on the training set: cuts the training set into 5 parts, uses 4 for train, 1 for validation, averages 5 results — selecting parameters this way doesn't depend on one split's luck. Its found best_estimator_ is then confirmed on independent validation set. Note: tuning only looks at CV scores and validation set, never the test set. More complete practices for parameter search are in Model Evaluation and Validation and Common Pitfalls and Anti-Patterns.
④ Iteration records: experiment log
Iteration must leave records, or three days later you'll forget "which version gave 0.81":
python
import json, datetime
log = {
"time": datetime.datetime.now().isoformat(),
"model": "RandomForest + GridSearch",
"best_params": gs.best_params_,
"cv_auc": float(gs.best_score_),
"val_auc": float(roc_auc_score(y_val, gs.best_estimator_.predict_proba(X_val)[:, 1])),
}
print(json.dumps(log, indent=2, ensure_ascii=False))Step 6: Evaluation and Error Analysis
Model selected; now look seriously at where it errs. Error analysis is the key step from score to understanding.
① Confusion matrix: what errors look like
python
from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay
y_val_prob = gs.best_estimator_.predict_proba(X_val)[:, 1]
y_val_pred = (y_val_prob >= 0.5).astype(int)
cm = confusion_matrix(y_val, y_val_pred)
disp = ConfusionMatrixDisplay(cm, display_labels=["Died", "Survived"])
disp.plot()
plt.title("Val Confusion Matrix (threshold 0.5)")
plt.show()
# Result interpretation:
# [[? ?] Row=true (Died/Survived), Column=predicted
# [? ?]]
# Top-left = correctly predicted died (TN), bottom-right = correctly predicted survived (TP)
# Top-right = mispredicted survived as died (FN, miss)
# Bottom-left = mispredicted died as survived (FP, false alarm)② Look at multiple metrics simultaneously: single metrics lie
| Metric | Formula | Question It Answers |
|---|---|---|
| Accuracy | (TP+TN)/all | Proportion correctly predicted overall |
| Precision | TP/(TP+FP) | Of those predicted "survived," the proportion truly survived |
| Recall | TP/(TP+FN) | Of truly survived, the proportion that were found |
| F1 | 2·P·R/(P+R) | Harmonic mean of precision and recall |
| AUC | Area under curve | Probability that, randomly picking a survived vs. died passenger, the model gives the survived one a higher score |
python
from sklearn.metrics import classification_report, precision_recall_curve
print(classification_report(y_val, y_val_pred, target_names=["Died", "Survived"]))Why you can't just look at accuracy
Titanic data is 62% died. A "all predict died" model already has 62% accuracy, and recall = 0 — it didn't find a single survivor. When classes are imbalanced, accuracy is an extremely misleading metric. First clarify what the business cares about: if the goal is "never miss a survivor," optimize recall; if the goal is "when predicting survival, make sure it's true," optimize precision. Full metric discussion is in Model Evaluation and Validation.
③ Threshold analysis: 0.5 isn't a golden rule
Classification models actually output probabilities; 0.5 is just the default cutoff line. Different businesses need different thresholds:
python
precisions, recalls, thresholds = precision_recall_curve(y_val, y_val_prob)
# Plot precision-recall varying with threshold
plt.plot(thresholds, precisions[:-1], label="Precision")
plt.plot(thresholds, recalls[:-1], label="Recall")
plt.xlabel("Threshold"); plt.legend(); plt.show()Want to "find as many survivors as possible" → lower the threshold (recall↑, precision↓); want to "only predict survival when it's reliable" → raise the threshold. Threshold is a business decision, not a model parameter — it can only be selected on the validation set, not the test set.
④ Look at specific error samples
The most insightful step in error analysis: print out predicted-wrong samples and look at them one by one to see why they're wrong.
python
# Put the validation set back together to look at error samples
val_df = X_val.copy()
val_df["true"] = y_val
val_df["predicted_prob"] = y_val_prob
val_df["predicted"] = y_val_pred
errors = val_df[val_df["true"] != val_df["predicted"]]
print("Number of errors:", len(errors))
print(errors.head(15).to_string())Looking at these error samples one by one, common findings are:
- The model systematically mispredicts a certain type of person (e.g., misclassifies all young males as survivors) → this feature combination still lacks info; continue feature engineering.
- Individual samples are extremely hard (e.g., "traveling alone, male, third class" but truly survived) → these are noise; no need to tune the model for them.
- This determines whether the next step is add features, switch models, or stop here.
Step 7: Result Presentation and Visualization
At the end of the project, present results honestly. A qualified presentation = one metrics table + one key chart + one reproducibility note.
① Metrics summary table
| Model | Train CV AUC | Val AUC | Val Accuracy | Notes |
|---|---|---|---|---|
| Majority-class baseline | — | 0.500 | 61.6% | All predict died |
| Logistic regression | 0.85 | 0.86 | 79.5% | Interpretable |
| Random forest | 0.88 | 0.87 | 81.1% | Needs tuning |
| Random forest (tuned) | 0.89 | 0.88 | 82.0% | Final selection |
② Honesty principles for presentations
- Report val scores, report the reason for final selection (why random forest over logistic regression: slightly higher AUC and more stable, at the cost of interpretability).
- Explicitly state limitations: only 891 samples, test set only 134 people, AUC's confidence interval is wide.
- Never write test-set scores into the "exploration process" — test-set scores only appear once in the final report.
③ Visualization trio
python
# (a) ROC curve: TPR vs FPR at all thresholds, AUC is the area under the curve
from sklearn.metrics import roc_curve
fpr, tpr, _ = roc_curve(y_val, y_val_prob)
plt.plot(fpr, tpr, label=f"AUC = {roc_auc_score(y_val, y_val_prob):.3f}")
plt.plot([0, 1], [0, 1], "--", color="gray") # random guessing baseline
plt.xlabel("FPR"); plt.ylabel("TPR"); plt.legend(); plt.show()
# (b) Feature importance bar chart (top 10)
# (c) Learning curve: samples vs score, judge "underfit or data shortage"
from sklearn.model_selection import learning_curve
sizes, train_scores, val_scores = learning_curve(
gs.best_estimator_, X_train, y_train,
cv=5, scoring="roc_auc", train_sizes=np.linspace(0.1, 1.0, 5),
)
plt.plot(sizes, train_scores.mean(1), label="train AUC")
plt.plot(sizes, val_scores.mean(1), label="CV AUC")
plt.xlabel("Training samples"); plt.legend(); plt.show()If the learning curve shows "high train score, low CV score, converging as sample size increases" → the problem is not enough data, not insufficient model; the next step is to find more data, not tune (this judgment standard is from Overfitting and Regularization).
④ Touch the test set one final time
python
from sklearn.metrics import roc_auc_score
test_prob = gs.best_estimator_.predict_proba(X_test)[:, 1]
print("Test AUC (final score):", roc_auc_score(y_test, test_prob))After running this line, record the score in the experiment log, and don't go touch the test set again. If you can't resist going back to change the model and re-measure — you've already been using the test set as a validation set, and the score will be inflated.
IV. Complete Code Repository Structure
A maintainable ML project's file organization follows "clear entry point, separated responsibilities, centralized parameters". Recommended directory structure:
titanic/
├── README.md # Project intro: goal, data source, how to reproduce, results
├── requirements.txt # Dependency list (generated by pip freeze)
├── config.py # All hyperparams, paths, random seeds centralized here
├── data/
│ ├── train.csv # raw data (read-only, never modify in place)
│ ├── processed/ # cleaned data (generated by scripts)
│ └── test.csv # Kaggle test data (optional)
├── notebooks/
│ ├── 01-eda.ipynb # exploratory analysis (exploration process, not in the reproducible path)
│ └── 02-experiments.ipynb # experiment log
├── src/
│ ├── __init__.py
│ ├── load_data.py # load raw data
│ ├── features.py # cleaning and feature engineering (the sole feature implementation)
│ ├── models.py # model definition, Pipeline, tuning logic
│ └── evaluate.py # metrics, confusion matrix, ROC, etc.
├── models/
│ └── best_model.joblib # serialized saved model
├── reports/
│ ├── figures/ # all chart outputs
│ └── results.md # experiment results summary table
└── train.py # main entry: load → features → train → evaluate → save| File/Directory | Responsibility | Why placed here |
|---|---|---|
config.py | Centralize random seeds, paths, hyperparams | Change params without flipping through code; experiments are reproducible |
src/features.py | The sole feature implementation | Training and online share the same feature code, preventing "train/online inconsistency" |
data/ read-only | Raw data never modified in place | Ensure every run starts from the same data |
notebooks/ | Exploration and analysis | Process content stays in notebooks; reusable logic sinks to src/ |
train.py | One-click full pipeline | Going from "runnable notebook" to "reproducible script" is a qualitative leap |
A golden test for reproducible experiments: delete the notebooks and models directories, keeping only requirements.txt, src/, config.py, train.py — can someone else reproduce your score on a new machine? If the answer is "yes," your engineering meets the bar.
Principles for migrating from notebooks to scripts
Notebooks suit exploration, but don't suit reproducibility (execution order is messy, implicit state is abundant). Mature practice: after tuning logic in a notebook, extract reusable parts into functions under src/; notebooks only "call and display results." This is exactly the practical landing of the "pipeline vs script" discussion in Data Engineering.
V. Common Pitfalls
The project is small but complete. Below are the most common pitfalls in this workflow, and the hardest to discover once hit:
Pitfall 1: Data Leakage (leakage)
Leakage points encountered in this workflow, checked one by one:
| Leakage Point | Symptom | Correct Practice |
|---|---|---|
| Fitting scaler / imputer on full data | Inflated val scores, crashes on deployment | Put in Pipeline or manually "fit training set first, then transform others" |
| Using full-data statistics in feature engineering before splitting (e.g., full-data median) | Same as above | Compute all statistics after splitting |
| Repeatedly using the test set for tuning | Final score higher than "true level" | Test set only touched once |
High-cardinality encoding for ID-like columns (Name / PassengerId) | Model "memorizes" ID-to-label mapping | Drop IDs directly |
"Split first, process later" is an iron rule. For a more complete list (including temporal leakage, train-serve inconsistency, and 25 more), see Common Pitfalls and Anti-Patterns.
Pitfall 2: Overfitting
Overfitting signals and handling for this project:
- Signal: train AUC 0.97, val 0.88 → the gap between train and val indicates the model is memorizing the training set.
- Handling: limit
max_depthandmin_samples_splitfor tree models, add regularization (reduceC), use CV for parameter selection. Mechanics are in Overfitting and Regularization. - Ultimate measure: this project has only 891 samples; any model easily overfits — sample size is a hard constraint — more features, stronger models, higher overfitting risk; be restrained with feature selection.
Pitfall 3: Misused Metrics
- Only looking at accuracy → fooled by "all predict died" on 38/62 imbalanced data.
- Treating AUC as the sole truth → AUC cares about ranking, not thresholds; the business cares about "precision at threshold 0.7+."
- Selecting metrics on the test set → metric selection itself is a form of tuning; must be decided in the validation phase.
Pitfall 4: Data Drift
After the model is deployed, the distribution of real data changes (e.g., changing seasons, new user groups). In a practice project, this is a next-step concern, but now is the time to write "training data limitations" in the report — patterns learned from 891 passengers in 1912 may not hold for passengers in 2024. Data's full lifecycle management is in Data Engineering.
VI. Advanced Directions
The Titanic project is the "minimum complete set" of the workflow; extending it into a real portfolio-worthy project can go in four directions:
- Deepen evaluation: add cross-validation (replace single hold-out), confidence intervals, calibration curves, upgrade evaluation from "reporting a number" to "having confidence in the score." Practical guide is in Evaluation in Practice.
- Switch problem and data: do regression (housing prices), multi-class (flower species), or switch to a larger, dirtier real dataset, exercising data cleaning skills. Dataset list is in Datasets and Tools Archive.
- More systematic training workflow: standardize the workflow with a more complete tutorial-style document, forming your own scaffold. See ML Tutorial and Evaluation in Practice.
- Make it a complete portfolio piece: add experiment tracking (compare multiple runs), model serialization and deployment (wrap a prediction interface with FastAPI), reproducible Docker environment, and extend it into a project you can clearly talk about in interviews/job-hunting. The complete extension roadmap is in Advanced Practice Projects.
bash
# Serialize the final model, preparing for deployment
import joblib
joblib.dump(gs.best_estimator_, "models/best_model.joblib")
# Load and predict at deployment time
model = joblib.load("models/best_model.joblib")
prob = model.predict_proba(new_passenger_features)[0, 1] # returns survival probabilityVII. Further Reading
- What is Machine Learning — the theoretical overview of four stages
- Overall Architecture Anatomy — scaling this small project up to a production-grade system
- Model Evaluation and Validation — complete theory of metrics, cross-validation, confidence intervals
- Feature Engineering — systematic methods for encoding, scaling, binning
- Overfitting and Regularization — math and practice of bias-variance trade-off
- Tree Models and Ensemble Learning — mechanics of random forest and gradient boosting
- Common Pitfalls and Anti-Patterns — the full version of Section V in this article (25 pitfalls)
- Advanced Practice Projects — extending a small project into a portfolio
- Datasets and Tools Archive — where to find data for the next practice project
References
- scikit-learn user guide: Model selection and evaluation —
train_test_split,GridSearchCV, evaluation metrics official docs - scikit-learn user guide: Pipelines and composite estimators — Pipeline and preprocessor official docs
- scikit-learn user guide: Preprocessing data — encoding and scaling official docs
- pandas user guide —
groupby,fillna,get_dummiesofficial docs - Kaggle: Titanic - Machine Learning from Disaster — source of this article's dataset, the first beginner competition question
- matplotlib docs — official docs for visualization code
- sklearn.metrics module docs —
roc_auc_score,confusion_matrix,learning_curve, and all other evaluation tools