Skip to content

Building an ML Project from Scratch

Quick overview Using the classic Titanic survival prediction as a small project, run through the complete workflow from data loading, EDA, feature engineering, baseline model to iterative improvement and error analysis: each step gives runnable code, explanation, repo directory structure, common pitfalls, and advanced directions.

Building an ML Project from Scratch ​

Reading ten tutorials is worth less than running a complete project end-to-end. This article takes you through every stage of a real project using the classic, minimal "Titanic survival prediction" — not "calling APIs to get results," but making clear why every step is done the way it is.

Many beginners, having learned many algorithms, still don't know what an ML project actually looks like: how are files organized? what's the order of steps? which parts can be skipped, which absolutely can't? This article walks the complete project workflow from start to finish using one thread (Titanic survival prediction). What you get isn't "0.81 score on this dataset," but a process skeleton that can be applied to any tabular classification problem.

A complete project has seven steps; first, the overall map:

Raw data
   │
   ▼
① Data loading and exploration (EDA) ── Answer: "what does the data look like, what problems?"
   ▼
② Data cleaning and feature engineering ── Answer: "which columns enter the model, how to turn them into numbers?"
   ▼
③ Split train / val / test ── Answer: "how to guarantee the evaluation is honest?"
   ▼
④ Baseline model ────────── Answer: "what's the best a simple solution can do?"
   ▼
⑤ Iterative improvement ─── Answer: "how much more can a more complex model add?"
   ▼
⑥ Evaluation and error analysis ── Answer: "where does the model err, and why?"
   ▼
⑦ Result presentation and visualization ─ Answer: "how to explain conclusions clearly?"

Throughout the process, we deliberately follow one discipline: the test set is only touched once, at the end. This discipline runs through the entire article, and it's the defense against #1 in Common Pitfalls: "data leakage."

Prerequisites

This article assumes you understand the basic concepts: what supervised learning is (Supervised Learning), how models are evaluated (Model Evaluation and Validation), what feature engineering is (Feature Engineering). If you need to fill basics, read What is Machine Learning and Overall Architecture Anatomy first, then come back.

I. Project Selection: A Simple but Real Problem ​

The first question: what to choose for a practice project? Many people's first project is "predicting stock prices with LSTM," which almost certainly fails — not because the technology is hard, but because the problem itself is unsolvable. There are three criteria for choosing a project, all required.

CriterionWhyConsequence of Violating
Has labels (has ground truth)You need to know "the correct answer" to evaluate whether the model learnedWithout labels, there's nothing to evaluate; you're just feeling good about yourself
Small scale (hundreds to thousands of samples)Each iteration runs in minutes, so you can focus on process and conceptsBig data projects spend half the time tuning Spark; you don't learn modeling
Interpretable (human-understandable feature meanings)You can use common sense to judge whether features are reasonable; error analysis works"Why it's wrong" for images/text needs additional tools; beginners can't get started

By these criteria, here's a comparison of three classic practice projects:

ProjectLabelsScaleInterpretabilitySuitability
Titanic survival predictionYes (Survived)891 training samples★★★ All features are plain language (age, cabin class, gender)★★★ Beginner's first choice
Boston / California housing price predictionYes (continuous values)~20k★★☆ Features have real-world meaning★★★ Best for regression practice
Spam email classificationYes (spam/ham)5k+ emails★☆☆ Features are word vectors; need text processing★★☆ Slightly advanced

This article chooses Titanic survival prediction, for three reasons:

  1. Clear labels: whether each passenger survived is a determined historical fact, no "label subjectivity" issue.
  2. Rich and interpretable features: age, gender, cabin class (Pclass), fare, boarding port — every column can be explained by common sense, suitable for error analysis.
  3. Small scale: 891 training samples; any model runs in milliseconds, so you can iterate freely.

What the Dataset Looks Like ​

The Titanic dataset (from the Kaggle same-named entry competition) in its raw form:

ColumnMeaningTypeHas Missing?
PassengerIdPassenger IDintNo
SurvivedWhether survived (1=yes, 0=no)int (label)No
PclassCabin class (1/2/3)intNo
NameNamestrNo
SexGender (male/female)strNo
AgeAgefloatYes (177 missing)
SibSpNumber of siblings/spouses aboardintNo
ParchNumber of parents/children aboardintNo
TicketTicket numberstrNo
FareFarefloatNo
CabinCabin numberstrYes (687 missing!)
EmbarkedBoarding port (C/Q/S)strYes (2 missing)

Note: these 891 rows are the training set for the Kaggle competition. In real projects, data is often scattered across different systems and formats, and loading and cleaning take up half the work — precisely because of this, the "clean dataset" training style is suitable for learning the workflow. For dataset search and evaluation methods, see Datasets and Tools Archive.

Define "success" before writing code

When translating a problem into a machine learning task, you must answer three questions:

  1. Task type: survival prediction is binary classification (survived / died).
  2. Evaluation metric: in this dataset, ~38% survived — this is class imbalanced. Accuracy is misleading (see Misused Metrics), so we use ROC-AUC as the primary metric, supplemented by precision / recall.
  3. Baseline: a model that "predicts all died" has ~62% accuracy. Any model of yours must significantly exceed this number, otherwise it means nothing was learned.

Project success is defined before code is written — this is the basic skill of Model Evaluation and Validation.

II. Environment Setup ​

1. Version Requirements ​

This article's code is based on the following versions (2024 stable versions, backwards compatible with older ones):

SoftwareVersionPurpose
Python3.10+Language itself
pandas2.xData loading and processing
scikit-learn1.3+Modeling, splitting, evaluation, tuning
matplotlib3.xVisualization
jupyterlab4.xInteractive exploration

2. Create a Virtual Environment (venv) ​

Never install dependencies into system Python. A virtual environment is an isolation container for project dependencies, ensuring no cross-contamination when switching machines / projects:

bash
# Enter project root directory
mkdir titanic && cd titanic

# Create virtual environment (activation commands differ between Windows and macOS/Linux)
python -m venv .venv
source .venv/bin/activate      # macOS / Linux
# .venv\Scripts\activate       # Windows (PowerShell)
# .venv\Scripts\activate.bat   # Windows (CMD)

# Install dependencies (one-line install)
pip install pandas scikit-learn matplotlib jupyterlab

# Freeze the dependency list, ensuring others can reproduce with one click
pip freeze > requirements.txt

Why use venv instead of global install

  • Different projects may depend on different versions of numpy/sklearn; global install forces you to "upgrade/downgrade back and forth."
  • requirements.txt generated by pip freeze is the first step to reproducible experiments — others get the exact same environment with pip install -r requirements.txt.
  • In team collaboration, reproducible environment = reproducible experiment results.

3. Start Jupyter and Verify ​

bash
jupyter lab

Create a new notebook, first run the "environment smoke test":

python
import pandas as pd, numpy as np, sklearn, matplotlib
print("pandas:", pd.__version__)
print("sklearn:", sklearn.__version__)
print("numpy:", np.__version__)

# One-line smoke test: confirm sklearn's core components all work
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
print("Environment OK")

When version numbers output, the environment is ready. Recommend putting EDA in notebooks, putting reusable logic in .py scripts (directory structure in Section IV) — notebooks suit exploration, scripts suit reuse and testing.

III. Complete Workflow Step by Step ​

Step 1: Data Loading and Exploration (EDA) ​

EDA (Exploratory Data Analysis) isn't about "looking at data"; it's about answering three questions: what problems does the data have (missing, types), what do individual features look like (distributions), and what's the relationship between features and labels (where is the signal)?

python
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt

# Load data (downloaded train.csv from Kaggle, also available from this site's dataset archive)
df = pd.read_csv("data/train.csv")

# ── ① Check shape and head ────────────────────────────────
print("Shape:", df.shape)              # (891, 12) → 891 passengers, 12 columns
df.head()

# ── ② Check per-column types and missing ──────────────────
df.info()

# ── ③ Check numeric column stats (count/mean/std/min/25%/50%/75%/max)──
df.describe()

# ── ④ Check missing values at a glance ────────────────────
df.isnull().sum()

df.info() immediately exposes three problems: Age is missing 177, Cabin is missing 687, Embarked is missing 2. df.describe() exposes a fourth problem: Fare's standard deviation is 49.7, mean 32 while median 14 — indicating severe right skew (a few first-class tickets with extremely high prices pulled the mean up).

Next, explore the target distribution:

python
# Survival ratio: 38.4%, meaning the majority class is "died" (62%)
print(df["Survived"].value_counts(normalize=True))
# 0    0.616162
# 1    0.383838

Then the relationship between features and labels — this is where EDA is most information-dense:

python
# Gender and survival: female survival rate 74.2%, male 18.9% — the strongest signal
print(df.groupby("Sex")["Survived"].mean())
# female    0.742038
# male      0.188908

# Cabin class and survival: first class 63.0%, second 47.3%, third 24.2%
print(df.groupby("Pclass")["Survived"].mean())
# 1    0.629630
# 2    0.472826
# 3    0.242363

# Age distribution (look at non-missing part for now)
df["Age"].hist(bins=30)
plt.title("Age Distribution")
plt.show()

Write down the EDA conclusions (not just glance at them); these are the basis for subsequent feature engineering:

FindingMeaningSubsequent Action
Survived ~38% survivalClass imbalanceUse AUC/recall, don't look at raw accuracy
Sex strongly correlated with survival (female 74% vs male 19%)Core signalMust encode into the model
Pclass strongly correlated with survival (decreasing 1→3)Core signalKeep, consider making it a categorical feature
Age missing 177, Cabin missing 687Severe missingnessNeeds imputation or exclusion (next step)
Fare severely right-skewedSkewed distributionLog transform or binning
Name / Ticket / PassengerIdHigh cardinality / identifiersLittle direct predictive power, but can extract features

Signs of a good EDA

A good EDA produces a findings list, where each finding leads to a subsequent decision. If you do EDA just by "printing a dozen tables and moving on," you haven't entered the right state. EDA is essentially hypothesis testing: you come with hypotheses about "what factors affect survival," and use data to validate or refute them.

Step 2: Data Cleaning and Feature Engineering ​

EDA findings need to be implemented in code. General principle: the output of cleaning and feature engineering is a unified X (feature matrix) + y (label), and all our models consume this pair.

① Missing value handling

ColumnMissing rateHandling methodReason
Age20%Fill with median (or group-fill by sex/cabin class)Age is right-skewed; median is more robust than mean
Embarked0.2%Fill with mode (most frequent S)Only 2 rows, negligible impact
Cabin77%Don't fill directly; convert to "has cabin number" binary featureMissingness itself may contain info (low survival rate for those without cabin records); filling with fake values introduces noise
python
# Age: group-fill is more fine-grained than global-fill (median by sex/cabin differs)
df["Age"] = df.groupby(["Sex", "Pclass"])["Age"].transform(
    lambda s: s.fillna(s.median())
)

# Embarked: mode fill
df["Embarked"] = df["Embarked"].fillna(df["Embarked"].mode()[0])

# Cabin: convert to binary feature "has cabin record" (higher information density than the cabin itself)
df["HasCabin"] = df["Cabin"].notna().astype(int)
df = df.drop(columns=["Cabin"])

The "leakage" red line for missing value handling

All quantities learned from data (median, mode, mean, min/max, regression coefficients) must only be computed from the training set, then applied to validation/test sets. Doing fillna(df["Age"].median()) on the full dataset is fine for practice, but in real projects, this is a form of data leakage — test-set information flows into the training set. The rigorous way is to put it in a Pipeline (see below), letting sklearn automatically "fit on the training set, transform on other sets."

② Categorical encoding

Sex and Embarked are strings; models only accept numbers:

python
# One-Hot encoding: Embarked has 3 classes → 3 0/1 columns (sklearn auto-drops the first to avoid collinearity)
df = pd.get_dummies(df, columns=["Sex", "Embarked"], drop_first=True)

Sex_male becomes 1/0, Embarked_Q, Embarked_S each 1/0. Note: Pclass has numeric form 1/2/3, but it's a rank (ordinal category) not a continuous quantity; treating it as a categorical feature is more stable — we use pd.get_dummies to expand Pclass into 3 columns, or directly keep the original value and let the tree model handle it (trees don't assume linear relationships of numbers).

③ Manual feature engineering: letting domain knowledge enter the model

The title (Mr/Mrs/Miss/Master) in Name is a classic feature engineering example — it strongly correlates with age and gender, especially useful for distinguishing age groups with missing values:

python
df["Title"] = df["Name"].str.extract(r" ([A-Za-z]+)\.", expand=False)
# Merge rare titles, reduce categories
df["Title"] = df["Title"].replace(
    ["Lady","Countess","Capt","Col","Don","Dr","Major","Rev","Sir","Jonkheer","Dona"],
    "Rare"
)
df["Title"] = df["Title"].replace(["Mlle","Ms"], "Miss")
df["Title"] = df["Title"].replace(["Mme"], "Mrs")
df = pd.get_dummies(df, columns=["Title"], drop_first=True)

Then construct a family size feature — "whole families together" is a significant factor in survival:

python
df["FamilySize"] = df["SibSp"] + df["Parch"] + 1   # +1 counts oneself
# Solo vs accompanied: solo survival rate significantly lower
df["IsAlone"] = (df["FamilySize"] == 1).astype(int)

④ Numerical scaling

Logistic regression (and any model optimized by gradient descent) is sensitive to feature scale: Fare is tens to hundreds, Age is teens — a 10× scale difference distorts the optimizer and regularization. Tree models don't care about scaling, but uniform scaling does no harm and makes the pipeline more general:

python
from sklearn.preprocessing import StandardScaler
# ⚠️ For demonstration only. The correct approach: put in a Pipeline, let scaler only fit on the training set
scaler = StandardScaler()
df["Age_scaled"]   = scaler.fit_transform(df[["Age"]])
df["Fare_scaled"]  = scaler.fit_transform(df[["Fare"]])

More techniques in feature engineering

This article only used imputation, encoding, scaling, and two manual features. A more systematic framework — binning, log transforms, time features, target encoding — is in Feature Engineering. The flip side of feature engineering is feature selection (removing redundancy, preventing overfitting), see Overfitting and Regularization.

Finally, assemble the modeling matrix:

python
# Feature column list (drop labels, IDs, and original columns that were already processed)
drop_cols = ["PassengerId", "Name", "Ticket", "Survived",
             "Age", "Fare", "SibSp", "Parch"]
X = df.drop(columns=drop_cols)
y = df["Survived"]
print("Feature matrix:", X.shape)   # (891, 20)
print(X.columns.tolist())

Step 3: Splitting Train / Val / Test ​

This step is the core anti-cheat measure. Three-layer splitting has different responsibilities:

Full data (891)
 ├── Training set train (70%) ──── model learns parameters here
 ├── Validation set val (15%) ───── here you select models / tune hyperparams (can look repeatedly)
 └── Test set test (15%) ────────── only touched once at the end (simulates "unseen future data")
python
from sklearn.model_selection import train_test_split

# Step 1: split off test set (only allowed to touch once at the end)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.15, random_state=42, stratify=y
)

# Step 2: split off validation set from the remainder
X_train, X_val, y_train, y_val = train_test_split(
    X_train, y_train, test_size=0.15 / 0.85, random_state=42, stratify=y_train
)

print(f"train {X_train.shape[0]} / val {X_val.shape[0]} / test {X_test.shape[0]}")

Two key points:

  1. stratify=y (stratified sampling): ensures the survival ratio in all three sets matches the original data (~38%). Without stratification, random splitting might make the training set's "female passenger ratio skewed," distorting evaluation. Stratification is mandatory under class imbalance.
  2. Split test first, then val: two-layer splitting must be done serially, rather than cutting the data into three parts and "arbitrarily designating one as test." The 0.15 / 0.85 above splits 15% from the remaining 85% — the ratio happens to be 15% of the total.

Why keep a separate test set? Isn't validation enough?

Because you repeatedly tune on the validation set, the validation set score gets "contaminated" by the tuning behavior itself — tune enough, and the good performance on the validation set might just be lucky overfitting. The test set is the only "judge" completely immune to your will. Even with small data, insist on this discipline; it's the defense against data leakage and evaluation traps.

Step 4: Baseline Models ​

Baselines have two levels: weak baseline (proving "the model learned something") and strong baseline (proving "the new model is worth deploying").

① Majority-class baseline: learn nothing

python
from sklearn.dummy import DummyClassifier
from sklearn.metrics import accuracy_score, roc_auc_score

dummy = DummyClassifier(strategy="most_frequent")
dummy.fit(X_train, y_train)

y_pred = dummy.predict(X_val)
print("Accuracy:", accuracy_score(y_val, y_pred))   # ≈ 0.616 (all predict died)
# Note: majority class has no "positive probability," AUC is meaningless (constantly 0.5)

All predict died, accuracy 61.6% — this is the lower bound every model must exceed.

② Logistic regression: the first real model

python
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

# Pipeline: scaler only fits on the training set; test/validation sets auto-use the training set's parameters
pipe_lr = Pipeline([
    ("scaler", StandardScaler()),
    ("lr", LogisticRegression(max_iter=1000, random_state=42)),
])

pipe_lr.fit(X_train, y_train)
print("Val accuracy:", pipe_lr.score(X_val, y_val))

Pipeline isn't a bonus — it forces the discipline of "fit on training, transform on others" into the structure, preventing the most common leakage mistake from manual handling.

③ See the model's "thoughts": logistic regression is interpretable

Logistic regression's value isn't just the score; it directly tells you the direction and intensity of each feature:

python
import numpy as np

coef = pd.Series(pipe_lr.named_steps["lr"].coef_[0], index=X_train.columns)
print(coef.sort_values(ascending=False).to_string())

# Typical output (illustrative):
# Title_Miss     2.74    ← Title "Miss" (young female) strongly points to survival
# Sex_male      -2.58    ← Male strongly points to death
# Title_Mr      -1.96    ← Title "Mr" (adult male) points to death
# Pclass_3      -0.98    ← Third class points to death
# ...

A positive coefficient means "increases survival probability when this feature increases," negative means the opposite. What the model learned aligns completely with EDA observations (higher survival rate for females / first class), which itself is a validation: the model didn't learn the wrong things. More on interpretability is in the trade-offs section of What is Machine Learning.

Step 5: Iterative Improvement ​

After the baseline "passes," it's time for more complex models. The order is always: simple baseline → medium model → tuning; jumping straight to the most complex model and then tuning is the beginner's most common waste.

① Random forest: tree models enter the field

python
from sklearn.ensemble import RandomForestClassifier

pipe_rf = Pipeline([
    ("rf", RandomForestClassifier(n_estimators=200, random_state=42)),
])
pipe_rf.fit(X_train, y_train)
print("Random Forest val AUC:",
      roc_auc_score(y_val, pipe_rf.predict_proba(X_val)[:, 1]))

Random forest doesn't need scaling, captures non-linearity (e.g., "age × cabin" interactions), and is usually a strong baseline for tabular data. See Tree Models and Ensemble Learning for tree model family mechanics.

② Feature importance: what the tree model tells you it used

python
importances = pd.Series(
    pipe_rf.named_steps["rf"].feature_importances_,
    index=X_train.columns,
).sort_values(ascending=False)

importances.plot.barh(figsize=(8, 8))
plt.title("Random Forest Feature Importance")
plt.show()

Age, Fare, Sex_male, FamilySize usually rank at the top. If a feature you consider important has importance near 0, it's worth going back and thinking: did the feature not get done right, or was the signal never there?

③ Hyperparameter tuning: search on the validation set

Random forest's tunable hyperparams are mainly: tree count n_estimators, max tree depth max_depth, minimum samples for splitting min_samples_split. Use GridSearchCV for cross-validation search:

python
from sklearn.model_selection import GridSearchCV

param_grid = {
    "rf__n_estimators": [100, 200, 400],
    "rf__max_depth": [3, 5, None],
    "rf__min_samples_split": [2, 5, 10],
}

gs = GridSearchCV(
    pipe_rf,
    param_grid,
    cv=5,                 # 5-fold cross-validation
    scoring="roc_auc",    # search by primary metric
    n_jobs=-1,
    verbose=1,
)
gs.fit(X_train, y_train)

print("Best params:", gs.best_params_)
print("Best CV AUC:", gs.best_score_)
print("Val AUC:", roc_auc_score(y_val, gs.best_estimator_.predict_proba(X_val)[:, 1]))

What cross-validation and grid search are

GridSearchCV does K-fold CV internally on the training set: cuts the training set into 5 parts, uses 4 for train, 1 for validation, averages 5 results — selecting parameters this way doesn't depend on one split's luck. Its found best_estimator_ is then confirmed on independent validation set. Note: tuning only looks at CV scores and validation set, never the test set. More complete practices for parameter search are in Model Evaluation and Validation and Common Pitfalls and Anti-Patterns.

④ Iteration records: experiment log

Iteration must leave records, or three days later you'll forget "which version gave 0.81":

python
import json, datetime

log = {
    "time": datetime.datetime.now().isoformat(),
    "model": "RandomForest + GridSearch",
    "best_params": gs.best_params_,
    "cv_auc": float(gs.best_score_),
    "val_auc": float(roc_auc_score(y_val, gs.best_estimator_.predict_proba(X_val)[:, 1])),
}
print(json.dumps(log, indent=2, ensure_ascii=False))

Step 6: Evaluation and Error Analysis ​

Model selected; now look seriously at where it errs. Error analysis is the key step from score to understanding.

① Confusion matrix: what errors look like

python
from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay

y_val_prob = gs.best_estimator_.predict_proba(X_val)[:, 1]
y_val_pred = (y_val_prob >= 0.5).astype(int)

cm = confusion_matrix(y_val, y_val_pred)
disp = ConfusionMatrixDisplay(cm, display_labels=["Died", "Survived"])
disp.plot()
plt.title("Val Confusion Matrix (threshold 0.5)")
plt.show()

# Result interpretation:
# [[?  ?]    Row=true (Died/Survived), Column=predicted
#  [?  ?]]
# Top-left = correctly predicted died (TN), bottom-right = correctly predicted survived (TP)
# Top-right = mispredicted survived as died (FN, miss)
# Bottom-left = mispredicted died as survived (FP, false alarm)

② Look at multiple metrics simultaneously: single metrics lie

MetricFormulaQuestion It Answers
Accuracy(TP+TN)/allProportion correctly predicted overall
PrecisionTP/(TP+FP)Of those predicted "survived," the proportion truly survived
RecallTP/(TP+FN)Of truly survived, the proportion that were found
F12·P·R/(P+R)Harmonic mean of precision and recall
AUCArea under curveProbability that, randomly picking a survived vs. died passenger, the model gives the survived one a higher score
python
from sklearn.metrics import classification_report, precision_recall_curve

print(classification_report(y_val, y_val_pred, target_names=["Died", "Survived"]))

Why you can't just look at accuracy

Titanic data is 62% died. A "all predict died" model already has 62% accuracy, and recall = 0 — it didn't find a single survivor. When classes are imbalanced, accuracy is an extremely misleading metric. First clarify what the business cares about: if the goal is "never miss a survivor," optimize recall; if the goal is "when predicting survival, make sure it's true," optimize precision. Full metric discussion is in Model Evaluation and Validation.

③ Threshold analysis: 0.5 isn't a golden rule

Classification models actually output probabilities; 0.5 is just the default cutoff line. Different businesses need different thresholds:

python
precisions, recalls, thresholds = precision_recall_curve(y_val, y_val_prob)

# Plot precision-recall varying with threshold
plt.plot(thresholds, precisions[:-1], label="Precision")
plt.plot(thresholds, recalls[:-1], label="Recall")
plt.xlabel("Threshold"); plt.legend(); plt.show()

Want to "find as many survivors as possible" → lower the threshold (recall↑, precision↓); want to "only predict survival when it's reliable" → raise the threshold. Threshold is a business decision, not a model parameter — it can only be selected on the validation set, not the test set.

④ Look at specific error samples

The most insightful step in error analysis: print out predicted-wrong samples and look at them one by one to see why they're wrong.

python
# Put the validation set back together to look at error samples
val_df = X_val.copy()
val_df["true"] = y_val
val_df["predicted_prob"] = y_val_prob
val_df["predicted"] = y_val_pred

errors = val_df[val_df["true"] != val_df["predicted"]]
print("Number of errors:", len(errors))
print(errors.head(15).to_string())

Looking at these error samples one by one, common findings are:

  • The model systematically mispredicts a certain type of person (e.g., misclassifies all young males as survivors) → this feature combination still lacks info; continue feature engineering.
  • Individual samples are extremely hard (e.g., "traveling alone, male, third class" but truly survived) → these are noise; no need to tune the model for them.
  • This determines whether the next step is add features, switch models, or stop here.

Step 7: Result Presentation and Visualization ​

At the end of the project, present results honestly. A qualified presentation = one metrics table + one key chart + one reproducibility note.

① Metrics summary table

ModelTrain CV AUCVal AUCVal AccuracyNotes
Majority-class baseline—0.50061.6%All predict died
Logistic regression0.850.8679.5%Interpretable
Random forest0.880.8781.1%Needs tuning
Random forest (tuned)0.890.8882.0%Final selection

② Honesty principles for presentations

  • Report val scores, report the reason for final selection (why random forest over logistic regression: slightly higher AUC and more stable, at the cost of interpretability).
  • Explicitly state limitations: only 891 samples, test set only 134 people, AUC's confidence interval is wide.
  • Never write test-set scores into the "exploration process" — test-set scores only appear once in the final report.

③ Visualization trio

python
# (a) ROC curve: TPR vs FPR at all thresholds, AUC is the area under the curve
from sklearn.metrics import roc_curve

fpr, tpr, _ = roc_curve(y_val, y_val_prob)
plt.plot(fpr, tpr, label=f"AUC = {roc_auc_score(y_val, y_val_prob):.3f}")
plt.plot([0, 1], [0, 1], "--", color="gray")   # random guessing baseline
plt.xlabel("FPR"); plt.ylabel("TPR"); plt.legend(); plt.show()

# (b) Feature importance bar chart (top 10)
# (c) Learning curve: samples vs score, judge "underfit or data shortage"
from sklearn.model_selection import learning_curve

sizes, train_scores, val_scores = learning_curve(
    gs.best_estimator_, X_train, y_train,
    cv=5, scoring="roc_auc", train_sizes=np.linspace(0.1, 1.0, 5),
)
plt.plot(sizes, train_scores.mean(1), label="train AUC")
plt.plot(sizes, val_scores.mean(1), label="CV AUC")
plt.xlabel("Training samples"); plt.legend(); plt.show()

If the learning curve shows "high train score, low CV score, converging as sample size increases" → the problem is not enough data, not insufficient model; the next step is to find more data, not tune (this judgment standard is from Overfitting and Regularization).

④ Touch the test set one final time

python
from sklearn.metrics import roc_auc_score

test_prob = gs.best_estimator_.predict_proba(X_test)[:, 1]
print("Test AUC (final score):", roc_auc_score(y_test, test_prob))

After running this line, record the score in the experiment log, and don't go touch the test set again. If you can't resist going back to change the model and re-measure — you've already been using the test set as a validation set, and the score will be inflated.

IV. Complete Code Repository Structure ​

A maintainable ML project's file organization follows "clear entry point, separated responsibilities, centralized parameters". Recommended directory structure:

titanic/
├── README.md                 # Project intro: goal, data source, how to reproduce, results
├── requirements.txt          # Dependency list (generated by pip freeze)
├── config.py                 # All hyperparams, paths, random seeds centralized here
├── data/
│   ├── train.csv             # raw data (read-only, never modify in place)
│   ├── processed/            # cleaned data (generated by scripts)
│   └── test.csv              # Kaggle test data (optional)
├── notebooks/
│   ├── 01-eda.ipynb          # exploratory analysis (exploration process, not in the reproducible path)
│   └── 02-experiments.ipynb  # experiment log
├── src/
│   ├── __init__.py
│   ├── load_data.py          # load raw data
│   ├── features.py           # cleaning and feature engineering (the sole feature implementation)
│   ├── models.py             # model definition, Pipeline, tuning logic
│   └── evaluate.py           # metrics, confusion matrix, ROC, etc.
├── models/
│   └── best_model.joblib     # serialized saved model
├── reports/
│   ├── figures/              # all chart outputs
│   └── results.md            # experiment results summary table
└── train.py                  # main entry: load → features → train → evaluate → save
File/DirectoryResponsibilityWhy placed here
config.pyCentralize random seeds, paths, hyperparamsChange params without flipping through code; experiments are reproducible
src/features.pyThe sole feature implementationTraining and online share the same feature code, preventing "train/online inconsistency"
data/ read-onlyRaw data never modified in placeEnsure every run starts from the same data
notebooks/Exploration and analysisProcess content stays in notebooks; reusable logic sinks to src/
train.pyOne-click full pipelineGoing from "runnable notebook" to "reproducible script" is a qualitative leap

A golden test for reproducible experiments: delete the notebooks and models directories, keeping only requirements.txt, src/, config.py, train.py — can someone else reproduce your score on a new machine? If the answer is "yes," your engineering meets the bar.

Principles for migrating from notebooks to scripts

Notebooks suit exploration, but don't suit reproducibility (execution order is messy, implicit state is abundant). Mature practice: after tuning logic in a notebook, extract reusable parts into functions under src/; notebooks only "call and display results." This is exactly the practical landing of the "pipeline vs script" discussion in Data Engineering.

V. Common Pitfalls ​

The project is small but complete. Below are the most common pitfalls in this workflow, and the hardest to discover once hit:

Pitfall 1: Data Leakage (leakage) ​

Leakage points encountered in this workflow, checked one by one:

Leakage PointSymptomCorrect Practice
Fitting scaler / imputer on full dataInflated val scores, crashes on deploymentPut in Pipeline or manually "fit training set first, then transform others"
Using full-data statistics in feature engineering before splitting (e.g., full-data median)Same as aboveCompute all statistics after splitting
Repeatedly using the test set for tuningFinal score higher than "true level"Test set only touched once
High-cardinality encoding for ID-like columns (Name / PassengerId)Model "memorizes" ID-to-label mappingDrop IDs directly

"Split first, process later" is an iron rule. For a more complete list (including temporal leakage, train-serve inconsistency, and 25 more), see Common Pitfalls and Anti-Patterns.

Pitfall 2: Overfitting ​

Overfitting signals and handling for this project:

  • Signal: train AUC 0.97, val 0.88 → the gap between train and val indicates the model is memorizing the training set.
  • Handling: limit max_depth and min_samples_split for tree models, add regularization (reduce C), use CV for parameter selection. Mechanics are in Overfitting and Regularization.
  • Ultimate measure: this project has only 891 samples; any model easily overfits — sample size is a hard constraint — more features, stronger models, higher overfitting risk; be restrained with feature selection.

Pitfall 3: Misused Metrics ​

  • Only looking at accuracy → fooled by "all predict died" on 38/62 imbalanced data.
  • Treating AUC as the sole truth → AUC cares about ranking, not thresholds; the business cares about "precision at threshold 0.7+."
  • Selecting metrics on the test set → metric selection itself is a form of tuning; must be decided in the validation phase.

Pitfall 4: Data Drift ​

After the model is deployed, the distribution of real data changes (e.g., changing seasons, new user groups). In a practice project, this is a next-step concern, but now is the time to write "training data limitations" in the report — patterns learned from 891 passengers in 1912 may not hold for passengers in 2024. Data's full lifecycle management is in Data Engineering.

VI. Advanced Directions ​

The Titanic project is the "minimum complete set" of the workflow; extending it into a real portfolio-worthy project can go in four directions:

  1. Deepen evaluation: add cross-validation (replace single hold-out), confidence intervals, calibration curves, upgrade evaluation from "reporting a number" to "having confidence in the score." Practical guide is in Evaluation in Practice.
  2. Switch problem and data: do regression (housing prices), multi-class (flower species), or switch to a larger, dirtier real dataset, exercising data cleaning skills. Dataset list is in Datasets and Tools Archive.
  3. More systematic training workflow: standardize the workflow with a more complete tutorial-style document, forming your own scaffold. See ML Tutorial and Evaluation in Practice.
  4. Make it a complete portfolio piece: add experiment tracking (compare multiple runs), model serialization and deployment (wrap a prediction interface with FastAPI), reproducible Docker environment, and extend it into a project you can clearly talk about in interviews/job-hunting. The complete extension roadmap is in Advanced Practice Projects.
bash
# Serialize the final model, preparing for deployment
import joblib
joblib.dump(gs.best_estimator_, "models/best_model.joblib")

# Load and predict at deployment time
model = joblib.load("models/best_model.joblib")
prob = model.predict_proba(new_passenger_features)[0, 1]   # returns survival probability

VII. Further Reading ​

References ​