Skip to content

Datasets & Tools Reference

Quick overview Machine learning datasets and tools reference: a categorized quick-reference table of datasets by task (Iris, MNIST, ImageNet, SQuAD, MovieLens, etc.), main acquisition channels, annotation and synthetic data tools, common libraries and data engineering tools, and practical advice on selecting datasets for exercises.

Datasets & Tools Reference ​

This page is the toolbox chapter of the Machine Learning Handbook: where to find data, what tools to use for annotation, what libraries for modeling, how to version and engineer data, and "which dataset to pick for your next exercise".

Data is the starting point of machine learning — and also its ceiling: feature quality determines the model's upper bound, and features come from the data itself. The Glossary teaches you to read the concepts; this page tells you "what tools you should have on hand and where to get them". Recommended to read alongside Data and Data Engineering and Feature Engineering.

How to use

This section is a quick-reference catalog. When doing exercises, first flip to the "Dataset Quick-Reference" to pick data, then download from the "Acquisition Channels", and model with the libraries in the "Common Libraries Toolbox"; as data accumulates, use the "Data Engineering Tools" to manage versioning, quality, and deployment.

Dataset Quick-Reference ​

All datasets below are real; scale and purpose are as described by their official sources. The "How to Access" column gives the official source; those marked with ⭐ are also built into scikit-learn / Hugging Face, loadable with a single import or one line of load_dataset.

Classification Tasks ​

DatasetScale & StructureTypical UseHow to Access
Iris ⭐150 samples, 4 features, 3 classes (setosa/versicolor/virginica)Introductory binary/multi-class, KNN, logistic regression teachingscikit-learn load_iris(); UCI
Titanic891 training records, ~12 features (with missing values)Introductory binary classification + full feature engineering exerciseKaggle competition
Adult (Census Income) ⭐48,842 samples, 14 featuresBinary income classification, class imbalance, fairness researchUCI; OpenML ID 1590
Wisconsin Breast Cancer ⭐569 samples, 30 featuresIntroductory binary classification, model comparisonscikit-learn load_breast_cancer(); UCI
Wine QualityRed 1,599 + White 4,898 samples, 11 chemical featuresDual exercise: multi-class / regressionUCI
Credit Card Fraud284,807 transactions, only 492 fraudulent (0.17%)Highly imbalanced classification, cost-sensitive learningKaggle dataset
Heart Disease303 samples, 14 featuresSmall-sample binary classification, cross-validation practiceUCI

Regression Tasks ​

DatasetScale & StructureTypical UseHow to Access
California Housing ⭐20,640 samples, 8 featuresIntroductory regression, linear model teachingscikit-learn fetch_california_housing()
House Prices (Kaggle)1,460 training samples, 79 featuresReal-world house price regression, feature engineering competitionKaggle competition
Boston Housing506 samples, 13 featuresClassic regression teaching (has historical controversy, removed from scikit-learn 1.2, use the OpenML version)OpenML ID 531
Medical Costs1,338 samples, 7 featuresInsurance cost prediction, categorical feature encodingKaggle dataset

The Lesson of Boston Housing

Boston Housing was removed from scikit-learn 1.2 because it contains a feature highly correlated with housing price ("proportion of Black population") and is outdated; the academic community generally recommends replacing it with California Housing. It reminds us: the ethics and timeliness of datasets are themselves part of data engineering.

Image Tasks ​

DatasetScale & StructureTypical UseHow to Access
MNIST (Handwritten Digits) ⭐70,000 28×28 grayscale images, 10 classesVision intro, the "Hello World"Original source; Kaggle, Hugging Face
Fashion-MNIST ⭐70,000 28×28 grayscale images, 10 clothing classesA harder alternative to MNISTGitHub; Hugging Face
CIFAR-10 / CIFAR-10060,000 32×32 color images each (10 classes / 100 classes)CNN intro, data augmentation exerciseOfficial page; Hugging Face
SVHN (Street View House Numbers)Over 600,000 32×32 color digit imagesReal-world OCR, transfer learningStanford UFLDL
Oxford-IIIT Pets7,349 images, 37 breedsSmall-data image classification / segmentation exerciseVGG official page
ImageNet (ILSVRC subset)~1.28M training images, 1,000 classesLarge-scale pretraining and model benchmarksOfficial site
COCO (Common Objects in Context)~330K annotated images, 80 object classes, 1.5M instancesDetection, segmentation, keypoint, image captioningOfficial site

Text & NLP ​

DatasetScale & StructureTypical UseHow to Access
IMDB Reviews ⭐50,000 entries (25k train / 25k test), binary sentiment classificationSentiment analysis intro, word embeddingsStanford official; Hugging Face
SST-2 (GLUE sub-task) ⭐67,349 training sentencesSentence-level binary sentiment classification, BERT evaluationGLUE benchmark; Hugging Face
AG News ⭐~120K training entries, 4 topic classesText classification, baseline comparisonHugging Face
SQuAD (Extractive QA)Version 2.0: ~150K questions, 50K Wikipedia articlesExtractive reading comprehension, BERT fine-tuningSQuAD official; Hugging Face
GLUE / SuperGLUE benchmarkMulti-task suite (including linguistic acceptability, entailment, coreference resolution, etc.)Universal evaluation of pre-trained modelsGLUE / SuperGLUE
THUCNews (Chinese News)740K Chinese news articles, 14 topic classesChinese text classificationGitHub mirror; Hugging Face

Recommendation Systems ​

DatasetScale & StructureTypical UseHow to Access
MovieLensThree scales: 100K / 1M / 25MCollaborative filtering, matrix factorization teaching standardGroupLens official
Netflix Prize100M ratings (original data no longer available, use mirrors)Historical benchmark for recommendation competitionsKaggle mirror
Amazon ReviewsOver 230M reviews, split across multiple domainsReview-based recommendation, implicit feedbackUCSD official
Last.fm Listening RecordsUser-track implicit interactionsImplicit feedback recommendation, graph-based methodsGroupLens page

Time Series ​

DatasetScale & StructureTypical UseHow to Access
Air Passengers144 monthly observations, 1949–1960Time series intro, ARIMA / naive baselineBuilt into R datasets package / statsmodels
ETTh1/ETTh2 (Transformer Temperature)2 years, every 15 min / hourly, 7 variablesLong-sequence prediction, Transformer time-series benchmarkETDataset GitHub
Jena ClimateEvery 10 min starting from 2003, 14 meteorological variablesMulti-variable time series, LSTM/Transformer exampleKeras example page; Kaggle mirror
M5 ForecastingHistorical sales for 30,490 Walmart productsMulti-level sales forecasting competitionKaggle competition
Stock Prices (Yahoo Finance)Historical OHLCV data for any tickerFinancial time series (note: for practice only, not trading)Pull with yfinance

The "Last Mile" of Time Series Data

Time series data inherently carries temporal order — never shuffle it randomly. Splits must use TimeSeriesSplit. When practicing with Air Passengers, start with a naive baseline like "previous value's prediction" — any model must beat that baseline before its value can be discussed.

Dataset Acquisition Channels ​

ChannelURLCharacteristicsBest For
Kagglekaggle.comCompetitions + datasets + Notebook, all in one; active community, rich solution write-upsCompetition practice, real business data, training a "competition mindset"
Hugging Face Datasetshuggingface.co/datasetsOne-line load_dataset() loading; seamless integration with Transformers; hosts community datasetsNLP/CV model fine-tuning, rapid prototyping
UCI ML Repositoryarchive.ics.uci.eduThe oldest teaching data warehouse, 600+ classic datasetsTeaching classics (Iris, Adult, Wine, etc.)
Papers with Codepaperswithcode.comBinds papers + code + benchmark scores; dataset pages include SOTA leaderboardsReproducing papers, checking current state-of-the-art for a task
OpenMLopenml.orgDatasets come with stable IDs and metadata, directly integrable with scikit-learnReproducible experiments, research benchmarks
Google Dataset Searchdatasetsearch.research.google.comWeb-wide dataset search engineFinding niche / domain-specific data
Awesome Public DatasetsGitHubOpen-source data lists organized by domainBrowsing by industry (finance, healthcare, geography…) for inspiration

The Right Way to Get Data from Kaggle

Beyond official competition data, Kaggle has a large number of user-uploaded datasets. Assess quality by three criteria: is the intended use clearly stated, has it been recently updated, and do notebooks contain validation scripts. Popular competitions (Titanic, House Prices) come with official evaluation metrics and leaderboards — the best "practice problems with known answers".

Annotation & Synthetic Data Tools ​

Supervised learning relies on labels, and manual annotation is expensive and slow. Tools fall into three categories: manual annotation platforms, weak supervision / automatic error correction, and synthetic data generation.

Manual Annotation ​

ToolURLOne-Liner
Label Studiolabelstud.ioOpen-source annotation platform covering images/text/audio/video/time series, supports multi-person collaboration and ML-assisted annotation, the top choice for personal practice
Labelboxlabelbox.comEnterprise-grade annotation and data management platform, mature for CV scenarios
Prodigyprodi.gyProgrammable annotation tool by Explosion, deeply integrated with spaCy/Transformers, great for NLP active learning
Scale AIscale.comCommercial crowdsourced annotation service, commonly seen in large-scale data like autonomous driving
AWS SageMaker Ground Truthofficial pageCloud-hosted annotation, supports built-in human review + automatic annotation pipelines

Weak Supervision & Label Correction ​

ToolURLOne-Liner
Snorkelsnorkel.orgProgrammatic weak supervision framework: write "labeling rules" as functions, automatically fuse to produce noisy training labels
Cleanlabcleanlab.aiUses confident learning to automatically discover and correct label errors in datasets
Argillaargilla.ioOpen-source data annotation and feedback tool for the LLM era, enables "review annotation" of model predictions

Synthetic Data ​

ToolURLOne-Liner
SDV (Synthetic Data Vault)docs.sdv.devGenerates new data from real data using generative models, preserving distribution and privacy
Fakerfaker.readthedocs.ioGenerates "realistic-looking" fake data like names/addresses/emails, extremely convenient for tabular/test data
Mimesismimesis.nameA high-performance alternative to Faker, supporting multiple languages
scikit-learn Synthesizersofficial docsmake_classification / make_regression / make_blobs — generate data with specified shape and difficulty in one line, a teaching gem

When Not to Rely on Synthetic Data

Synthetic data can never replace the "irregularities" in real data — missing values, dirty values, long tails, hidden biases are precisely the most valuable parts of real-world data. Synthetic data is suitable for: privacy-constrained scenarios, supplementing long-tail classes, and testing code correctness. For learning feature engineering, always use real data (see Feature Engineering).

Common Libraries Toolbox ​

Ordered by usage frequency, with one sentence of "when to use it" per category.

Data Processing & Visualization

  • NumPy — the foundation of numerical computing in Python: multi-dimensional arrays and vectorized operations, the bedrock of every library.
  • pandas — the de facto standard for tabular data, DataFrame for cleaning, pivoting, grouping, and joining.
  • Polars — a faster DataFrame implemented in Rust, several times faster than pandas at 100M+ rows, with a more modern API.
  • Dask — pandas' big-data twin: use it for distributed/lazy computation when data exceeds memory.
  • matplotlib — the most basic plotting library, can plot anything but verbose API.
  • seaborn — a statistical plotting library built on matplotlib, one-liners for distribution/correlation plots.

Traditional Machine Learning

  • scikit-learn — the standard library for traditional ML: models, preprocessing, cross-validation, metrics, and Pipeline in one place, use it for learning algorithms.
  • XGBoost — the de facto standard for gradient boosting trees, the champion tool for tabular data.
  • LightGBM — Microsoft's faster gradient boosting, often outperforms earlier XGBoost at large scale.
  • statsmodels — dedicated to statistical modeling and testing: p-values for linear regression, time series ARIMA, hypothesis testing.

Deep Learning

  • PyTorch — the current mainstream deep learning framework in both research and industry: dynamic graph, largest ecosystem. See the site's Deep Learning Fundamentals and Optimization sections.
  • TensorFlow / Keras — the veteran framework; Keras' high-level API is beginner-friendly, production deployment (TFLite/TF Serving) is mature.
  • Hugging Face Transformers (library) — a unified interface for loading pre-trained models: BERT, GPT, LLaMA, T5 usable in one line, with a companion Trainer that simplifies fine-tuning.
  • Hugging Face Datasets (library) — a dataset library paired with Transformers: lazy loading, memory mapping, sharding — processing large text datasets without blowing up memory.

The Golden Rule for Choosing Libraries

Use scikit-learn for learning principles, XGBoost/LightGBM for production, PyTorch for deep learning, Transformers for NLP. Don't use deep learning for tabular data, and don't use sklearn to train large models — tool selection itself is an engineering decision.

Data Engineering Tools ​

When data volume grows, "having data" isn't enough — you also need to "manage data". This toolchain turns data into a sustainably evolving asset; the full workflow is in Data and Data Engineering.

ToolURLProblem It Solves
DVC (Data Version Control)dvc.orgManage data the same way you manage code with Git: dataset versioning, rollback, reproducible experiments
Feastfeast.devOpen-source feature store: training features and online inference features share the same definition,solving "training/online inconsistency"
Apache Airflowairflow.apache.orgThe most popular batch-processing workflow scheduler: data pipeline scheduling with timing/dependencies
Prefectwww.prefect.ioA more approachable modern data flow orchestrator, natively Python with strong dynamic scheduling
Dagsterdagster.ioAn orchestration framework built around data assets, treating "data dependencies" as a first-class citizen
Great Expectationsgreatexpectations.ioData quality testing framework: declare "what data should look like", automatically validates after pipelines run
dbtwww.getdbt.comEngineering tool for data transformations using SQL, the standard tool for analytics engineers
MLflowmlflow.orgUnified entry point for experiment tracking + model registry + deployment, the first stop for MLOps adoption
DuckDBduckdb.orgEmbedded analytical database, run SQL directly on pandas/Parquet files — a single-machine big-data analysis tool

Don't Roll Out the Full Suite From Day One

For beginners and small-to-medium projects, DVC + MLflow + Great Expectations — a three-piece set covering the most critical pain points of "versioning, experiments, and quality" — is sufficient, and all are open-source with single-machine deployment. Hold off on Airflow/Feast until you truly have deployment and multi-team collaboration needs; introducing toolchains too early is an anti-pattern.

How to Choose Datasets for Practice ​

Choosing data is the first step of practice, and where many learners get stuck. Here are four pieces of experience.

1. Pick Difficulty by Learning Phase ​

  • Beginner (first 3 projects): pick small, clean, well-documented ones — Iris, Titanic, MNIST. The goal is to run through the full workflow, not to chase scores.
  • Advanced (after mastering fundamentals): switch to data with noise, missing values, and real-world pain points — House Prices (79 features teaching you feature engineering), Credit Card Fraud (highly imbalanced teaching you to change metrics).
  • Further advanced: dig into your chosen direction — CV: CIFAR-10 → ImageNet subset; NLP: IMDB → SQuAD; Recommendation: MovieLens; Time series: Air Passengers → ETTh. See Learning Paths for route planning.

2. Pick Tasks by Practice Goal ​

Skill to PracticeRecommended DataWhy
Feature engineeringHouse Prices, TitanicMany missing values, categorical features, long-tail distributions — the most operations to perform
Model tuningCIFAR-10, IMDBMassive community solution write-ups, easy to compare your hyperparameter choices
Deep learningMNIST → Fashion-MNIST → CIFAR-10Smooth difficulty progression, each level adding "real problems" (e.g., data augmentation)
Engineering deploymentAny real data + your own data pipelineThe focus is not on model accuracy but on DVC versioning, reproducibility, and monitorability
Interview portfolio piecePick one where you can tell a "business story"See Portfolio Projects and Build Your Own

3. Check Whether the Data Is "Worth Your Time" ​

  • Licensing: Kaggle competition data has clear competition agreements; UCI/OpenML are mostly permissive. Always confirm licensing terms before commercial or public use.
  • Source and timeliness: Pre-2010 datasets are fine for teaching but not for discussing "current state". Be wary of data with questionable annotation quality (e.g., scraped data).
  • Is there still "something to do": If a dataset has already been squeezed to 99%+ accuracy, don't expect it to produce results — go to Papers with Code for harder benchmarks.
  • Are you genuinely interested: Willingness to invest another 20 hours into "this topic" is the greatest driver for completing a project.

4. Build Your Own "Data Garden" ​

Keep notes on good datasets you've used, pitfalls you've encountered (which dataset was dirty, which had misannotations), and verified loading code. This habit, paired with the Awesome Resources and Glossary, will make your learning accumulation grow thicker and thicker.

Further Reading ​

All links on this page point to official sources. Dataset scale information is current as described on each official page; licensing terms should be verified before downloading.