Theme
Datasets & Tool Archive
In one sentence: a practical archive answering "where to get data and what tools to use" — classic datasets grouped by modality (size, task, source, notes) + one-line profiles of common tools + decision principles for choosing datasets.
Timeliness note: Dataset sizes, versions, and other information are current as of 2026-08. Verify before citing. Download links all point to official channels.
Data is the raw material of deep learning: the model's ceiling is determined by the data. See Data & Data Engineering. This page addresses three practical questions: which dataset to use for practice, where to download, and what tools to run with.
1. Image Datasets
| Dataset | Size | Task | How to Get | Notes |
|---|---|---|---|---|
| MNIST | 70K 28×28 handwritten digits | 10-class classification | Official site / direct download via Torchvision | Too "clean" — you can't overfit on it. Good for running through a pipeline, not for chasing SOTA scores |
| CIFAR-10/100 | 60K 32×32 color images | 10/100-class classification | Official site / Torchvision | Low resolution, but the standard battleground for quickly validating network architectures |
| ImageNet | ~14M images, 22K classes | Image classification, pre-training benchmark | Official site (apply for download) | Large volume (hundreds of GB). Academic use requires agreeing to terms. Most people do pre-training comparisons on ImageNet-1K |
| COCO | 330K images, 2.5M annotated instances | Detection / segmentation / keypoints / captions | Download files from official site | Annotation format (JSON) needs parsing. TorchVision has built-in loaders |
| Open Images | 9M+ images | Detection / segmentation / visual relationships | Google Cloud storage shards | Large scale, higher annotation noise than COCO, suited for pre-training |
Tip: For all image datasets, first run through the pipeline using TorchVision's built-in interfaces before considering full downloads. When processing large-scale data, use DataLoader multi-processing and caching. See Training Recipes & Hyperparameter Tuning for details.
2. Text Datasets
| Dataset | Size | Task | How to Get | Notes |
|---|---|---|---|---|
| GLUE | A suite of 9 NLU tasks (includes SST-2, MNLI, QQP, etc.) | Sentence classification, natural language inference, similarity, etc. | HuggingFace datasets | For Chinese, check out CLUE as a counterpart. A classic benchmark for general language capability |
| SuperGLUE | A harder version of GLUE (8 more difficult tasks) | Reasoning, common sense, etc. | HuggingFace datasets | Fewer samples — pay attention to metric computation details |
| SQuAD | 100K+ QA pairs (SQuAD1.1); includes unanswerable questions (SQuAD2.0) | Extractive QA | Official site / HuggingFace | Version 2.0 adds unanswerable samples, testing the model's ability to "refuse to answer" |
| Wikipedia (English) | Billions of tokens | Pre-training corpus, RAG knowledge source | Official dumps, released monthly | Enormous volume — usually download subsets on demand. Commonly used as a knowledge base for RAG practice. See Large Language Models (LLMs) |
| Common Crawl | Hundreds of TB of web data | Large-scale pre-training corpus | Official monthly snapshots | Requires extensive cleaning and deduplication. Hard for individuals to use in full — usually use subsets |
| The Pile | 820GB multi-source text | Open-source pre-training corpus | Official site / HuggingFace | Diverse components (papers, GitHub, books, conversations, etc.). A common corpus for large-scale pre-training research |
Tip: Chinese corpora can be obtained from HuggingFace's Chinese datasets (e.g., C4 Chinese version, CLUECorpus). Any pre-training corpus needs deduplication, filtering, and privacy cleaning. See Data & Data Engineering.
3. Speech & Audio Datasets
| Dataset | Size | Task | How to Get | Notes |
|---|---|---|---|---|
| LibriSpeech | ~1,000 hours of English read speech | Automatic Speech Recognition (ASR) | OpenSLR shard downloads | 16kHz sampling rate, high-quality annotation alignment — the standard ASR benchmark |
| Common Voice | Multilingual crowdsourced speech (continuously growing) | ASR, multilingual | Mozilla official download | High noise, diverse accents — realistic but uneven annotation quality. See Speech & Audio |
4. Multimodal & Molecular/Graph Datasets
| Dataset | Size | Task | How to Get | Notes |
|---|---|---|---|---|
| LAION-5B | 5.8B image-text pairs | Image-text pre-training, CLIP training | Official shard-index downloads | Unfiltered crawled data — contains harmful/low-quality content. Research use requires attention to compliance and content review |
| Conceptual Captions (CC) | ~3.3M image-text pairs | Image-text pre-training | GitHub official link | Requires downloading original images separately; link validity needs self-checking |
| QM9 | 134K organic molecules | Molecular property regression | Official site | A standard entry-level dataset for graph/molecular ML. See Graph Neural Networks |
| OGB (Open Graph Benchmark) | Multi-task graph benchmarks (node / edge / graph-level) | Graph learning benchmark | Official site / HuggingFace | Has official evaluation protocols and leaderboards. Don't tune on the test set |
5. Tool Profiles
| Tool | One-line Description | Use Case |
|---|---|---|
| PyTorch | The dominant deep learning framework | Main modeling and training framework |
| TorchVision | PyTorch's official vision toolkit | Image dataset loading, pre-trained models, transforms |
| Hugging Face Transformers | Unified API for loading thousands of pre-trained models | Text/vision/speech model loading and fine-tuning |
| Hugging Face datasets | Dataset loading and processing library | Stream large corpora, unified preprocessing |
| diffusers | Official diffusion model toolkit | Generative model training and inference. See Diffusion Models & Generative AI |
| peft | Parameter-efficient fine-tuning toolkit | LoRA and other fine-tuning methods out of the box |
| accelerate | Distributed training simplifier | Run multi-GPU / mixed-precision training with a few lines of code |
| ONNX Runtime | Cross-platform inference engine | Model export and deployment. See MLOps & Model Deployment |
| TensorBoard / W&B | Visualization and experiment tracking | Metric monitoring, experiment comparison |
| OpenAI CLIP | Image-text contrastive learning model | Image-text retrieval, zero-shot classification, feature extraction |
Tool selection advice
Beginners should use the "PyTorch + HuggingFace family (Transformers/datasets/peft/accelerate)" end-to-end. For research efficiency, learn JAX. Add ONNX / inference frameworks for production deployment. A more complete selection methodology is at How to Choose Frameworks & Tools.
6. Principles for Choosing Datasets
Principle 1: Task Before Data
First clarify "what problem to solve and what metric to use for evaluation," then choose data. If the metric is wrong, no dataset size matters. See Deep Learning Evaluation & Experimentation.
Principle 2: Size Matches the Stage
- Learning principles / running through a pipeline: MNIST, CIFAR — validate ideas within minutes;
- Serious projects / portfolio: COCO, GLUE, LibriSpeech — scale and annotation quality sufficient for resume writing;
- Pre-training / research: ImageNet, The Pile, LAION — requires compute and engineering support.
Principle 3: Watch for Three Categories of Data Problems
- Leakage: Training/test mixing, temporal leakage of future information — inflates metrics. See Common Pitfalls & Anti-patterns;
- Distribution: Dataset distribution differs from the real-world scenario — the model fails upon deployment. See Evaluation in Practice;
- Compliance: Licensing and ethical issues for crawled data, person images, and copyrighted text — especially critical for datasets like LAION. See Interpretability & Fairness.
Principle 4: Prioritize "Ecosystem-Rich" Datasets
Datasets with official loaders (TorchVision / HF datasets), public baseline scores, and community discussion make your results comparable and interpretable — which is why ImageNet and GLUE have become de facto standards.
Further Reading
- Data & Data Engineering — Data cleaning, annotation, and imbalance handling
- Training Recipes & Hyperparameter Tuning — Data loading and training pipeline optimization
- Deep Learning Evaluation & Experimentation — Choose the right metrics before choosing data
- Large Language Models (LLMs) — Pre-training corpora and RAG data practices
- How to Choose Frameworks & Tools — Toolchain selection methodology
- Portfolio Projects — Use datasets to build projects worth writing on your resume