Appearance
Datasets and Tools Reference
This page is the toolbox of the AI Hot Concepts Handbook: where to source data for hands-on practice, which toolchain to reach for, and how to steer clear of copyright and compliance pitfalls. Use it alongside the Models and Leaderboards Quick Reference — this page covers "data and tools", that one covers "how to pick a model".
Data determines a model's ceiling; tools determine your efficiency. In LLM projects, what really separates outcomes is rarely a single line of code — it's what data you fed in and which tools you used to turn that data into training signal. Every stage of the large language model pipeline — pretraining, fine-tuning, alignment, evaluation — maps to dedicated public datasets. This page organizes those datasets into quick-reference tables by purpose, then groups the common tools by role, and closes with selection and compliance advice.
How to use this page
Start from what you want to do (pretraining? fine-tuning? evaluation? building RAG?) and pick data from the matching section; then choose a framework and vector database from "Tool Reference"; and before any real work, read "Copyright and Compliance Notes" end to end. Companion tutorials on this site: Build a RAG App from Scratch, Fine-Tune Your Own LLM, Set Up LLM Evaluations.
Dataset Reference: By Use Case
Pretraining Corpora
Pretraining corpora are the "textbooks" of every large model. They are measured in terabytes and come almost entirely from web crawls or book/paper aggregations, which makes them inherently gray in copyright terms — the focus of several ongoing LLM lawsuits; see "Copyright and Compliance Notes" below.
| Dataset | Scale | Use | License | Link |
|---|---|---|---|---|
| Common Crawl | Billions of web pages per month; PB-scale total storage (as of 2025) | The bedrock of pretraining corpora — nearly every mainstream model has trained on it; can be filtered by domain | Data free to download; page content remains copyrighted by the original sites — you must judge permitted use yourself | Website |
| The Pile | ~800 GB, composed of 22 sub-datasets (including Books3) | The mainstream pick for academic pretraining and reproduction studies; highly diverse composition | Curated release under MIT; some subsets (especially Books3) are copyright-contested | GitHub |
| RedPajama | v1 ~1.2T tokens; v2 ~30T tokens (with GPT-4 synthetic and human-filtered samples) | Reproducing the LLaMA training recipe; RedPajama-Data-1T includes a Chinese subset | Curated release under Apache-2.0 / CC-BY-4.0; raw content copyright stays with the original sites | GitHub |
The copyright baggage of pretraining corpora
Books3 has drawn lawsuits from rights holders (the 2023 Kafka case), and LAION has been sued by multiple artists. "Downloadable online" does not mean "free for commercial use." Research use can be more relaxed; for products, verify every source individually. For the full risk discussion, see AI Safety and Governance.
Instruction-Tuning Data
A pretrained model only knows how to "continue text"; instruction-tuning data teaches it to "follow instructions." This tier of data is mostly synthesized by large models or crowdsourced, and it is the easiest entry point for personal fine-tuning.
| Dataset | Scale | Use | License | Link |
|---|---|---|---|---|
| Alpaca | 52K instruction-response pairs (synthesized by GPT-3.5 from seed tasks) | Entry-level SFT — the classic starting point for "fine-tune your own little assistant" | CC BY-NC 4.0; the official copy has been taken down, community mirrors remain | GitHub |
| ShareGPT | ~90K real user conversations shared by the community | Conversational SFT with broad coverage | No formal license (scraped user data — evaluate carefully before any commercial use) | Website |
| OpenAssistant (OASST1) | ~46K multilingual crowdsourced conversation trees (with preference annotations) | The go-to open-source dual-purpose set for conversational SFT and preference data | Apache-2.0 | Hugging Face |
| UltraChat | ~1.5M synthetic multi-turn conversations | Large-scale conversational SFT; trained community mainstays such as Zephyr | CC BY-NC 4.0 | Hugging Face |
Preference Data (for Alignment)
Preference data is the fuel of alignment (RLHF and DPO): each row is a human judgment that "for the same input, answer A beats answer B." It decides whether a model is "willing" to speak in line with human values.
| Dataset | Scale | Use | License | Link |
|---|---|---|---|---|
| Anthropic HH-RLHF | ~160K "helpful/harmless" preference comparisons | The de facto standard for RLHF/DPO training and alignment research | MIT (defer to official labeling) | Hugging Face |
| OpenAI Summarize from Feedback | ~90K summarization preference comparisons | RLHF and preference-model training for summarization tasks | As per official release | GitHub |
| Stanford SHP | ~380K community preferences on "which answer is better" | Reward-model and DPO training (a community favorite) | Academic license (defer to official terms) | Hugging Face |
A note on "OpenAI-related" data
The labeler comparison data used in OpenAI's InstructGPT paper (Ouyang et al., 2022) was never fully released. What researchers can actually get is mainly the Summarize from Feedback dataset above, plus third-party repackaged preference pairs. Don't claim in a paper that you used "OpenAI alignment data" without a traceable source.
Evaluation Benchmarks
Evaluation benchmarks are the "exam papers" used to quantify model capability. The LLM Evaluations and Benchmarks page covers methodology in depth; here are the four most-cited exams — also the names that appear most often in model launch news.
| Dataset | Scale | Use | License | Link |
|---|---|---|---|---|
| MMLU | ~14K multiple-choice questions across 57 subjects | The standard test of knowledge breadth — measures "general education" | MIT | GitHub |
| GSM8K | 8.5K grade-school math word problems (~1.3K in the test split) | Math reasoning evaluation | CC BY 4.0 | GitHub |
| HumanEval | 164 hand-written programming problems (with unit tests) | Code generation evaluation | MIT | GitHub |
| C-Eval | 13,948 Chinese questions across 52 subjects | Chinese-language LLM evaluation — measures a model's "Chinese chops" | CC BY-NC-SA 4.0 | GitHub |
Benchmark saturation alert
MMLU has been pushed to around 90% by frontier models; questions everyone can answer no longer separate models. When reading reports, pay more attention to fresh benchmarks that haven't saturated yet (e.g., GPQA, LiveBench) and to evaluation sets you build from your own business scenarios. Don't stare at a single leaderboard.
Retrieval / RAG Data
Retrieval-augmented generation (RAG) needs "document collections + queries" data to train retrievers and evaluate retrieval quality. The hallmark of this tier is standard query-document relevance annotations, so you can compute MRR, NDCG, and similar metrics directly.
| Dataset | Scale | Use | License | Link |
|---|---|---|---|---|
| MS MARCO | ~1M queries + 8.8M passages | The standard exam for passage retrieval and ranking — BERT-era retrieval models all ran on it | CC BY 4.0 | Website |
| Natural Questions | ~300K real Google search questions (paired with Wikipedia-passage answers) | Open-domain QA and long-document retrieval | CC BY-SA 3.0 | Website |
| KILT | 5 task types across 6 datasets (ELI5, FEVER, Wizard of Wikipedia, etc.) | Unified evaluation for knowledge-intensive tasks, covering QA / fact verification / dialogue | Licenses vary by subset — defer to official terms | GitHub |
Rule of thumb
Use MS MARCO as your retrieval baseline (the community has published the most solutions for it), and turn to KILT for knowledge-intensive QA. If your business runs on Chinese-language document collections, don't forget to build your own eval set of "business docs + human-labeled relevance" — public data can never replace your domain. Related reading: Knowledge Graphs and Knowledge Injection.
Multimodal Data
Multimodal models need paired data — image↔text and speech↔text. The Multimodal Models page covers the principles; here are the three most-used open corpora.
| Dataset | Scale | Use | License | Link |
|---|---|---|---|---|
| LAION-5B | 5.85 billion image-text pairs (URLs scraped from Common Crawl) | Contrastive image-text pretraining (CLIP, Stable Diffusion-style models) | Annotations released with the project; image copyright stays with owners — heavily disputed | LAION blog |
| COCO | ~330K annotated images, 80 object categories, 1.5M instances | Detection / segmentation / image captioning; standard multimodal evaluation kit | CC BY 4.0 | Website |
| Common Voice | Crowdsourced speech in 100+ languages (as of 2025), total hours keep growing | Speech recognition (ASR) training and evaluation | Each sample individually declared CC0 or CC-BY | Website |
Tool Reference
Frameworks
| Tool | Maintainer | One-line positioning | License | Link |
|---|---|---|---|---|
| Transformers | Hugging Face | Unified interface for loading / fine-tuning / inference with mainstream models; the de facto standard | Apache-2.0 | Docs |
| LangChain | LangChain team | LLM application orchestration: chains, tool calling, RAG, agent scaffolding | MIT | Website |
| LlamaIndex | LlamaIndex team | A data-centric RAG framework: loading, indexing, and querying in one package | MIT | Website |
| vLLM | vLLM community | High-throughput inference serving (PagedAttention, continuous batching); the top pick for production deployment | Apache-2.0 | Docs |
| peft | Hugging Face | Parameter-efficient fine-tuning (LoRA/QLoRA, etc.) — runs on a single consumer GPU | Apache-2.0 | Docs |
| trl | Hugging Face | Ready-made implementations of SFT, DPO, PPO and other training and alignment pipelines | Apache-2.0 | Docs |
Vector Databases
Vector databases handle "semantic similarity" retrieval — the foundation of RAG. For the full mechanics, see Vector Databases and Semantic Search.
| Tool | Type | Best for | License | Link |
|---|---|---|---|---|
| FAISS | Single-machine retrieval library | Prototyping and similarity search up to the million-vector scale; plays seamlessly with Python/GPUs | MIT | GitHub |
| Milvus | Distributed vector database | Very large scale, multi-replica, cloud-native production deployments | Apache-2.0 | Website |
| Qdrant | Vector database (Rust) | Fast and easy to deploy, with built-in metadata filtering and high availability | Apache-2.0 | Website |
| pgvector | PostgreSQL extension | Keep vectors inside your existing business database; transactions and retrieval managed together | PostgreSQL License | GitHub |
Rule of thumb
FAISS for single-machine prototypes, pgvector when vectors must live alongside your business data, Qdrant or Milvus when you're serious about a product. Don't introduce a distributed vector database for a few tens of thousands of vectors — over-engineering is the most common anti-pattern.
Annotation & Evaluation
| Tool | One-line positioning | License | Link |
|---|---|---|---|
| Argilla | Open-source data annotation and feedback platform for LLMs — review-style labeling of model outputs, serving SFT and preference data construction | Apache-2.0 | Docs |
| promptfoo | Prompt regression testing and red-teaming tool: define cases on the command line, run batch scoring, catch regressions; pairs well with prompt engineering | MIT | Website |
Deployment
| Tool | One-line positioning | License | Link |
|---|---|---|---|
| Ollama | Download and run open models locally with a single command (built-in API and model registry); the least-fuss option for personal and prototype deployments | MIT | Website |
| Docker | Package model services and environments as containers; the production standard | Apache-2.0 | Website |
Selection Guide: What to Use When
More tools is not better — pick the smallest viable combination for the task:
| Your task | Recommended combo | Further reading |
|---|---|---|
| Run a concept demo / try open models locally | Ollama + Transformers | Deployment and Inference Optimization in Practice |
| Fine-tune your own small model | peft + trl + an SFT dataset (Alpaca/UltraChat) | Fine-Tune Your Own LLM |
| Build a RAG Q&A app | LlamaIndex or LangChain + FAISS/pgvector + bge embeddings | Build a RAG App from Scratch |
| Ship a high-throughput inference service | vLLM + Docker | Deployment and Inference Optimization in Practice |
| Set up evaluation and regression testing | promptfoo + benchmark datasets | Set Up LLM Evaluations |
| Build SFT / preference data | Argilla annotation + your own business data | Alignment: RLHF and DPO |
Copyright and Compliance Notes
- A license is not free rein:
CC BY-NCbans commercial use,CC BY-SArequires derivatives to stay equally open, and only Apache/MIT come close to "use it freely." The first thing to do after downloading a dataset is read the license field. - Scraped data is high-risk: Common Crawl, Books3, ShareGPT, LAION — for data "grabbed from the web," copyright belongs to the original content owners, and nearly every lawsuit lands here. Fine for research, use with caution in production, and for commercial use go with licensed data or a synthetic data + cleaning pipeline.
- Don't misuse user data: ShareGPT involves real user conversations; using it directly in a commercial product violates platform terms and may also violate privacy regulations such as GDPR.
- Self-check list before shipping a product: (1) archive data sources and licenses in writing; (2) anonymize sensitive personal information; (3) review commercial-use terms line by line (especially CC BY-NC and competition data agreements); (4) re-check periodically, because licenses can change.
Compliance is not just the legal team's job
Startup teams often push compliance onto legal, but every step of data selection is already a compliance decision. One extra day spent reading licenses can save a year of litigation. For more on risk governance, see AI Safety and Governance.
Further Reading
- Models and Leaderboards Quick Reference — data picked out? now choose the model
- LLM Evaluations and Benchmarks — the methodology and pitfalls behind evaluation benchmarks
- Fine-Tuning and PEFT (LoRA) — how instruction data becomes model capability
- Retrieval-Augmented Generation (RAG) — how to use retrieval data and pipelines
- Vector Databases and Semantic Search — the principles behind vector DB selection
- Build a RAG App from Scratch and Fine-Tune Your Own LLM — hands-on practice entry points
- Curated Resource List and Glossary — more learning materials and term definitions
- Learning Paths — fit this page's tools into an overall learning route
References
- Common Crawl official website — monthly crawl archives and documentation; the source of record for scale figures
- The Pile: an 800GB open pretraining corpus — EleutherAI paper (Gao et al., 2021)
- Stanford Alpaca project page — original source of the 52K instruction data
- Anthropic HH-RLHF dataset — the standard reference for preference data
- MMLU paper — Measuring Massive Multitask Language Understanding (Hendrycks et al., 2020)
- C-Eval website — official description of the Chinese evaluation benchmark
- MS MARCO website — official documentation for the retrieval evaluation data
- LAION-5B release announcement — scale and construction methodology
- Hugging Face Datasets — where most of the datasets above are officially hosted
- vLLM official documentation — the authoritative reference for the inference deployment tool
Dataset scale and license information on this page follows each project's official pages as of mid-2025; re-verify license terms and versions before you start.