Skip to content

Datasets and Tools Reference

At a glance A reference catalog of LLM datasets and tools: quick-lookup dataset tables organized by purpose (pretraining, instruction tuning, preference, evaluation, RAG, multimodal), plus tool recommendations for frameworks, vector databases, annotation, and deployment — with copyright and compliance notes.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

Datasets and Tools Reference ​

This page is the toolbox of the AI Hot Concepts Handbook: where to source data for hands-on practice, which toolchain to reach for, and how to steer clear of copyright and compliance pitfalls. Use it alongside the Models and Leaderboards Quick Reference — this page covers "data and tools", that one covers "how to pick a model".

Data determines a model's ceiling; tools determine your efficiency. In LLM projects, what really separates outcomes is rarely a single line of code — it's what data you fed in and which tools you used to turn that data into training signal. Every stage of the large language model pipeline — pretraining, fine-tuning, alignment, evaluation — maps to dedicated public datasets. This page organizes those datasets into quick-reference tables by purpose, then groups the common tools by role, and closes with selection and compliance advice.

How to use this page

Start from what you want to do (pretraining? fine-tuning? evaluation? building RAG?) and pick data from the matching section; then choose a framework and vector database from "Tool Reference"; and before any real work, read "Copyright and Compliance Notes" end to end. Companion tutorials on this site: Build a RAG App from Scratch, Fine-Tune Your Own LLM, Set Up LLM Evaluations.

Dataset Reference: By Use Case ​

Pretraining Corpora ​

Pretraining corpora are the "textbooks" of every large model. They are measured in terabytes and come almost entirely from web crawls or book/paper aggregations, which makes them inherently gray in copyright terms — the focus of several ongoing LLM lawsuits; see "Copyright and Compliance Notes" below.

DatasetScaleUseLicenseLink
Common CrawlBillions of web pages per month; PB-scale total storage (as of 2025)The bedrock of pretraining corpora — nearly every mainstream model has trained on it; can be filtered by domainData free to download; page content remains copyrighted by the original sites — you must judge permitted use yourselfWebsite
The Pile~800 GB, composed of 22 sub-datasets (including Books3)The mainstream pick for academic pretraining and reproduction studies; highly diverse compositionCurated release under MIT; some subsets (especially Books3) are copyright-contestedGitHub
RedPajamav1 ~1.2T tokens; v2 ~30T tokens (with GPT-4 synthetic and human-filtered samples)Reproducing the LLaMA training recipe; RedPajama-Data-1T includes a Chinese subsetCurated release under Apache-2.0 / CC-BY-4.0; raw content copyright stays with the original sitesGitHub

The copyright baggage of pretraining corpora

Books3 has drawn lawsuits from rights holders (the 2023 Kafka case), and LAION has been sued by multiple artists. "Downloadable online" does not mean "free for commercial use." Research use can be more relaxed; for products, verify every source individually. For the full risk discussion, see AI Safety and Governance.

Instruction-Tuning Data ​

A pretrained model only knows how to "continue text"; instruction-tuning data teaches it to "follow instructions." This tier of data is mostly synthesized by large models or crowdsourced, and it is the easiest entry point for personal fine-tuning.

DatasetScaleUseLicenseLink
Alpaca52K instruction-response pairs (synthesized by GPT-3.5 from seed tasks)Entry-level SFT — the classic starting point for "fine-tune your own little assistant"CC BY-NC 4.0; the official copy has been taken down, community mirrors remainGitHub
ShareGPT~90K real user conversations shared by the communityConversational SFT with broad coverageNo formal license (scraped user data — evaluate carefully before any commercial use)Website
OpenAssistant (OASST1)~46K multilingual crowdsourced conversation trees (with preference annotations)The go-to open-source dual-purpose set for conversational SFT and preference dataApache-2.0Hugging Face
UltraChat~1.5M synthetic multi-turn conversationsLarge-scale conversational SFT; trained community mainstays such as ZephyrCC BY-NC 4.0Hugging Face

Preference Data (for Alignment) ​

Preference data is the fuel of alignment (RLHF and DPO): each row is a human judgment that "for the same input, answer A beats answer B." It decides whether a model is "willing" to speak in line with human values.

DatasetScaleUseLicenseLink
Anthropic HH-RLHF~160K "helpful/harmless" preference comparisonsThe de facto standard for RLHF/DPO training and alignment researchMIT (defer to official labeling)Hugging Face
OpenAI Summarize from Feedback~90K summarization preference comparisonsRLHF and preference-model training for summarization tasksAs per official releaseGitHub
Stanford SHP~380K community preferences on "which answer is better"Reward-model and DPO training (a community favorite)Academic license (defer to official terms)Hugging Face

A note on "OpenAI-related" data

The labeler comparison data used in OpenAI's InstructGPT paper (Ouyang et al., 2022) was never fully released. What researchers can actually get is mainly the Summarize from Feedback dataset above, plus third-party repackaged preference pairs. Don't claim in a paper that you used "OpenAI alignment data" without a traceable source.

Evaluation Benchmarks ​

Evaluation benchmarks are the "exam papers" used to quantify model capability. The LLM Evaluations and Benchmarks page covers methodology in depth; here are the four most-cited exams — also the names that appear most often in model launch news.

DatasetScaleUseLicenseLink
MMLU~14K multiple-choice questions across 57 subjectsThe standard test of knowledge breadth — measures "general education"MITGitHub
GSM8K8.5K grade-school math word problems (~1.3K in the test split)Math reasoning evaluationCC BY 4.0GitHub
HumanEval164 hand-written programming problems (with unit tests)Code generation evaluationMITGitHub
C-Eval13,948 Chinese questions across 52 subjectsChinese-language LLM evaluation — measures a model's "Chinese chops"CC BY-NC-SA 4.0GitHub

Benchmark saturation alert

MMLU has been pushed to around 90% by frontier models; questions everyone can answer no longer separate models. When reading reports, pay more attention to fresh benchmarks that haven't saturated yet (e.g., GPQA, LiveBench) and to evaluation sets you build from your own business scenarios. Don't stare at a single leaderboard.

Retrieval / RAG Data ​

Retrieval-augmented generation (RAG) needs "document collections + queries" data to train retrievers and evaluate retrieval quality. The hallmark of this tier is standard query-document relevance annotations, so you can compute MRR, NDCG, and similar metrics directly.

DatasetScaleUseLicenseLink
MS MARCO~1M queries + 8.8M passagesThe standard exam for passage retrieval and ranking — BERT-era retrieval models all ran on itCC BY 4.0Website
Natural Questions~300K real Google search questions (paired with Wikipedia-passage answers)Open-domain QA and long-document retrievalCC BY-SA 3.0Website
KILT5 task types across 6 datasets (ELI5, FEVER, Wizard of Wikipedia, etc.)Unified evaluation for knowledge-intensive tasks, covering QA / fact verification / dialogueLicenses vary by subset — defer to official termsGitHub

Rule of thumb

Use MS MARCO as your retrieval baseline (the community has published the most solutions for it), and turn to KILT for knowledge-intensive QA. If your business runs on Chinese-language document collections, don't forget to build your own eval set of "business docs + human-labeled relevance" — public data can never replace your domain. Related reading: Knowledge Graphs and Knowledge Injection.

Multimodal Data ​

Multimodal models need paired data — image↔text and speech↔text. The Multimodal Models page covers the principles; here are the three most-used open corpora.

DatasetScaleUseLicenseLink
LAION-5B5.85 billion image-text pairs (URLs scraped from Common Crawl)Contrastive image-text pretraining (CLIP, Stable Diffusion-style models)Annotations released with the project; image copyright stays with owners — heavily disputedLAION blog
COCO~330K annotated images, 80 object categories, 1.5M instancesDetection / segmentation / image captioning; standard multimodal evaluation kitCC BY 4.0Website
Common VoiceCrowdsourced speech in 100+ languages (as of 2025), total hours keep growingSpeech recognition (ASR) training and evaluationEach sample individually declared CC0 or CC-BYWebsite

Tool Reference ​

Frameworks ​

ToolMaintainerOne-line positioningLicenseLink
TransformersHugging FaceUnified interface for loading / fine-tuning / inference with mainstream models; the de facto standardApache-2.0Docs
LangChainLangChain teamLLM application orchestration: chains, tool calling, RAG, agent scaffoldingMITWebsite
LlamaIndexLlamaIndex teamA data-centric RAG framework: loading, indexing, and querying in one packageMITWebsite
vLLMvLLM communityHigh-throughput inference serving (PagedAttention, continuous batching); the top pick for production deploymentApache-2.0Docs
peftHugging FaceParameter-efficient fine-tuning (LoRA/QLoRA, etc.) — runs on a single consumer GPUApache-2.0Docs
trlHugging FaceReady-made implementations of SFT, DPO, PPO and other training and alignment pipelinesApache-2.0Docs

Vector Databases ​

Vector databases handle "semantic similarity" retrieval — the foundation of RAG. For the full mechanics, see Vector Databases and Semantic Search.

ToolTypeBest forLicenseLink
FAISSSingle-machine retrieval libraryPrototyping and similarity search up to the million-vector scale; plays seamlessly with Python/GPUsMITGitHub
MilvusDistributed vector databaseVery large scale, multi-replica, cloud-native production deploymentsApache-2.0Website
QdrantVector database (Rust)Fast and easy to deploy, with built-in metadata filtering and high availabilityApache-2.0Website
pgvectorPostgreSQL extensionKeep vectors inside your existing business database; transactions and retrieval managed togetherPostgreSQL LicenseGitHub

Rule of thumb

FAISS for single-machine prototypes, pgvector when vectors must live alongside your business data, Qdrant or Milvus when you're serious about a product. Don't introduce a distributed vector database for a few tens of thousands of vectors — over-engineering is the most common anti-pattern.

Annotation & Evaluation ​

ToolOne-line positioningLicenseLink
ArgillaOpen-source data annotation and feedback platform for LLMs — review-style labeling of model outputs, serving SFT and preference data constructionApache-2.0Docs
promptfooPrompt regression testing and red-teaming tool: define cases on the command line, run batch scoring, catch regressions; pairs well with prompt engineeringMITWebsite

Deployment ​

ToolOne-line positioningLicenseLink
OllamaDownload and run open models locally with a single command (built-in API and model registry); the least-fuss option for personal and prototype deploymentsMITWebsite
DockerPackage model services and environments as containers; the production standardApache-2.0Website

Selection Guide: What to Use When ​

More tools is not better — pick the smallest viable combination for the task:

Your taskRecommended comboFurther reading
Run a concept demo / try open models locallyOllama + TransformersDeployment and Inference Optimization in Practice
Fine-tune your own small modelpeft + trl + an SFT dataset (Alpaca/UltraChat)Fine-Tune Your Own LLM
Build a RAG Q&A appLlamaIndex or LangChain + FAISS/pgvector + bge embeddingsBuild a RAG App from Scratch
Ship a high-throughput inference servicevLLM + DockerDeployment and Inference Optimization in Practice
Set up evaluation and regression testingpromptfoo + benchmark datasetsSet Up LLM Evaluations
Build SFT / preference dataArgilla annotation + your own business dataAlignment: RLHF and DPO
  • A license is not free rein: CC BY-NC bans commercial use, CC BY-SA requires derivatives to stay equally open, and only Apache/MIT come close to "use it freely." The first thing to do after downloading a dataset is read the license field.
  • Scraped data is high-risk: Common Crawl, Books3, ShareGPT, LAION — for data "grabbed from the web," copyright belongs to the original content owners, and nearly every lawsuit lands here. Fine for research, use with caution in production, and for commercial use go with licensed data or a synthetic data + cleaning pipeline.
  • Don't misuse user data: ShareGPT involves real user conversations; using it directly in a commercial product violates platform terms and may also violate privacy regulations such as GDPR.
  • Self-check list before shipping a product: (1) archive data sources and licenses in writing; (2) anonymize sensitive personal information; (3) review commercial-use terms line by line (especially CC BY-NC and competition data agreements); (4) re-check periodically, because licenses can change.

Compliance is not just the legal team's job

Startup teams often push compliance onto legal, but every step of data selection is already a compliance decision. One extra day spent reading licenses can save a year of litigation. For more on risk governance, see AI Safety and Governance.

Further Reading ​

References ​

Dataset scale and license information on this page follows each project's official pages as of mid-2025; re-verify license terms and versions before you start.