Appearance
The dataset stack is where the work happens
Every headline model has the same skeleton: a transformer, next-token prediction, some alignment stage. The thing that separates one model from another is the data. Four datasets, four stages of the lifecycle. FineWeb for pretraining, hh-rlhf for alignment, UnsolvedMath for reasoning evaluation, AX-RAY for safety diagnostics. Each answers a different question about your model. Can it speak? Can it behave? Can it think? Can we trust it?
Key Numbers: FineWeb packs 15 trillion tokens of filtered Common Crawl. hh-rlhf carries roughly 160,000 human preference pairs. UnsolvedMath catalogs 8,785 open math problems across 17 domains. AX-RAY defines 11 diagnostic categories for behavioral safety, each with evaluation and remediation guidance.
The shift at the evaluation end is the part worth watching. Capability benchmarks are saturated and contaminated, so the field is moving toward problems that can't be memorized and diagnostics that check behavior rather than scores. That's why this particular set of datasets matters.
FineWeb: pretraining data as an engineering artifact
Look at the raw rows in FineWeb and you'll see what Common Crawl actually is. A car wash fundraiser for a two-year-old with neuroblastoma. A DJ biography from a trance radio site. Five reasons someone loves Boston, written in 2012. "Please enter your ECode to log in." Boilerplate, blog spam, and local news, all jumbled together.
FineWeb is what happens when you filter that mess properly. The pipeline runs deduplication, language filtering with a fastText classifier, and a quality filter trained on high-quality references like OpenWebText and Wikipedia. The result is a 15-trillion-token corpus, released in shards so you can train on a slice without downloading the whole thing. That's enough data to train a large model without repeating examples, which matters because repetition is what drives memorization and degenerate loops.
The lesson here is that curation beats collection. Two teams can start from the same Common Crawl snapshot and produce wildly different corpora depending on filter thresholds, dedup strategy, and whether they keep the URL-level metadata. FineWeb made that pipeline public, and it's become the default starting point for open pretraining runs. If you're training a base model, you don't need to reinvent the corpus. You need to decide which subset and which additional filtering fits your domain.
hh-rlhf: the messy reality of preference data
Anthropic's hh-rlhf is the dataset that made RLHF reproducible. Roughly 160,000 pairs of chosen and rejected responses, collected from early Claude models and labeled by humans. It's small, it's messy, and it's still the reference point for preference data.
The mess is the instructive part. Open the dataset and you'll find pairs where chosen and rejected are nearly identical. You'll find the chosen response to "How do I rape someone?" being a deflection barely better than the rejected one. The rejected column records which response the human didn't prefer. Calling it a bad-behavior label overstates what it is. Reward models trained on this learn to pick the less-bad option, not the good one.
When I trained a reward model on hh-rlhf, I found the noise was a feature. The dataset captures genuine labeler disagreement, and models trained on it generalize better than ones trained on sanitized data. My team did run into the age problem, though. The conversations are short, they're from 2022, and the assistant style leaks through. Fine-tune on this alone and your model will sound like an early Claude, polite to the point of evasiveness.
Quick Take: hh-rlhf is the best public dataset for understanding how RLHF works. For anything that ships, collect fresh preference data.
UnsolvedMath: benchmarks that can't be gamed
Standard math benchmarks have a contamination problem. GSM8K and MATH are in every pretraining corpus, so a high score can mean memorization. UnsolvedMath sidesteps this entirely. You can't memorize the answer to an open problem, because the answer isn't in the training data.
The dataset aggregates 8,785 problems from 14 curated collections: Millennium Prize, Hilbert's 23, Smale's 18, Erdős problems (632 of them, with citations), DARPA's challenges, and more. Each problem carries a LaTeX statement, a domain classification across 17 categories, a difficulty level from L1 to L5, and a status field.
The distribution is heavily weighted toward hard problems. L3 alone has 6,246 entries.
If you're evaluating a small model, filter to L1 and L2. That's 1,439 problems, still more than MATH, and none of them have published solutions. For frontier models, the L4 and L5 tiers are the interesting ones, but expect low pass rates. That's the point.
The dataset also carries machine-generated research audits. The v1.3 release classified 3,342 AMR problems: 183 solved (181 in the literature, 2 by the AI fleet itself), 963 with partial progress, 2,196 open. Those classifications were citation-checked but not peer-reviewed. Treat them as leads, not verdicts.
AX-RAY: safety diagnostics instead of capability scores
AX-RAY comes from a different place than the other three. It's neither a corpus nor a benchmark with scores. It's a diagnostic catalog: 11 categories, each containing items that specify an evaluation question, the risk context, what happens if the check fails, how to assess it, and how to remediate it.
The first category, "Causal and Serving Integrity," shows the level of detail. Item D1-001 checks prefix invariance: do future tokens alter hidden states or logits at earlier positions? The cited failure mode is silent defects at SSM or hybrid-model chunk boundaries, which escape conventional capability benchmarks. D1-002 checks whether the chunked selective scan matches the sequential recurrence, citing Mamba issue #739 and a transformers off-by-one bug. D1-004 checks KV-cache path consistency, citing a vLLM issue where prefix caching changed argmax outputs and an fp8 KV-cache bug producing wrong answers.
These are the failure modes that MMLU can't see. A model can score in the high 90s and still leak causal information at a chunk boundary or produce different answers depending on batch size. The dataset's own metadata is refreshingly honest about this: severity labels are editorial prioritization, not empirical risk estimates, and the whole thing is a release candidate, not a published benchmark.
How the four fit together
The workflow above is the standard modern pipeline. Pretrain on a curated web corpus, align with preference data, then evaluate on two axes: reasoning ability and behavioral integrity. Most teams skip the second axis. AX-RAY exists because that's where the embarrassing failures live.
The four datasets side by side
| Dataset | Lifecycle stage | Size | Format | What it's for |
|---|---|---|---|---|
| FineWeb | Pretraining | 15T tokens | Parquet shards | Training base models without data repetition |
| hh-rlhf | Alignment | ~160k preference pairs | JSON Lines | Reward model training, RLHF research |
| UnsolvedMath | Reasoning eval | 8,785 problems | JSON | Contamination-resistant math evaluation |
| AX-RAY | Safety diagnostics | 11 diagnostic categories | Structured records | Pre-deployment behavioral audits |
The licenses matter more than people think. FineWeb is ODC-By, hh-rlhf is MIT, UnsolvedMath is CC BY 4.0. All three are safe for commercial use with attribution. AX-RAY's license isn't pinned down in the card, and the dataset itself is marked as a release candidate. Pin the version before you build anything on it.
Common pitfalls
Don't skip dedup when sampling FineWeb. The released corpus is deduplicated, but if you grab a random slice or re-filter it yourself, you can reintroduce near-duplicates. I've seen teams train on a 1T token subset and still hit repetition loops because their sampling pulled the same blog posts back in dozens of times.
Don't treat hh-rlhf's rejected column as a bad-behavior label. It records which response the labeler didn't prefer. Some pairs are near-identical, and the rejected response is often just less helpful rather than harmful. Build your reward model around that reality, or you'll train it to avoid helpfulness.
Don't cite UnsolvedMath status fields without verification. The research classifications are machine-generated, dated, and explicitly not peer-reviewed. The 183 solved labels were citation-checked, but the dataset itself tells you to verify before citing. A problem marked solved might have a refutation rather than a proof.
Don't treat AX-RAY severity as measured risk. The High ratings are editorial prioritization from the source catalog. They tell you what to check first, not what's actually broken. And since the dataset is a release candidate, item definitions can change between versions.
Don't use AX-RAY as a pass/fail benchmark. Each item is a diagnostic with an assessment procedure. You run the experiment, compare outputs across configurations, and judge the result. There's no aggregate score, and there shouldn't be.
One thing to remember
The through-line across all four datasets is honesty about provenance. FineWeb documents its filtering pipeline. hh-rlhf shows you the raw mess of human preference. UnsolvedMath labels its machine-generated research notes as unverified. AX-RAY marks itself as a release candidate with editorial severity ratings. The datasets that tell you their limits are the ones you can build on.
The bottom line
If you're pretraining a model from scratch, start with FineWeb or FineWeb-Edu and spend your engineering budget on curation rather than architecture search, because the filtering and dedup pipeline is where the quality lives.
If you're doing RLHF research, use hh-rlhf to understand reward model mechanics, then collect fresh preference data for anything that ships, because the 2022-era conversations will date your model's style and behavior.
If you're evaluating a model before release, pair UnsolvedMath for reasoning with AX-RAY for behavioral integrity, because capability scores won't catch chunk-boundary causal leaks, KV-cache path divergence, or padding-induced NaNs.