Dataset · the-pile
The Pile
Builder: EleutherAI
Availability: unknown · checked 2026-09-24 · source
Homepage says 'The Pile is hosted by the Eye'; the-eye.eu did not respond when checked, and the HF repo EleutherAI/pile holds only a loader script pointing there. Reports that the dataset was taken down were not verified in-session.
Raw record: /data/datasets/the-pile.json
Fields
- id
- the-pile
- identifiers
- huggingface
- EleutherAI/pileEleutherAI/the_pilethe_pile
- builder
- EleutherAI
- release_date
- 2020-12-31partial · sourcenote: arXiv v1 date of the paper.
- availability
- unknown · checked 2026-09-24 · sourcenote: Homepage says 'The Pile is hosted by the Eye'; the-eye.eu did not respond when checked, and the HF repo EleutherAI/pile holds only a loader script pointing there. Reports that the dataset was taken down were not verified in-session.
- content
- 825 GiB of English text from 22 sources, built for training large language models.recorded · source
- primary_sources
- record_history
- date:2026-09-24 · change:created from primary sources (dataset records ruling) · by:wilson-pruitt + claude ·
Parents
No edges recorded.
Children
Training data
- ← trained_on GPT-Neo 2.7B declared source 420B tokens.
- ← trained_on GPT-J 6B declared source 402B tokens.
- ← trained_on GPT-NeoX-20B declared source
- ← trained_on Pythia 6.9B declared source 300B tokens; a twin model was trained on the deduplicated Pile.
- ← trained_on OPT 6.7B declared source A subset of the Pile, alongside RoBERTa corpus data and PushShift.io Reddit (not recorded as datasets).
- ← trained_on gpt-neo-1.3B declared_by_uploader source Card metadata datasets: EleutherAI/pile. Consistent with the gpt-neo-2-7b record, same family/release.
- ← trained_on gpt-neo-125m declared_by_uploader source Card metadata datasets: EleutherAI/pile. Consistent with the gpt-neo-2-7b record, same family/release.
- ← trained_on opt-1.3b declared source Paper: trained on ~180B tokens incl. 'a subset of the Pile'; also RoBERTa corpus subsets and PushShift.io Reddit (not recorded as datasets). Same mix as opt-6-7b -- not stated per-size in the card, but the paper describes one training recipe for the whole OPT suite.
- ← trained_on opt-125m declared source Paper: trained on ~180B tokens incl. 'a subset of the Pile'; also RoBERTa corpus subsets and PushShift.io Reddit (not recorded as datasets). Same mix as opt-6-7b -- not stated per-size in the card, but the paper describes one training recipe for the whole OPT suite.
- ← trained_on opt-13b declared source Paper: trained on ~180B tokens incl. 'a subset of the Pile'; also RoBERTa corpus subsets and PushShift.io Reddit (not recorded as datasets). Same mix as opt-6-7b -- not stated per-size in the card, but the paper describes one training recipe for the whole OPT suite.
- ← trained_on opt-30b declared source Paper: trained on ~180B tokens incl. 'a subset of the Pile'; also RoBERTa corpus subsets and PushShift.io Reddit (not recorded as datasets). Same mix as opt-6-7b -- not stated per-size in the card, but the paper describes one training recipe for the whole OPT suite.
- ← trained_on opt-66b declared source Paper: trained on ~180B tokens incl. 'a subset of the Pile'; also RoBERTa corpus subsets and PushShift.io Reddit (not recorded as datasets). Same mix as opt-6-7b -- not stated per-size in the card, but the paper describes one training recipe for the whole OPT suite.
- ← trained_on pythia-1.4b declared_by_uploader source HF card metadata datasets: EleutherAI/the_pile. Often incomplete; confirm which training stage used it.
- ← trained_on pythia-12b declared_by_uploader source HF card metadata datasets: EleutherAI/pile. Often incomplete; confirm which training stage used it.
- ← trained_on pythia-160m declared_by_uploader source HF card metadata datasets: EleutherAI/pile. Often incomplete; confirm which training stage used it.
- ← trained_on pythia-1b declared_by_uploader source HF card metadata datasets: the_pile. Often incomplete; confirm which training stage used it.
- ← trained_on pythia-2.8b declared_by_uploader source HF card metadata datasets: EleutherAI/pile. Often incomplete; confirm which training stage used it.
- ← trained_on pythia-410m declared_by_uploader source HF card metadata datasets: EleutherAI/pile. Often incomplete; confirm which training stage used it.
- ← trained_on pythia-70m declared_by_uploader source HF card metadata datasets: EleutherAI/pile. Often incomplete; confirm which training stage used it.
- ← trained_on Cerebras-GPT 13B declared source Paper abstract and sec. 2.2: trained on the Pile with the provided train/test/validation splits; no deduplication.
- ← trained_on GPT-JT-6B-v1 declared source Card: 2.62B tokens with UL2 loss on the Pile, and 55% of the second-stage mix.
Read in
No station on the reading path has touched this record yet.