Dataset · starcoderdata
StarCoderData
Builder: BigCode
Availability: gated · checked 2026-09-24 · source
HF API reports gated: auto (approval step); the card text could not be read without accepting terms.
Raw record: /data/datasets/starcoderdata.json
Fields
- id
- starcoderdata
- identifiers
- huggingface
- bigcode/starcoderdata
- builder
- BigCode
- release_date
- 2023-03-30partial · sourcenote: HF dataset repo creation date, not the builder's own release announcement.
- availability
- gated · checked 2026-09-24 · sourcenote: HF API reports gated: auto (approval step); the card text could not be read without accepting terms.
- content
- TinyLlama paper: 'the training data of StarCoder ... contains code data in 86 programming languages. In addition to code, it also includes GitHub issues and text-code pairs that involve natural languages.' StarCoder paper: StarCoderBase trained on 1 trillion tokens from The Stack v1.2 (permissively licensed code, 44 opt-outs at processing time).partial · sourcenote: Primary card was gated; description rests on the TinyLlama and StarCoder papers.
- primary_sources
- record_history
- date:2026-09-24 · change:created from primary sources (tranche 2 dataset records) · by:wilson-pruitt + claude ·
Parents
No edges recorded.
Children
Training data
- ← trained_on StarCoderBase declared source Dataset card (public API description): 'This is the dataset used for training StarCoder and StarCoderBase. It contains 783GB of code in 86 programming languages ... approximately 250 Billion tokens.' NOTE the 250B-token figure differs from the 1T tokens seen in training (multiple epochs per the paper); the dataset file itself is gated.
- ← trained_on TinyLlama-1.1B (intermediate-step-1431k-3T) declared source Card metadata `datasets: bigcode/starcoderdata`; paper: 'the training data of StarCoder', code-related samples only. `starcoderdata` is gated.
- ← trained_on StarCoder declared source Dataset card (public API description): 'the dataset used for training StarCoder and StarCoderBase'. StarCoder's own fine-tuning used the Python subset of that training data (paper sec. 5.6).
Read in
No station on the reading path has touched this record yet.