Dataset · alpaca-52k
Alpaca instruction data (52K)
Builder: Stanford CRFM (Tatsu Lab)
Availability: available · checked 2026-09-24 · source
Raw record: /data/datasets/alpaca-52k.json
Fields
- id
- alpaca-52k
- identifiers
- huggingface
- tatsu-lab/alpaca
- builder
- Stanford CRFM (Tatsu Lab)
- release_date
- 2023-03-13recorded · source
- availability
- available · checked 2026-09-24 · source
- content
- 52K instruction-following examples generated self-instruct style with text-davinci-003.recorded · source
- primary_sources
- record_history
- date:2026-09-24 · change:created from primary sources (dataset records ruling) · by:wilson-pruitt + claude ·
Parents
Influence without weights
- distilled_from_outputs → text-davinci-003 declared source Stanford CRFM generated the data itself with text-davinci-003.
Children
Training data
- ← trained_on Alpaca 7B declared source Fine-tuning data.
- ← trained_on Dolly v1 6B declared source Card: 'fine-tuned on a ~52K record instruction corpus (Stanford Alpaca)'; blog: 'data from Alpaca'. The Alpaca set is text-davinci-003 output (see the `alpaca-52k` dataset edge), so the closed-model path runs through it.
- ← trained_on MPT-7B-Chat declared source Archived card lists Alpaca (link tatsu-lab/alpaca).
- ← trained_on StableLM-Tuned-Alpha-7B declared source Card training-dataset section lists Alpaca first; uploader is the developer.
Read in
No station on the reading path has touched this record yet.