Dataset · alpaca-cleaned
Alpaca-Cleaned
Builder: gururise / yahma (community)
Availability: available · checked 2026-09-24 · source
Raw record: /data/datasets/alpaca-cleaned.json
Fields
- id
- alpaca-cleaned
- identifiers
- huggingface
- yahma/alpaca-cleaned
- builder
- gururise / yahma (community)
- release_date
- 2023-03-24partial · sourcenote: HF dataset repo creation date, not the builder's own release announcement.
- availability
- available · checked 2026-09-24 · source
- content
- Card: 'a cleaned version of the original Alpaca Dataset released by Stanford', fixing hallucinations, empty outputs and similar issues. Derivation from `alpaca-52k` is recorded in prose (no dataset-to-dataset edge). The card refers to the original outputs as GPT3's.recorded · source
- primary_sources
- record_history
- date:2026-09-24 · change:created from primary sources (tranche 2 dataset records) · by:wilson-pruitt + claude ·
Parents
Influence without weights
- distilled_from_outputs → text-davinci-003 declared source Inherited through Alpaca: Stanford CRFM generated the original 52K with text-davinci-003; the cleaner's card describes this set as a cleaned version of it. Cleaning changed some outputs; the fraction is not_recorded.
Children
Training data
- ← trained_on alpaca-lora-7b declared source Card metadata: datasets: [yahma/alpaca-cleaned]; README training command --data_path 'yahma/alpaca-cleaned'. Uploader tloen is the developer. Card prose instead says 'Stanford Alpaca dataset'; see training_data note. Reviewer: if you prefer the prose reading, swap to alpaca-52k.
Read in
No station on the reading path has touched this record yet.