Dataset · tulu-v2-sft-mixture
Tulu V2 SFT mixture
Builder: Allen Institute for AI
Availability: available · checked 2026-09-24 · source
Raw record: /data/datasets/tulu-v2-sft-mixture.json
Fields
- id
- tulu-v2-sft-mixture
- identifiers
- huggingface
- allenai/tulu-v2-sft-mixture
- builder
- Allen Institute for AI
- release_date
- 2023-11-13partial · sourcenote: HF dataset repo creation date, not the builder's own release announcement.
- availability
- available · checked 2026-09-24 · source
- content
- Card: a mix of FLAN v2 (50,000 + 50,000 CoT), OpenAssistant 1 (7,708), ShareGPT (114,046), GPT4-Alpaca (20,000, 'distilled GPT-4 data'), Code-Alpaca (20,022), LIMA (1,030), WizardLM Evol Instruct (30,000), Open-Orca (30,000 'generated by GPT-4'), hardcoded prompts (140) and science data (7,544). ODC-BY overall, with different licenses on subsets.recorded · source
- primary_sources
- record_history
- date:2026-09-24 · change:created from primary sources (tranche 2 dataset records) · by:wilson-pruitt + claude ·
Parents
Influence without weights
- distilled_from_outputs → GPT-4 (Mar. 2023) declared source Card names GPT4-Alpaca as 'distilled GPT-4 data' and 30,000 Open-Orca samples 'generated by GPT-4'. Only these subsets are asserted; ShareGPT and WizardLM subsets also hold other models' outputs, not linked here.
Children
Training data
- ← trained_on Tulu 2 7B declared source Card metadata `datasets: allenai/tulu-v2-sft-mixture`; paper: finetuned on Tulu-V2-mix. Ai2 built the mixture. The distillation inside it (GPT4-Alpaca, OpenOrca GPT-4 subset) is recorded on the dataset.
Read in
No station on the reading path has touched this record yet.