Dataset · orca-2-data
Orca 2 dataset (~817K)
Builder: Microsoft Research (Orca 2 authors)
Availability: unknown · checked 2026-09-24 · source
Neither paper says the training data was released: Orca 1 announces only a weight diff ('publicly release a diff of the model weights'), Orca 2 says 'We make Orca 2 weights publicly available'. A search of the Hugging Face microsoft org for 'orca' on 2026-09-24 returned only microsoft/orca-math-word-problems-200k and microsoft/orca-agentinstruct-1M-v1, neither of which is described as this data. Absence from a search is not proof of non-release, so recorded unknown, not never_released. OpenOrca (id openorca) is a third-party reproduction of Orca 1's recipe, not this data.
Raw record: /data/datasets/orca-2-data.json
Fields
- id
- orca-2-data
- builder
- Microsoft Research (Orca 2 authors)
- release_date
- 2023-11-18partial · sourcenote: arXiv v1 date of the Orca 2 paper; the data itself has no separate release date.
- availability
- unknown · checked 2026-09-24 · sourcenote: Neither paper says the training data was released: Orca 1 announces only a weight diff ('publicly release a diff of the model weights'), Orca 2 says 'We make Orca 2 weights publicly available'. A search of the Hugging Face microsoft org for 'orca' on 2026-09-24 returned only microsoft/orca-math-word-problems-200k and microsoft/orca-agentinstruct-1M-v1, neither of which is described as this data. Absence from a search is not proof of non-release, so recorded unknown, not never_released. OpenOrca (id openorca) is a third-party reproduction of Orca 1's recipe, not this data.
- content
- 'For Orca 2, we created a new dataset with ~817K training instances, which we will refer as Orca 2 dataset.' Four main sources: FLAN-v2-derived cautious-reasoning prompts (~602K), a few-shot set (55K, re-purposed from Orca 1 data), math (~160K), and fully synthetic doctor-patient conversations ('synthetically created 2000 Doctor-Patient Conversations with GPT-4'). Section 4.2 calls the combined Orca 1 GPT-4 set plus this one '~1.8 million GPT-4 data'.recorded · sourcenote: The paper's own labelling of this data as GPT-4 data is at section 4.2 (compute paragraph) and the progressive-learning paragraph; it does not give a per-source teacher breakdown for the four sources.
- primary_sources
- record_history
- date:2026-09-24 · change:created from primary sources (ruling: 3-tier panel (2/3 dataset routing), see session log) · by:wilson-pruitt + claude (Sonnet 5) ·
Parents
Influence without weights
- distilled_from_outputs → GPT-4 (Mar. 2023) declared source Orca 2 paper section 4.2: 'training on ~1.8 million GPT-4 data' = Orca 1's 1M GPT-4 data plus Orca 2's 817K. The paper names GPT-4 for the doctor-patient conversations explicitly; for the other three sources it gives no separate teacher line, so this edge rests on the section 4.2 labelling. Tag `declared` is the paper's aggregate label ('~1.8 million GPT-4 data'), not a per-source statement: only the doctor-patient portion is attributed to GPT-4 by name, so for the rest GPT-4 authorship is the paper's own summary, not separately sourced (Wilson's ruling, 2026-09-24: keep declared, note the limit).
Children
Training data
- ← trained_on Orca 2 7B declared source Same passage: 'Orca 2's 817K data'; section 4: 'we created a new dataset with ~817K training instances, which we will refer as Orca 2 dataset.'
Read in
No station on the reading path has touched this record yet.