Model · orca-2-7b
Orca 2 7B
Developer: Microsoft Research
Availability: available · checked 2026-09-24 · source
Raw record: /data/models/orca-2-7b.json
Fields
- id
- orca-2-7b
- identifiers
- huggingface
- microsoft/Orca-2-7b
- developer
- Microsoft Research
- release_date
- 2023-11-18recorded · sourcenote: arXiv v1 submission date; the abstract announces the weights ('We make Orca 2 weights publicly available at aka.ms/orca-lm'). The HF repo's first commit is 2023-11-21 (HF API); repo creation 2023-11-14.
- weights_status
- open
- availability
- available · checked 2026-09-24 · source
- license
- Microsoft Research Licenserecorded · sourcenote: Card: 'Orca 2 is licensed under the Microsoft Research License.' It adds: 'Llama 2 is licensed under the LLAMA 2 Community License', which applies to the base.
- architecture
- family
- decoder_only
- note
- family read from config 'architectures': ['LlamaForCausalLM'].
- n_layers
- 32recorded · source
- hidden_size
- 4096recorded · source
- n_heads
- 32recorded · source
- vocab_size
- 32003recorded · source
- context_length
- 4096recorded · source
- n_kv_heads
- 32recorded · source
- positional_encoding
- rotary (RoPE)recorded · sourcenote: Propagated from llama-2-7b: the card says 'Please refer to LLaMA-2 technical report for details on the model architecture' and the config matches.
- training_data
- Synthetic data created to improve small-model reasoning. Paper (progressive learning): start from the LLaMA-2-7B checkpoint; fine-tune on the FLAN-v2 train split (1 epoch); then on 5 million ChatGPT examples from Orca 1 (3 epochs); then on 1 million GPT-4 examples from Orca 1 plus Orca 2's 817K examples (4 epochs). Card: all synthetic data moderated with Azure content filters.recorded · sourcenote: Paper section 4.2, 'Progressive Learning'. Loss is computed only on teacher-generated tokens.
- techniques
- primary_sources
- record_history
- date:2026-09-24 · change:ingested as candidate from HF (microsoft/Orca-2-7b@60e31e6bdcf582ad103b807cb74b73ee1d2c4b17) · by:ingest_hf.py ·date:2026-09-24 · change:preparer: developer, release date (arXiv), license, positional encoding (propagated from llama-2-7b), training data, 1 edge accepted + 2 PANEL-NEEDED; sources: card, arXiv 2311.11045 · by:claude (preparer, Sonnet 5) ·date:2026-09-24 · change:ruling: 3-tier panel (2/3 dataset routing), see session log: 2 direct edges rejected, 3 trained_on edges added · by:wilson-pruitt + claude (Sonnet 5) ·date:2026-09-24 · change:reviewed and promoted from staging (4 edge(s) accepted) · by:Wilson Pruitt ·
Parents
Weights descend
- fine_tuned_from → Llama 2 7B declared source Paper 4.2: 'We start with LLaMA-2-7B or LLaMA-2-13B checkpoint and finetune it'; card: 'Orca 2 is a finetuned version of LLAMA-2'.
Training data
- trained_on → FLAN-v2 Collection (Flan 2022) declared source Orca 2 paper 4.2: 'We start with LLaMA-2-7B or LLaMA-2-13B checkpoint and finetune it on the train split of FLAN-v2 dataset for one epoch.'
- trained_on → Orca 1 explanation-tuning data (FLAN-5M / FLAN-1M) declared source Orca 2 paper 4.2: 'We then train on 5 million ChatGPT data from Orca 1 for 3 epochs. Then we train on the combination of 1 million GPT-4 data from Orca 1 and Orca 2's 817K data for 4 epochs.'
- trained_on → Orca 2 dataset (~817K) declared source Same passage: 'Orca 2's 817K data'; section 4: 'we created a new dataset with ~817K training instances, which we will refer as Orca 2 dataset.'
Children
No edges recorded.
Read in
No station on the reading path has touched this record yet.