Stemma Machinarum

Model · orca-2-7b

Orca 2 7B

Developer: Microsoft Research
Availability: available · checked 2026-09-24 · source

Raw record: /data/models/orca-2-7b.json

Fields

id
orca-2-7b
identifiers
huggingface
microsoft/Orca-2-7b
developer
Microsoft Research
release_date
2023-11-18recorded · source
note:  arXiv v1 submission date; the abstract announces the weights ('We make Orca 2 weights publicly available at aka.ms/orca-lm'). The HF repo's first commit is 2023-11-21 (HF API); repo creation 2023-11-14.
weights_status
open
availability
available · checked 2026-09-24 · source
license
Microsoft Research Licenserecorded · source
note:  Card: 'Orca 2 is licensed under the Microsoft Research License.' It adds: 'Llama 2 is licensed under the LLAMA 2 Community License', which applies to the base.
architecture
family
decoder_only
note
family read from config 'architectures': ['LlamaForCausalLM'].
n_layers
32recorded · source
hidden_size
4096recorded · source
n_heads
32recorded · source
vocab_size
32003recorded · source
context_length
4096recorded · source
n_kv_heads
32recorded · source
positional_encoding
rotary (RoPE)recorded · source
note:  Propagated from llama-2-7b: the card says 'Please refer to LLaMA-2 technical report for details on the model architecture' and the config matches.
training_data
Synthetic data created to improve small-model reasoning. Paper (progressive learning): start from the LLaMA-2-7B checkpoint; fine-tune on the FLAN-v2 train split (1 epoch); then on 5 million ChatGPT examples from Orca 1 (3 epochs); then on 1 million GPT-4 examples from Orca 1 plus Orca 2's 817K examples (4 epochs). Card: all synthetic data moderated with Azure content filters.recorded · source
note:  Paper section 4.2, 'Progressive Learning'. Loss is computed only on teacher-generated tokens.
techniques
primary_sources
https://huggingface.co/microsoft/Orca-2-7b
https://huggingface.co/microsoft/Orca-2-7b/raw/60e31e6bdcf582ad103b807cb74b73ee1d2c4b17/config.json
https://huggingface.co/microsoft/Orca-2-7b/blob/60e31e6bdcf582ad103b807cb74b73ee1d2c4b17/README.md
https://arxiv.org/abs/2311.11045
record_history
date:2026-09-24 · change:ingested as candidate from HF (microsoft/Orca-2-7b@60e31e6bdcf582ad103b807cb74b73ee1d2c4b17) · by:ingest_hf.py ·
date:2026-09-24 · change:preparer: developer, release date (arXiv), license, positional encoding (propagated from llama-2-7b), training data, 1 edge accepted + 2 PANEL-NEEDED; sources: card, arXiv 2311.11045 · by:claude (preparer, Sonnet 5) ·
date:2026-09-24 · change:ruling: 3-tier panel (2/3 dataset routing), see session log: 2 direct edges rejected, 3 trained_on edges added · by:wilson-pruitt + claude (Sonnet 5) ·
date:2026-09-24 · change:reviewed and promoted from staging (4 edge(s) accepted) · by:Wilson Pruitt ·

Parents

Weights descend

Training data

Children

No edges recorded.

Read in

No station on the reading path has touched this record yet.