Model · tulu-2-dpo-7b
Tulu 2 DPO 7B
Developer: Allen Institute for AI (Ai2) and University of Washington (paper authors Ivison, Wang, et al.)
Availability: available · checked 2026-09-24 · source
Raw record: /data/models/tulu-2-dpo-7b.json
Fields
- id
- tulu-2-dpo-7b
- identifiers
- huggingface
- allenai/tulu-2-dpo-7b
- developer
- Allen Institute for AI (Ai2) and University of Washington (paper authors Ivison, Wang, et al.)
- release_date
- 2023-11-13partial · sourcenote: HF repo creation date (2023-11-13); paper arXiv 2311.10702 dated Nov 2023 (v2 20 Nov).
- weights_status
- open
- availability
- available · checked 2026-09-24 · source
- license
- AI2 ImpACT Low-risk license (card); base weights also subject to the Llama 2 Community Licensepartial · sourcenote: Card metadata: `license: other`, `license_name: ai2-impact-license-low-risk`, link allenai.org/impact-license; card text 'License: AI2 ImpACT Low-risk license'. License page text could not be read this session (JavaScript-rendered).
- architecture
- family
- decoder_only
- note
- family read from config 'architectures': ['LlamaForCausalLM'].
- n_layers
- 32recorded · source
- hidden_size
- 4096recorded · source
- n_heads
- 32recorded · source
- vocab_size
- 32000recorded · source
- context_length
- 8192recorded · sourcenote: Config value; see tulu-2-7b.
- n_kv_heads
- 32recorded · source
- positional_encoding
- rotary (RoPE)recorded · sourcenote: Architecture unchanged from llama-2-7b (fine-tune, not a structural change).
- training_data
- Tulu V2 mix SFT (see tulu-2-7b) followed by DPO on a filtered, binarized UltraFeedback (HuggingFaceH4/ultrafeedback_binarized), following the Zephyr-Beta recipe.recorded · sourcenote: Paper (arXiv 2311.10702, section 2.2): 'For DPO training, we follow the Zephyr-Beta approach: we train on a filtered and binarized form of UltraFeedback for three epochs', learning rate 5e-7. Card: 'DPO Recipe: The DPO recipe is from the Zephyr Beta model'.
- techniques
- primary_sources
- record_history
- date:2026-09-24 · change:ingested as candidate from HF (allenai/tulu-2-dpo-7b@b57ef95260b6d4e726adf64518af038e5673f126) · by:ingest_hf.py ·date:2026-09-24 · change:preparer: developer, license (partial), training data, positional encoding (propagated), 4 edges (fine_tuned_from tulu-2-7b, trained_on ultrafeedback; 2 rejected); sources: card, arXiv 2311.10702 · by:claude (preparer, Sonnet 5) ·date:2026-09-24 · change:reviewed and promoted from staging (2 edge(s) accepted) · by:Wilson Pruitt ·
Parents
Weights descend
- fine_tuned_from → Tulu 2 7B declared source Paper abstract: 'a LLaMA-2 70B model finetuned on Tulu-V2-mix and further trained using direct preference optimization (DPO)'; Table 3 compares Tulu V2 models 'with and without DPO finetuning' at 7B, 13B, 70B. The paper states this for the family, not in a 7B-specific sentence, so the 7B DPO checkpoint's start point is read from that framing. Card metadata instead names only Llama-2-7b-hf as base_model (the original base, not the SFT stage).
Training data
- trained_on → UltraFeedback declared source Card metadata `datasets: HuggingFaceH4/ultrafeedback_binarized` (a filtered, binarized UltraFeedback); paper: DPO on 'a filtered and binarized form of UltraFeedback'. Recorded like Zephyr: trained_on the UltraFeedback record, which itself carries feedback_from gpt-4. Tulu's own paper notes UltraFeedback used TruthfulQA prompts (contamination caveat for evaluation).
Children
No edges recorded.
Read in
No station on the reading path has touched this record yet.