Model · starling-rm-7b-alpha
Starling-RM-7B-alpha
Developer: Banghua Zhu, Evan Frick, Tianhao Wu, Hanlin Zhu and Jiantao Jiao (Berkeley NEST; card 'Developed by')
Availability: available · checked 2026-09-24 · source
Raw record: /data/models/starling-rm-7b-alpha.json
Fields
- id
- starling-rm-7b-alpha
- identifiers
- huggingface
- berkeley-nest/Starling-RM-7B-alpha
- developer
- Banghua Zhu, Evan Frick, Tianhao Wu, Hanlin Zhu and Jiantao Jiao (Berkeley NEST; card 'Developed by')
- release_date
- 2023-11-25partial · sourcenote: HF repo creation date; may predate or postdate public release.
- weights_status
- open
- availability
- available · checked 2026-09-24 · source
- license
- Apache-2.0 per card metadata, 'under the condition that the model is not used to compete with OpenAI'partial · sourcenote: Card: 'License: Apache-2.0 license under the condition that the model is not used to compete with OpenAI'. The policy card's License section adds non-commercial and distillation-licence language; this card was not seen to say so. Reviewer confirm.
- architecture
- family
- decoder_only
- n_layers
- 32recorded · source
- hidden_size
- 4096recorded · source
- n_heads
- 32recorded · source
- vocab_size
- 32000recorded · source
- context_length
- 4096recorded · source
- n_kv_heads
- 32recorded · source
- positional_encoding
- nullnot_recorded
- note
- Card: 'we remove the last layer of Llama2-7B Chat, and concatenate a linear layer that outputs scalar for any pair of input prompt and response.' So a Llama 2 decoder backbone with a scalar reward head in place of the language-model head; HF config 'architectures' did not resolve to a family, and the head is not counted in the dims below.
- training_data
- Nectar (GPT-4-ranked 7-wise comparisons), with the K-wise maximum likelihood estimator.recorded · sourcenote: Card: 'We train the reward model with preference dataset berkeley-nest/Nectar, with the K-wise maximum likelihood estimator proposed in [this paper]'; 'since the preference dataset ... is based on GPT-4 preference, the reward model is likely to be biased towards GPT-4's own preference'.
- techniques
- primary_sources
- record_history
- date:2026-09-24 · change:ingested as candidate from HF (berkeley-nest/Starling-RM-7B-alpha@6c6b4d5627834fe010d2c001632de2b94db81d66) · by:ingest_hf.py ·date:2026-09-24 · change:ruling: 3-tier panel (3/3, feedback_from via reward-model record), see session log: record created as the reward-model node; developer, license, architecture family, training data, 2 edges filled from card and blog · by:wilson-pruitt + claude (Sonnet 5) ·date:2026-09-24 · change:reviewed and promoted from staging (2 edge(s) accepted) · by:Wilson Pruitt ·
Parents
Weights descend
- fine_tuned_from → Llama 2-Chat 7B declared source Card: 'Starling-RM-7B-alpha is a reward model trained from Llama2-7B-Chat'; 'Finetuned from model: Llama2-7B-Chat'. Blog: 'Our reward model is fine-tuned from Llama2-7B-Chat'. The final layer is replaced by a scalar head, so weights below it descend.
Training data
- trained_on → Nectar declared source Card: 'We train the reward model with preference dataset berkeley-nest/Nectar'. Blog: 'trained with our K-wise loss on the Nectar dataset'.
Children
Influence without weights
- ← feedback_from Starling-LM-7B-alpha declared source Card: 'a language model trained from Openchat 3.5 with reward model berkeley-nest/Starling-RM-7B-alpha and policy optimization method APA'. Blog: 'we fine-tuned the Openchat 3.5 language model using the learned reward model.' Developers' own card and blog; the uploader IS the developer. Parent is a model whose scores, not text, shaped the policy. PARENT starling-rm-7b-alpha IS STAGED (promote it first; it in turn needs nothing unpromoted except openchat-3-5 for this record).
Read in
No station on the reading path has touched this record yet.