Model · starling-lm-7b-alpha
Starling-LM-7B-alpha
Developer: Banghua Zhu, Evan Frick, Tianhao Wu, Hanlin Zhu and Jiantao Jiao (Berkeley NEST; card 'Developed by')
Availability: available · checked 2026-09-24 · source
Raw record: /data/models/starling-lm-7b-alpha.json
Fields
- id
- starling-lm-7b-alpha
- identifiers
- huggingface
- berkeley-nest/Starling-LM-7B-alpha
- developer
- Banghua Zhu, Evan Frick, Tianhao Wu, Hanlin Zhu and Jiantao Jiao (Berkeley NEST; card 'Developed by')
- release_date
- 2023-11-25partial · sourcenote: HF repo creation date; may predate or postdate public release.
- weights_status
- open
- availability
- available · checked 2026-09-24 · source
- license
- Apache-2.0 per card metadata and model summary, 'under the condition that the model is not used to compete with OpenAI'; the card's License section also says research preview, non-commercial, subject to LLaMA/OpenAI/ShareGPT termspartial · sourcenote: The card contradicts itself: metadata `license: apache-2.0` and 'License: Apache-2.0 license under the condition that the model is not used to compete with OpenAI', but the later License section: 'research preview intended for non-commercial use only, subject to the data distillation License of LLaMA, Terms of Use of the data generated by OpenAI, and Privacy Practices of ShareGPT'. Not reconciled; reviewer check.
- architecture
- family
- decoder_only
- note
- family read from config 'architectures': ['MistralForCausalLM'].
- n_layers
- 32recorded · source
- hidden_size
- 4096recorded · source
- n_heads
- 32recorded · source
- vocab_size
- 32002recorded · source
- context_length
- 8192recorded · source
- n_kv_heads
- 8recorded · source
- positional_encoding
- rotary (RoPE)recorded · sourcenote: Architecture unchanged from mistral-7b-v0-1 (fine-tune, not a structural change).
- training_data
- Nectar (GPT-4-ranked 7-wise comparisons) used to train the reward model Starling-RM-7B-alpha (fine-tuned from Llama2-7B-Chat); the policy was then tuned from OpenChat 3.5 against that reward model with advantage-induced policy alignment (APA).partial · sourcenote: Card: 'trained from Openchat 3.5 with reward model berkeley-nest/Starling-RM-7B-alpha and policy optimization method APA'. Blog (starling.cs.berkeley.edu): 'training a reward model and conducting online RL based on the existing Nectar Dataset'; 'Our reward model is fine-tuned from Llama2-7B-Chat'. Whether the policy's own prompts came from Nectar is not stated in the sources read.
- techniques
- primary_sources
- record_history
- date:2026-09-24 · change:ingested as candidate from HF (berkeley-nest/Starling-LM-7B-alpha@1dddf3b95bc1391f6307299eb1c162c194bde9bd) · by:ingest_hf.py ·date:2026-09-24 · change:preparer: developer, license (partial), training data (partial), positional encoding (propagated), 2 edges (fine_tuned_from openchat-3-5; trained_on nectar left PANEL-NEEDED); sources: card, Starling blog · by:claude (preparer, Sonnet 5) ·date:2026-09-24 · change:ruling: 3-tier panel (3/3, feedback_from via reward-model record), see session log: trained_on nectar rejected, feedback_from starling-rm-7b-alpha added; evidence tags set to declared (uploader is the developer) · by:wilson-pruitt + claude (Sonnet 5) ·date:2026-09-24 · change:reviewed and promoted from staging (2 edge(s) accepted) · by:Wilson Pruitt ·
Parents
Weights descend
- fine_tuned_from → OpenChat 3.5 declared source Card: 'Finetuned from model: Openchat 3.5 (based on Mistral-7B-v0.1)'; 'a language model trained from Openchat 3.5 with reward model ... and policy optimization method APA'. Developers' own card. PARENT openchat-3-5 IS STAGED (promote it first).
Influence without weights
- feedback_from → Starling-RM-7B-alpha declared source Card: 'a language model trained from Openchat 3.5 with reward model berkeley-nest/Starling-RM-7B-alpha and policy optimization method APA'. Blog: 'we fine-tuned the Openchat 3.5 language model using the learned reward model.' Developers' own card and blog; the uploader IS the developer. Parent is a model whose scores, not text, shaped the policy. PARENT starling-rm-7b-alpha IS STAGED (promote it first; it in turn needs nothing unpromoted except openchat-3-5 for this record).
Children
No edges recorded.
Read in
No station on the reading path has touched this record yet.