Stemma Machinarum

Model · starling-lm-7b-alpha

Starling-LM-7B-alpha

Developer: Banghua Zhu, Evan Frick, Tianhao Wu, Hanlin Zhu and Jiantao Jiao (Berkeley NEST; card 'Developed by')
Availability: available · checked 2026-09-24 · source

Raw record: /data/models/starling-lm-7b-alpha.json

Fields

id
starling-lm-7b-alpha
identifiers
huggingface
berkeley-nest/Starling-LM-7B-alpha
developer
Banghua Zhu, Evan Frick, Tianhao Wu, Hanlin Zhu and Jiantao Jiao (Berkeley NEST; card 'Developed by')
release_date
2023-11-25partial · source
note:  HF repo creation date; may predate or postdate public release.
weights_status
open
availability
available · checked 2026-09-24 · source
license
Apache-2.0 per card metadata and model summary, 'under the condition that the model is not used to compete with OpenAI'; the card's License section also says research preview, non-commercial, subject to LLaMA/OpenAI/ShareGPT termspartial · source
note:  The card contradicts itself: metadata `license: apache-2.0` and 'License: Apache-2.0 license under the condition that the model is not used to compete with OpenAI', but the later License section: 'research preview intended for non-commercial use only, subject to the data distillation License of LLaMA, Terms of Use of the data generated by OpenAI, and Privacy Practices of ShareGPT'. Not reconciled; reviewer check.
architecture
family
decoder_only
note
family read from config 'architectures': ['MistralForCausalLM'].
n_layers
32recorded · source
hidden_size
4096recorded · source
n_heads
32recorded · source
vocab_size
32002recorded · source
context_length
8192recorded · source
n_kv_heads
8recorded · source
positional_encoding
rotary (RoPE)recorded · source
note:  Architecture unchanged from mistral-7b-v0-1 (fine-tune, not a structural change).
training_data
Nectar (GPT-4-ranked 7-wise comparisons) used to train the reward model Starling-RM-7B-alpha (fine-tuned from Llama2-7B-Chat); the policy was then tuned from OpenChat 3.5 against that reward model with advantage-induced policy alignment (APA).partial · source
note:  Card: 'trained from Openchat 3.5 with reward model berkeley-nest/Starling-RM-7B-alpha and policy optimization method APA'. Blog (starling.cs.berkeley.edu): 'training a reward model and conducting online RL based on the existing Nectar Dataset'; 'Our reward model is fine-tuned from Llama2-7B-Chat'. Whether the policy's own prompts came from Nectar is not stated in the sources read.
techniques
primary_sources
https://huggingface.co/berkeley-nest/Starling-LM-7B-alpha
https://huggingface.co/berkeley-nest/Starling-LM-7B-alpha/raw/1dddf3b95bc1391f6307299eb1c162c194bde9bd/config.json
record_history
date:2026-09-24 · change:ingested as candidate from HF (berkeley-nest/Starling-LM-7B-alpha@1dddf3b95bc1391f6307299eb1c162c194bde9bd) · by:ingest_hf.py ·
date:2026-09-24 · change:preparer: developer, license (partial), training data (partial), positional encoding (propagated), 2 edges (fine_tuned_from openchat-3-5; trained_on nectar left PANEL-NEEDED); sources: card, Starling blog · by:claude (preparer, Sonnet 5) ·
date:2026-09-24 · change:ruling: 3-tier panel (3/3, feedback_from via reward-model record), see session log: trained_on nectar rejected, feedback_from starling-rm-7b-alpha added; evidence tags set to declared (uploader is the developer) · by:wilson-pruitt + claude (Sonnet 5) ·
date:2026-09-24 · change:reviewed and promoted from staging (2 edge(s) accepted) · by:Wilson Pruitt ·

Parents

Weights descend

Influence without weights

Children

No edges recorded.

Read in

No station on the reading path has touched this record yet.