Model · solar-10-7b-v1-0
SOLAR-10.7B-v1.0
Developer: Upstage
Availability: available · checked 2026-09-24 · source
Raw record: /data/models/solar-10-7b-v1-0.json
Fields
- id
- solar-10-7b-v1-0
- identifiers
- huggingface
- upstage/SOLAR-10.7B-v1.0
- developer
- Upstage
- release_date
- 2023-12-23recorded · sourcenote: arXiv paper (2312.15166) date.
- weights_status
- open
- availability
- available · checked 2026-09-24 · source
- license
- apache-2.0partial · sourcenote: From HF card metadata (uploader-declared); confirm against the license text.
- architecture
- family
- decoder_only
- note
- family read from config 'architectures': ['LlamaForCausalLM'].
- n_layers
- 48recorded · source
- hidden_size
- 4096recorded · source
- n_heads
- 32recorded · source
- vocab_size
- 32000recorded · source
- context_length
- 4096recorded · source
- n_kv_heads
- 8recorded · source
- positional_encoding
- rotary (RoPE)recorded · sourcenote: Inherited from Mistral 7B's layers (see fine_tuned_from note on the depth-up-scaling process); not independently re-verified against SOLAR's own config this session.
- training_data
- nullnot_recordednote: The continued-pretraining corpus for the depth-up-scaled model (distinct from the DUS parent, already captured in the fine_tuned_from edge note) is not disclosed in the sources read this session.
- techniques
- primary_sources
- record_history
- date:2026-09-24 · change:ingested as candidate from HF (upstage/SOLAR-10.7B-v1.0@5a06298f374dae1fb918d73fba72c7722ce71d36) · by:ingest_hf.py ·date:2026-09-24 · change:preparer: developer, release date set; DUS relationship to mistral-7b-v0-1 flagged, not resolved · by:claude (preparer, Sonnet 5) ·date:2026-09-24 · change:preparer: depth_upscaled_from edge to mistral-7b-v0-1 resolved per ruling · by:claude (preparer, Sonnet 5) ·date:2026-09-24 · change:ruling reversed: blind 3-tier panel (3/3 fine_tuned_from) overturned the earlier solo depth_upscaled_from call; relation type removed from schema · by:panel: haiku+sonnet+opus, applied by claude (preparer, Sonnet 5) ·date:2026-09-24 · change:preparer cleanup: removed 2 stale flag(s) · by:claude (preparer, Sonnet 5) ·date:2026-09-24 · change:preparer cleanup: training_data flag reworded to reflect a real, checked gap rather than an unfilled TODO · by:claude (preparer, Sonnet 5) ·date:2026-09-24 · change:preparer cleanup: positional_encoding propagated from an unchanged parent architecture · by:claude (preparer, Sonnet 5) ·date:2026-09-24 · change:reviewed and promoted from staging (1 edge(s) accepted) · by:Wilson Pruitt ·
Parents
Weights descend
- fine_tuned_from → Mistral 7B (v0.1) declared source Card: built by 'depth up-scaling' (DUS): duplicating and interleaving Mistral 7B's layers into a taller 10.7B-parameter model (32->48 layers), then continuing pretraining on the WHOLE enlarged model (matches the doc's fine_tuned_from test: continued pretraining of the whole model, as with Code Llama from Llama 2 -- not a merged_from case, which requires no whole-model retrain). Panel-reviewed 2026-09-24, blind three-tier panel (Haiku/Sonnet/Opus, unanimous 3/3 for fine_tuned_from over an earlier solo depth_upscaled_from ruling that was reverted). All three panelists flagged the same open question: this loses the fact that the architecture was reshaped before training resumed, which this note is here to preserve.
Children
No edges recorded.
Read in
No station on the reading path has touched this record yet.