Model · mixtral-8x7b-v0-1
Mixtral-8x7B-v0.1
Developer: Mistral AI
Availability: available · checked 2026-09-24 · source
Raw record: /data/models/mixtral-8x7b-v0-1.json
Fields
- id
- mixtral-8x7b-v0-1
- identifiers
- huggingface
- mistralai/Mixtral-8x7B-v0.1
- developer
- Mistral AI
- release_date
- 2023-12-11recorded · source
- weights_status
- open
- availability
- available · checked 2026-09-24 · source
- license
- Apache 2.0recorded · source
- architecture
- family
- decoder_only_moe
- note
- Sparse mixture-of-experts. Paper: 'Mixtral has the same architecture as Mistral 7B, with the difference that each layer is composed of 8 feedforward blocks (i.e. experts).' A router selects 2 of 8 experts per token per layer. 46.7B total parameters, 12.9B active per token. Pretrained from scratch, not fine-tuned from Mistral 7B weights (Mistral AI release post).
- n_experts
- 8recorded · source
- n_experts_per_tok
- 2recorded · source
- n_layers
- nullnot_recorded
- hidden_size
- nullnot_recorded
- n_heads
- nullnot_recorded
- vocab_size
- nullnot_recorded
- positional_encoding
- nullnot_recordednote: Config not fetched this session; likely rotary as in Mistral 7B, per the 'same architecture' claim, but not independently confirmed.
- training_data
- Pretrained from scratch; neither the paper nor the release post discloses the training corpus.recorded · source
- techniques
- primary_sources
- record_history
- date:2026-09-24 · change:ingested as candidate from HF (mistralai/Mixtral-8x7B-v0.1@fc7ac94680e38d7348cfa806e51218e6273104b0) · by:ingest_hf.py ·date:2026-09-24 · change:preparer: developer, architecture (MoE, from paper), license, release date, same_architecture_retrained to mistral-7b-v0-1 added · by:claude (preparer, Sonnet 5) ·date:2026-09-24 · change:preparer cleanup: removed 2 stale flag(s), filled training_data · by:claude (preparer, Sonnet 5) ·date:2026-09-24 · change:reviewed and promoted from staging (1 edge(s) accepted) · by:Wilson Pruitt ·
Parents
Design
- same_architecture_retrained → Mistral 7B (v0.1) declared source Paper: 'Mixtral has the same architecture as Mistral 7B, with the difference that each layer is composed of 8 feedforward blocks.' Pretrained from scratch -- no weights passed, so no fine_tuned_from edge.
Children
Weights descend
- ← fine_tuned_from Mixtral-8x7B-Instruct-v0.1 declared source Uploader (mistralai) is the developer of both this and the base model.
Read in
No station on the reading path has touched this record yet.