The stemma
The stemma
Two families so far, drawn by hand. Each is its own drawing: a model appears in the family its weights descend from, so a Llama-2-derived reward model that feeds a Mistral-family policy (Starling) is shown as a dashed link back to the LLaMA drawing, not repeated.
The LLaMA family
Textual critics draw the family tree of a text’s manuscripts as a stemma: the oldest copies at the top, each copy below the one it was made from. This is the same drawing for one family of language models, from LLaMA (February 2023) through Llama 2, Code Llama, TinyLlama, and the fine-tunes, adapters and merges built on them.
Time runs down the page. The hollow circles at the top left are closed models: their weights were never published, so, like a lost manuscript, they are known only through what descends from them. The dashed lines are what stemmatics calls contamination: text written or graded by those closed models became the training data for Alpaca, Vicuna, Baize, Orca, Tulu and others, so those open models draw on two lines at once.
On a phone the families are drawn as one indented column: each model sits under the model it came from, siblings in date order. The closed models and datasets sit in the left margin, and their lines run into each model that trained on them.
Show the full drawing (scrolls sideways)
Every line is a recorded edge and links to its source; every node links to its record.
The Mistral family
Mistral 7B (September 2023) and its fine-tunes, quantization, and Mixtral. This family shows a case the LLaMA drawing does not: feedback from a reward model (Starling-RM, trained on GPT-4’s rankings) rather than from a dataset directly, and a cross-family parent — Starling-RM was itself fine-tuned from Llama 2-Chat 7B, drawn back in the LLaMA family above.
Show the full drawing (scrolls sideways)
Every line is a recorded edge and links to its source; every node links to its record.
Key
- Weights descend: the child was fine-tuned from the parent’s parameters.
- Trained on: a model learned from a dataset.
- Contamination: a dataset holds text another model wrote (or graded).
- Design only: a successor or declared imitation, with no weights passed.
- Blue: declared by the uploader, not the original developer. Plain ink means declared by the developer.
- Hollow: a closed model, known only by its outputs, like a lost exemplar.
- Dashed stub with a hollow ?: this model’s parents are not recorded in Stemma.
- Square: a dataset.
- Dashed ring, coloured like a link: this record belongs to a different family, drawn there in full.
How to read it
- A node sits at its release date. A child is never drawn level with its parent, so a few nodes sit slightly below their true month; the date under each name is the recorded one.
- ShareGPT and SlimPajama have no recorded release date, so each is placed just above the first model trained on it.
- A dashed stub ending in a hollow ? (MythoLogic and MythoMax) marks a model whose parents are not recorded in Stemma. Both are merges of other models, but those parents are not records here, so no edge is drawn. The stub marks a connection the graph does not yet hold; it does not say the model has no parents.
- The dashed ring in the Mistral drawing (Llama 2-Chat 7B) is a different kind of stub: that record exists in Stemma, but it is drawn in full only in the LLaMA family, to avoid repeating its own lineage here. Follow it to see where it comes from.
- Line form tells you the relation; colour tells you the evidence. The full definitions are in the method.
What this drawing is not
It is one family, laid out by hand, as a test of the drawing conventions. The whole network (every record, laid out from the data) is the next step. The data behind it is already public: edges.jsonl and the records.