Stemma Machinarum
A genealogy of machine-learning models, starting with open-weight language models. It records which model descends from which, what changed at each step, and the evidence for every claim.
The name comes from textual criticism: a stemma is the family tree of manuscript copies, and stemmatics already has a term, contamination, for a copy that draws on two lines at once. Distillation and model merging are contamination in exactly that sense, so the structure here is a network, not a strict tree.
It is built for people and for LLM agents alike. More about the project.
See the whole family →The LLaMA family, drawn as a manuscript stemma. Dashed lines are contamination: text written by closed models (hollow) that became training data for open ones.
Last data change: 2026-09-24. Counts are computed from the data at build time.
The Graph
Every model and dataset, every lineage edge, each with an evidence tag and a source. Raw JSON is served as-is.
The Notebook
The reading path through the primary sources, with dated notes and hands-on exercises. Each station links to the records it produced.
Follow the path · 0 of 6 stations read · 0 notes
Evidence tags
Every edge carries one. Definitions are in the method.