{"child": "alpaca-7b", "parent": "llama-7b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://crfm.stanford.edu/2023/03/13/alpaca.html", "note": "Stanford CRFM: \"fine-tuned from Meta's LLaMA 7B model\""}
{"child": "llama-2-7b-chat", "parent": "llama-2-7b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://arxiv.org/abs/2307.09288", "note": "Paper sec. 1: 'Llama 2-Chat, a fine-tuned version of Llama 2 that is optimized for dialogue use cases.'"}
{"child": "llama-2-7b", "parent": "llama-7b", "relation": "successor_in_series", "evidence": "declared", "source": "https://arxiv.org/abs/2307.09288", "note": "Paper: 'Llama 2, an updated version of Llama 1'. Trained from scratch on a new data mix, NOT initialized from Llama 1 weights."}
{"child": "code-llama-7b", "parent": "llama-2-7b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://arxiv.org/abs/2308.12950", "note": "Paper sec. 2: 'We train Code Llama on 500B tokens during the initial phase, starting from the 7B, 13B, and 34B versions of Llama 2.' Continued pretraining rather than instruction tuning."}
{"child": "vicuna-7b-v1-3", "parent": "llama-7b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://github.com/lm-sys/FastChat/blob/main/docs/vicuna_weights_version.md", "note": "FastChat docs: v1.3 base model 'Llama 1'; model card: 'Finetuned from: LLaMA'. Neither names the size; 7B-to-7B pairing rests on the model name and matching config dims."}
{"child": "zephyr-7b-beta", "parent": "mistral-7b-v0-1", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://huggingface.co/HuggingFaceH4/zephyr-7b-beta", "note": "Developer's own model card: 'fine-tuned version of mistralai/Mistral-7B-v0.1'."}
{"child": "gpt-neox-20b", "parent": "gpt-j-6b", "relation": "same_architecture_retrained", "evidence": "declared", "source": "https://arxiv.org/abs/2204.06745", "note": "Paper sec. 2.1: 'Our model architecture is almost identical to that of GPT-J', though it uses GPT-3 as the stated reference 'because there is no canonical published reference on the design of GPT-J.' No weights passed."}
{"child": "alpaca-52k", "parent": "text-davinci-003", "relation": "distilled_from_outputs", "evidence": "declared", "source": "https://crfm.stanford.edu/2023/03/13/alpaca.html", "note": "Stanford CRFM generated the data itself with text-davinci-003."}
{"child": "alpaca-7b", "parent": "alpaca-52k", "relation": "trained_on", "evidence": "declared", "source": "https://crfm.stanford.edu/2023/03/13/alpaca.html", "note": "Fine-tuning data."}
{"child": "sharegpt-vicuna", "parent": "chatgpt", "relation": "distilled_from_outputs", "evidence": "declared", "source": "https://lmsys.org/blog/2023-03-30-vicuna/", "note": "ShareGPT.com is a site 'where users can share their ChatGPT conversations.' User-curated, not generated by LMSYS."}
{"child": "vicuna-7b-v1-3", "parent": "sharegpt-vicuna", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/lmsys/vicuna-7b-v1.3", "note": "Card: 'around 125K conversations collected from ShareGPT.com'."}
{"child": "ultrachat", "parent": "chatgpt", "relation": "distilled_from_outputs", "evidence": "declared", "source": "https://arxiv.org/abs/2305.14233", "note": "Builder's own paper: both sides of each dialogue generated by 'ChatGPT Turbo APIs'. Which ChatGPT snapshot is not_recorded."}
{"child": "zephyr-7b-beta", "parent": "ultrachat", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/HuggingFaceH4/zephyr-7b-beta", "note": "Distilled SFT step; card: 'filtered and preprocessed' UltraChat."}
{"child": "ultrafeedback", "parent": "gpt-4", "relation": "feedback_from", "evidence": "declared", "source": "https://arxiv.org/abs/2310.01377", "note": "Builder's own paper: GPT-4 employed 'to offer detailed feedback in both numerical and textual forms.'"}
{"child": "zephyr-7b-beta", "parent": "ultrafeedback", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/HuggingFaceH4/zephyr-7b-beta", "note": "DPO step on GPT-4's rankings of model completions."}
{"child": "gpt2-xl", "parent": "webtext", "relation": "trained_on", "evidence": "declared", "source": "https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf"}
{"child": "gpt-neo-2-7b", "parent": "the-pile", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/EleutherAI/gpt-neo-2.7B", "note": "420B tokens."}
{"child": "gpt-j-6b", "parent": "the-pile", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/EleutherAI/gpt-j-6b", "note": "402B tokens."}
{"child": "gpt-neox-20b", "parent": "the-pile", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2204.06745"}
{"child": "pythia-6-9b", "parent": "the-pile", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2304.01373", "note": "300B tokens; a twin model was trained on the deduplicated Pile."}
{"child": "opt-6-7b", "parent": "the-pile", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2205.01068", "note": "A subset of the Pile, alongside RoBERTa corpus data and PushShift.io Reddit (not recorded as datasets)."}
{"child": "bloom-7b1", "parent": "roots", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2211.05100"}
{"child": "falcon-7b", "parent": "refinedweb", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/tiiuae/falcon-7b", "note": "RefinedWeb-English is 79% of 1,500B training tokens; the rest is curated corpora (not recorded as datasets)."}
{"child": "gpt-3", "parent": "gpt2-xl", "relation": "same_architecture_retrained", "evidence": "declared", "source": "https://arxiv.org/abs/2005.14165", "note": "Paper: 'the same model and architecture as GPT-2', except alternating dense and locally banded sparse attention. gpt2-xl stands in for the GPT-2 family."}
{"child": "gpt-neo-2-7b", "parent": "gpt-3", "relation": "same_architecture_retrained", "evidence": "declared", "source": "https://huggingface.co/EleutherAI/gpt-neo-2.7B", "note": "Card: 'designed using EleutherAI's replication of the GPT-3 architecture.' Like GPT-3, alternates dense (global) and local attention layers."}
{"child": "gpt-neox-20b", "parent": "gpt-3", "relation": "design_follows", "evidence": "declared", "source": "https://arxiv.org/abs/2204.06745", "note": "Paper: architecture 'largely follows that of GPT-3'."}
{"child": "pythia-6-9b", "parent": "gpt-3", "relation": "design_follows", "evidence": "declared", "source": "https://arxiv.org/abs/2304.01373", "note": "Paper: 'Our model architecture and hyperparameters largely follow Brown et al. (2020)'."}
{"child": "opt-6-7b", "parent": "gpt-3", "relation": "design_follows", "evidence": "declared", "source": "https://arxiv.org/abs/2205.01068", "note": "Paper: trained 'to roughly match the performance and sizes of the GPT-3 class of models'; hyperparameters 'largely follow Brown et al. (2020)'."}
{"child": "falcon-7b", "parent": "gpt-3", "relation": "design_follows", "evidence": "declared", "source": "https://huggingface.co/tiiuae/falcon-7b", "note": "Card: 'broadly adapted from the GPT-3 paper'."}
{"child": "bloom", "parent": "roots", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2211.05100", "note": "Paper describes one training recipe (ROOTS corpus) for the whole BLOOM suite, incl. the 176B model. Same as bloom-7b1."}
{"child": "bloom-560m", "parent": "roots", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2211.05100", "note": "Paper describes one training recipe (ROOTS corpus) for the whole BLOOM suite, incl. the 560M model. Same as bloom-7b1."}
{"child": "bloomz-7b1", "parent": "bloom-7b1", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://huggingface.co/bigscience/bloomz-7b1", "note": "Card: 'Architecture: Same as bloom-7b1'; fine-tuned 1000 steps / 4.19B tokens on xP3 (crosslingual task mixture, 46 languages)."}
{"child": "codellama-7b-instruct", "parent": "code-llama-7b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://arxiv.org/abs/2308.12950", "note": "Paper describes Code Llama - Instruct as a further fine-tuned stage of the base Code Llama models for instruction following."}
{"child": "codellama-7b-python", "parent": "code-llama-7b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://arxiv.org/abs/2308.12950", "note": "Paper: Code Llama - Python models are 'further specialized on 100B tokens using a Python-heavy dataset', starting from the base Code Llama checkpoints."}
{"child": "falcon-40b", "parent": "refinedweb", "relation": "trained_on", "evidence": "declared_by_uploader", "source": "https://huggingface.co/tiiuae/falcon-40b/blob/05ab2ee8d6b593bdbab17d728de5c028a7a94d83/README.md", "note": "HF card metadata datasets: tiiuae/falcon-refinedweb. Often incomplete; confirm which training stage used it."}
{"child": "falcon-40b", "parent": "gpt-3", "relation": "design_follows", "evidence": "declared_by_uploader", "source": "https://huggingface.co/tiiuae/falcon-40b", "note": "Card verified this session: 'The architecture is broadly adapted from the GPT-3 paper (Brown et al., 2020), with the following differences...' -- same wording as falcon-7b."}
{"child": "falcon-7b-instruct", "parent": "refinedweb", "relation": "trained_on", "evidence": "declared_by_uploader", "source": "https://huggingface.co/tiiuae/falcon-7b-instruct/blob/8782b5c5d8c9290412416618f36a133653e85285/README.md", "note": "Card: RefinedWeb-English is 5% (13M of 250M) of the fine-tuning mix."}
{"child": "falcon-7b-instruct", "parent": "falcon-7b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://huggingface.co/tiiuae/falcon-7b-instruct", "note": "Card: 'built by TII based on Falcon-7B and finetuned on a mixture of chat/instruct datasets.'"}
{"child": "gpt-neo-1-3b", "parent": "the-pile", "relation": "trained_on", "evidence": "declared_by_uploader", "source": "https://huggingface.co/EleutherAI/gpt-neo-1.3B/blob/dbe59a7f4a88d01d1ba9798d78dbe3fe038792c8/README.md", "note": "Card metadata datasets: EleutherAI/pile. Consistent with the gpt-neo-2-7b record, same family/release."}
{"child": "gpt-neo-125m", "parent": "the-pile", "relation": "trained_on", "evidence": "declared_by_uploader", "source": "https://huggingface.co/EleutherAI/gpt-neo-125m/blob/21def0189f5705e2521767faed922f1f15e7d7db/README.md", "note": "Card metadata datasets: EleutherAI/pile. Consistent with the gpt-neo-2-7b record, same family/release."}
{"child": "gpt2", "parent": "webtext", "relation": "trained_on", "evidence": "declared", "source": "https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf"}
{"child": "gpt2-large", "parent": "webtext", "relation": "trained_on", "evidence": "declared", "source": "https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf"}
{"child": "gpt2-medium", "parent": "webtext", "relation": "trained_on", "evidence": "declared", "source": "https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf"}
{"child": "llama-2-13b", "parent": "llama-7b", "relation": "successor_in_series", "evidence": "declared", "source": "https://arxiv.org/abs/2307.09288", "note": "Paper: 'Llama 2, an updated version of Llama 1... trained on a new mix of publicly available data', not initialized from Llama 1 weights. Same edge as llama-2-7b."}
{"child": "llama-2-70b", "parent": "llama-7b", "relation": "successor_in_series", "evidence": "declared", "source": "https://arxiv.org/abs/2307.09288", "note": "Paper: 'Llama 2, an updated version of Llama 1... trained on a new mix of publicly available data', not initialized from Llama 1 weights. Same edge as llama-2-7b."}
{"child": "mistral-7b-instruct-v0-1", "parent": "mistral-7b-v0-1", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.1/blob/ec5deb64f2c6e6fa90c1abf74a91d5c93a9669ca/README.md", "note": "Uploader (mistralai) is the developer of both this and the base model, so the base_model card claim counts as declared, not just declared_by_uploader."}
{"child": "mistral-7b-openorca", "parent": "mistral-7b-v0-1", "relation": "fine_tuned_from", "evidence": "declared_by_uploader", "source": "https://huggingface.co/Open-Orca/Mistral-7B-OpenOrca", "note": "Card: 'fine-tuned on top of Mistral 7B'."}
{"child": "mixtral-8x7b-v0-1", "parent": "mistral-7b-v0-1", "relation": "same_architecture_retrained", "evidence": "declared", "source": "https://arxiv.org/abs/2310.06825", "note": "Paper: 'Mixtral has the same architecture as Mistral 7B, with the difference that each layer is composed of 8 feedforward blocks.' Pretrained from scratch -- no weights passed, so no fine_tuned_from edge."}
{"child": "mpt-7b-instruct", "parent": "mpt-7b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://www.databricks.com/blog/mpt-7b", "note": "Post: 'Built by finetuning MPT-7B on a dataset we also release, derived from Databricks Dolly-15k and Anthropic's Helpful and Harmless datasets.'"}
{"child": "nous-hermes-llama-2-7b", "parent": "llama-2-7b", "relation": "fine_tuned_from", "evidence": "declared_by_uploader", "source": "https://huggingface.co/NousResearch/Nous-Hermes-llama-2-7b", "note": "Card: fine-tuned version of Llama-2-7b."}
{"child": "open-llama-7b", "parent": "llama-7b", "relation": "design_follows", "evidence": "declared", "source": "https://huggingface.co/openlm-research/open_llama_7b", "note": "Card: 'a permissively licensed open source reproduction of Meta AI's LLaMA', trained from scratch on different data (RedPajama, not LLaMA's original mix). No weights passed."}
{"child": "openhermes-2-5-mistral-7b", "parent": "mistral-7b-v0-1", "relation": "fine_tuned_from", "evidence": "declared_by_uploader", "source": "https://huggingface.co/teknium/OpenHermes-2.5-Mistral-7B/blob/24c0bea14d53e6f67f1fbe2eca5bfe7cae389b33/README.md", "note": "HF card metadata base_model: mistralai/Mistral-7B-v0.1. Relation 'finetune' is the Hub's inference (tag), not stated in the card."}
{"child": "opt-1-3b", "parent": "the-pile", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2205.01068", "note": "Paper: trained on ~180B tokens incl. 'a subset of the Pile'; also RoBERTa corpus subsets and PushShift.io Reddit (not recorded as datasets). Same mix as opt-6-7b -- not stated per-size in the card, but the paper describes one training recipe for the whole OPT suite."}
{"child": "opt-125m", "parent": "the-pile", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2205.01068", "note": "Paper: trained on ~180B tokens incl. 'a subset of the Pile'; also RoBERTa corpus subsets and PushShift.io Reddit (not recorded as datasets). Same mix as opt-6-7b -- not stated per-size in the card, but the paper describes one training recipe for the whole OPT suite."}
{"child": "opt-13b", "parent": "the-pile", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2205.01068", "note": "Paper: trained on ~180B tokens incl. 'a subset of the Pile'; also RoBERTa corpus subsets and PushShift.io Reddit (not recorded as datasets). Same mix as opt-6-7b -- not stated per-size in the card, but the paper describes one training recipe for the whole OPT suite."}
{"child": "opt-30b", "parent": "the-pile", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2205.01068", "note": "Paper: trained on ~180B tokens incl. 'a subset of the Pile'; also RoBERTa corpus subsets and PushShift.io Reddit (not recorded as datasets). Same mix as opt-6-7b -- not stated per-size in the card, but the paper describes one training recipe for the whole OPT suite."}
{"child": "opt-66b", "parent": "the-pile", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2205.01068", "note": "Paper: trained on ~180B tokens incl. 'a subset of the Pile'; also RoBERTa corpus subsets and PushShift.io Reddit (not recorded as datasets). Same mix as opt-6-7b -- not stated per-size in the card, but the paper describes one training recipe for the whole OPT suite."}
{"child": "phi-2", "parent": "refinedweb", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/microsoft/phi-2", "note": "Card: training mix includes 'filtered web data from Falcon RefinedWeb and SlimPajama, which was assessed by AOAI GPT-4'."}
{"child": "pythia-1-4b", "parent": "the-pile", "relation": "trained_on", "evidence": "declared_by_uploader", "source": "https://huggingface.co/EleutherAI/pythia-1.4b/blob/fedc38a16eea3bd36a96b906d78d11d2ce18ed79/README.md", "note": "HF card metadata datasets: EleutherAI/the_pile. Often incomplete; confirm which training stage used it."}
{"child": "pythia-12b", "parent": "the-pile", "relation": "trained_on", "evidence": "declared_by_uploader", "source": "https://huggingface.co/EleutherAI/pythia-12b/blob/bb1e3e710cdf6b524461d543cfb5ba773f0a81b6/README.md", "note": "HF card metadata datasets: EleutherAI/pile. Often incomplete; confirm which training stage used it."}
{"child": "pythia-160m", "parent": "the-pile", "relation": "trained_on", "evidence": "declared_by_uploader", "source": "https://huggingface.co/EleutherAI/pythia-160m/blob/50f5173d932e8e61f858120bcb800b97af589f46/README.md", "note": "HF card metadata datasets: EleutherAI/pile. Often incomplete; confirm which training stage used it."}
{"child": "pythia-1b", "parent": "the-pile", "relation": "trained_on", "evidence": "declared_by_uploader", "source": "https://huggingface.co/EleutherAI/pythia-1b/blob/f73d7dcc545c8bd326d8559c8ef84ffe92fea6b2/README.md", "note": "HF card metadata datasets: the_pile. Often incomplete; confirm which training stage used it."}
{"child": "pythia-2-8b", "parent": "the-pile", "relation": "trained_on", "evidence": "declared_by_uploader", "source": "https://huggingface.co/EleutherAI/pythia-2.8b/blob/2a259cdd96a4beb1cdf467512e3904197345f6a9/README.md", "note": "HF card metadata datasets: EleutherAI/pile. Often incomplete; confirm which training stage used it."}
{"child": "pythia-410m", "parent": "the-pile", "relation": "trained_on", "evidence": "declared_by_uploader", "source": "https://huggingface.co/EleutherAI/pythia-410m/blob/9879c9b5f8bea9051dcb0e68dff21493d67e9d4f/README.md", "note": "HF card metadata datasets: EleutherAI/pile. Often incomplete; confirm which training stage used it."}
{"child": "pythia-70m", "parent": "the-pile", "relation": "trained_on", "evidence": "declared_by_uploader", "source": "https://huggingface.co/EleutherAI/pythia-70m/blob/a39f36b100fe8a5377810d56c3f4789b9c53ac42/README.md", "note": "HF card metadata datasets: EleutherAI/pile. Often incomplete; confirm which training stage used it."}
{"child": "solar-10-7b-v1-0", "parent": "mistral-7b-v0-1", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://huggingface.co/upstage/SOLAR-10.7B-v1.0", "note": "Card: built by 'depth up-scaling' (DUS): duplicating and interleaving Mistral 7B's layers into a taller 10.7B-parameter model (32->48 layers), then continuing pretraining on the WHOLE enlarged model (matches the doc's fine_tuned_from test: continued pretraining of the whole model, as with Code Llama from Llama 2 -- not a merged_from case, which requires no whole-model retrain). Panel-reviewed 2026-09-24, blind three-tier panel (Haiku/Sonnet/Opus, unanimous 3/3 for fine_tuned_from over an earlier solo depth_upscaled_from ruling that was reverted). All three panelists flagged the same open question: this loses the fact that the architecture was reshaped before training resumed, which this note is here to preserve."}
{"child": "vicuna-7b-v1-5", "parent": "llama-2-7b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://github.com/lm-sys/FastChat/blob/main/docs/vicuna_weights_version.md", "note": "Version doc: v1.5 base model is Llama 2. Card doesn't carry base_model metadata; sourced to FastChat's own release notes, same as v1.3."}
{"child": "vicuna-7b-v1-5", "parent": "sharegpt-vicuna", "relation": "trained_on", "evidence": "declared", "source": "https://github.com/lm-sys/FastChat/blob/main/docs/vicuna_weights_version.md", "note": "Same ShareGPT collection as v1.3; version doc gives 370M training tokens, same figure as v1.1/v1.3."}
{"child": "zephyr-7b-alpha", "parent": "mistral-7b-v0-1", "relation": "fine_tuned_from", "evidence": "declared_by_uploader", "source": "https://huggingface.co/HuggingFaceH4/zephyr-7b-alpha/blob/014792bbb59d04ced3b5a9b8b4dfc926655d958f/README.md", "note": "HF card metadata base_model: mistralai/Mistral-7B-v0.1. Relation 'finetune' is the Hub's inference (tag), not stated in the card."}
{"child": "zephyr-7b-alpha", "parent": "ultrachat", "relation": "trained_on", "evidence": "declared_by_uploader", "source": "https://huggingface.co/HuggingFaceH4/zephyr-7b-alpha/blob/014792bbb59d04ced3b5a9b8b4dfc926655d958f/README.md", "note": "HF card metadata datasets: stingning/ultrachat. Often incomplete; confirm which training stage used it."}
{"child": "zephyr-7b-alpha", "parent": "ultrafeedback", "relation": "trained_on", "evidence": "declared_by_uploader", "source": "https://huggingface.co/HuggingFaceH4/zephyr-7b-alpha/blob/014792bbb59d04ced3b5a9b8b4dfc926655d958f/README.md", "note": "HF card metadata datasets: openbmb/UltraFeedback. Often incomplete; confirm which training stage used it."}
{"child": "dolly-v2-12b", "parent": "pythia-12b", "relation": "fine_tuned_from", "evidence": "declared", "source": "http://web.archive.org/web/20230602091912/https://www.databricks.com/blog/2023/04/12/dolly-first-open-commercially-viable-instruction-tuned-llm", "note": "Blog: 'a 12B parameter language model based on the EleutherAI pythia model family', 'based on EleutherAI's pythia-12b'."}
{"child": "falcon-40b-instruct", "parent": "refinedweb", "relation": "trained_on", "evidence": "declared_by_uploader", "source": "https://huggingface.co/tiiuae/falcon-40b-instruct/blob/ecb78d97ac356d098e79f0db222c9ce7c5d9ee5f/README.md", "note": "Card: '150M tokens from Baize mixed with 5% of RefinedWeb data.'"}
{"child": "falcon-40b-instruct", "parent": "falcon-40b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://huggingface.co/tiiuae/falcon-40b-instruct", "note": "Card: 'Finetuned from model: Falcon-40B'."}
{"child": "mistral-7b-instruct-v0-1-gptq", "parent": "mistral-7b-instruct-v0-1", "relation": "quantized_from", "evidence": "declared_by_uploader", "source": "https://huggingface.co/TheBloke/Mistral-7B-Instruct-v0.1-GPTQ/blob/6ae1e4ae2cfbaf107c705ed722ec243b4f88014d/README.md", "note": "HF card metadata base_model: mistralai/Mistral-7B-Instruct-v0.1. Relation 'quantized' is the Hub's inference (tag), not stated in the card."}
{"child": "mixtral-8x7b-instruct-v0-1", "parent": "mixtral-8x7b-v0-1", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1/blob/eba92302a2861cdc0098cc54bc9f17cb2c47eb61/README.md", "note": "Uploader (mistralai) is the developer of both this and the base model."}
{"child": "vicuna-13b-v1-5", "parent": "llama-2-13b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://github.com/lm-sys/FastChat/blob/main/docs/vicuna_weights_version.md", "note": "Same as vicuna-7b-v1-5, 13B size."}
{"child": "vicuna-13b-v1-5", "parent": "sharegpt-vicuna", "relation": "trained_on", "evidence": "declared", "source": "https://github.com/lm-sys/FastChat/blob/main/docs/vicuna_weights_version.md"}
{"child": "wizardlm-13b-v1-2", "parent": "llama-2-13b", "relation": "fine_tuned_from", "evidence": "declared_by_uploader", "source": "https://huggingface.co/WizardLMTeam/WizardLM-13B-V1.2", "note": "Card: 'this model is trained from Llama-2 13b'."}
{"child": "openorca", "parent": "gpt-4", "relation": "distilled_from_outputs", "evidence": "declared", "source": "https://huggingface.co/datasets/Open-Orca/OpenOrca", "note": "Card: '~1M GPT-4 completions'. Builder's own card."}
{"child": "openorca", "parent": "chatgpt", "relation": "distilled_from_outputs", "evidence": "declared", "source": "https://huggingface.co/datasets/Open-Orca/OpenOrca", "note": "Card: '~3.2M GPT-3.5 completions'. Linked to the unversioned ChatGPT stub; the exact GPT-3.5 model is not_recorded on the card."}
{"child": "openhermes-2-5", "parent": "gpt-4", "relation": "distilled_from_outputs", "evidence": "declared", "source": "https://huggingface.co/datasets/teknium/OpenHermes-2.5", "note": "Builder's card tags the set 'GPT-4' and 'Distillation' and several sources are marked GPT-4-only. The card does not give a per-source split, so the share generated by GPT-4 is not_recorded; other closed-model outputs may also be present."}
{"child": "baize", "parent": "chatgpt", "relation": "distilled_from_outputs", "evidence": "declared", "source": "https://arxiv.org/abs/2304.01196", "note": "Paper: 'leveraging ChatGPT (gpt-3.5-turbo)' for self-chat; snapshot not_recorded."}
{"child": "gpt4all", "parent": "chatgpt", "relation": "distilled_from_outputs", "evidence": "declared", "source": "https://s3.amazonaws.com/static.nomic.ai/gpt4all/2023_GPT4All_Technical_Report.pdf", "note": "Technical report: 'We collected roughly one million prompt-response pairs using the GPT-3.5-Turbo OpenAI API'; snapshot not_recorded."}
{"child": "nectar", "parent": "gpt-4", "relation": "feedback_from", "evidence": "declared", "source": "https://huggingface.co/datasets/berkeley-nest/Nectar", "note": "Card: 'generated through GPT-4-based ranking'. The candidate answers were written by several models; only GPT-4's judgments are asserted here."}
{"child": "alpaca-cleaned", "parent": "text-davinci-003", "relation": "distilled_from_outputs", "evidence": "declared", "source": "https://crfm.stanford.edu/2023/03/13/alpaca.html", "note": "Inherited through Alpaca: Stanford CRFM generated the original 52K with text-davinci-003; the cleaner's card describes this set as a cleaned version of it. Cleaning changed some outputs; the fraction is not_recorded."}
{"child": "tulu-v2-sft-mixture", "parent": "gpt-4", "relation": "distilled_from_outputs", "evidence": "declared", "source": "https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture", "note": "Card names GPT4-Alpaca as 'distilled GPT-4 data' and 30,000 Open-Orca samples 'generated by GPT-4'. Only these subsets are asserted; ShareGPT and WizardLM subsets also hold other models' outputs, not linked here."}
{"child": "evol-instruct-70k", "parent": "chatgpt", "relation": "distilled_from_outputs", "evidence": "declared", "source": "https://arxiv.org/abs/2304.12244", "note": "Builders' own paper: evolution 'using Azure OpenAI ChatGPT API. Then, we leverage ChatGPT to generate responses.' Which ChatGPT snapshot is not_recorded (paper says gpt-3.5-turbo from the Azure portal)."}
{"child": "code-evol-instruct", "parent": "chatgpt", "relation": "distilled_from_outputs", "evidence": "declared", "source": "https://arxiv.org/abs/2306.08568", "note": "Builders' own paper: 'OpenAI's gpt3.5turbo is used to evolve the dataset and generate responses.' Snapshot not_recorded."}
{"child": "gpt4all-j-prompt-generations", "parent": "chatgpt", "relation": "distilled_from_outputs", "evidence": "declared", "source": "https://static.nomic.ai/gpt4all/2023_GPT4All-J_Technical_Report_2.pdf", "note": "Builders' report: 'The assistant data was gathered from OpenAI's GPT3.5-Turbo'; the dataset 'is a superset of the original 400k points GPT4All dataset'. v1.3 also adds ShareGPT and Dolly rows, which this edge does not cover. Snapshot not_recorded."}
{"child": "redpajama-incite-7b-base", "parent": "redpajama-data-1t", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/togethercomputer/RedPajama-INCITE-7B-Base/blob/78f7e482443971f4873ba3239f0ac810a367833b/README.md", "note": "Developer's card (Together is the uploader and developer): 'Training Data: Please refer to togethercomputer/RedPajama-Data-1T'; also in card metadata `datasets`."}
{"child": "orca-1-data", "parent": "chatgpt", "relation": "distilled_from_outputs", "evidence": "declared", "source": "https://arxiv.org/abs/2306.02707", "note": "Builder's own paper: 'We use Azure OpenAI API to collect ChatGPT (GPT-3.5-turbo) responses to FLAN-5M'. Snapshot not_recorded."}
{"child": "orca-1-data", "parent": "gpt-4", "relation": "distilled_from_outputs", "evidence": "declared", "source": "https://arxiv.org/abs/2306.02707", "note": "Builder's own paper: '... and GPT-4 responses to FLAN-1M' (1M queries sampled from the 5M)."}
{"child": "orca-2-data", "parent": "gpt-4", "relation": "distilled_from_outputs", "evidence": "declared", "source": "https://arxiv.org/abs/2311.11045", "note": "Orca 2 paper section 4.2: 'training on ~1.8 million GPT-4 data' = Orca 1's 1M GPT-4 data plus Orca 2's 817K. The paper names GPT-4 for the doctor-patient conversations explicitly; for the other three sources it gives no separate teacher line, so this edge rests on the section 4.2 labelling. Tag `declared` is the paper's aggregate label ('~1.8 million GPT-4 data'), not a per-source statement: only the doctor-patient portion is attributed to GPT-4 by name, so for the rest GPT-4 authorship is the paper's own summary, not separately sourced (Wilson's ruling, 2026-09-24: keep declared, note the limit)."}
{"child": "baize-sdf", "parent": "chatgpt", "relation": "feedback_from", "evidence": "declared", "source": "https://arxiv.org/abs/2304.01196", "note": "Paper: 'we use the resulted Baize v1.5 models to generate four responses for each instruction from the Quora dataset mentioned in Table 2. We then engage ChatGPT using the prompt provided in Appendix C to rank generate responses for self-distillation. Finally, we select the best response ranked by ChatGPT to finetune the model.' ChatGPT ranked candidates; the selected text is Baize v1.5's own output, so this is feedback (judgments), not distilled_from_outputs. ChatGPT snapshot not_recorded."}
{"child": "alpaca-lora-7b", "parent": "llama-7b", "relation": "adapter_on", "evidence": "declared", "source": "https://huggingface.co/tloen/alpaca-lora-7b/blob/12103d6baae1b320aa60631b38acb6ea094a0539/README.md", "note": "Card: 'This repo contains a low-rank adapter for LLaMA-7b fit on the [Stanford Alpaca] dataset.' adapter_config.json: peft_type LORA, r 16, targets q_proj,k_proj,v_proj,o_proj; base_model_name_or_path decapoda-research/llama-7b-hf (a third-party re-upload of LLaMA-7B; the card names LLaMA-7b). The record holds adapter weights only; a runnable model needs LLaMA-7B."}
{"child": "alpaca-lora-7b", "parent": "alpaca-cleaned", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/tloen/alpaca-lora-7b/blob/12103d6baae1b320aa60631b38acb6ea094a0539/README.md", "note": "Card metadata: datasets: [yahma/alpaca-cleaned]; README training command --data_path 'yahma/alpaca-cleaned'. Uploader tloen is the developer. Card prose instead says 'Stanford Alpaca dataset'; see training_data note. Reviewer: if you prefer the prose reading, swap to alpaca-52k."}
{"child": "baize-v2-7b", "parent": "llama-7b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://huggingface.co/project-baize/baize-v2-7b/blob/e4731c2c2671e2d0b47b5eba08c753ca21671fab/README.md", "note": "Card: 'Baize is an open-source chat model fine-tuned with LoRA. This model is a 7B Baize-v2 ... This checkpoint has been merged with LLaMA so it's ready for use.' The LoRA is merged into LLaMA, so the weights descend from LLaMA. Size (7B, so llama-7b) is from the model's name; the README's v1 merge example uses base huggyllama/llama-7b. Intermediate step not recorded: the Baize paper (arXiv 2304.01196, Table 3) has v2 built on Baize v1.5, itself LLaMA-7B fine-tuned, and v1.5 is not recorded in Stemma. This edge is true but skips that step (Wilson's ruling, 2026-09-24: keep, note the gap)."}
{"child": "baize-v2-7b", "parent": "baize", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/project-baize/baize-v2-7b/blob/e4731c2c2671e2d0b47b5eba08c753ca21671fab/README.md", "note": "Card: trained on self-chat data (SFT + SDF); paper: 'leveraging ChatGPT to engage in a conversation with itself'. The `baize` dataset record describes the v1 corpus; the README ships separate v1 and v2 collection scripts (collect.py, collect_v2.py), so v2's exact dialogues are not pinned down. Reviewer: confirm the family-level edge is acceptable."}
{"child": "baize-v2-7b", "parent": "baize-sdf", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2304.01196", "note": "Card: v2 trained with 'self-distillation with feedback (SDF)'. Paper Table 3 lists Baize-v2-7B as SDF applied to Baize-v1.5-7B: 'we apply new LoRA modules to all linear layers in Baize v1.5'. SDF data is Baize v1.5's own generations for Quora instructions, ranked by ChatGPT."}
{"child": "cerebras-gpt-13b", "parent": "the-pile", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2304.03208", "note": "Paper abstract and sec. 2.2: trained on the Pile with the provided train/test/validation splits; no deduplication."}
{"child": "cerebras-gpt-13b", "parent": "gpt-3", "relation": "design_follows", "evidence": "declared", "source": "https://arxiv.org/abs/2304.03208", "note": "Paper: 'Cerebras-GPT models have a GPT-3-like architecture, an autoregressive transformer decoder model. The main difference is that unlike GPT-3, which uses alternating dense and sparse-banded attention, we use dense attention in all decoder blocks.' 'GPT-3-like' with a stated exception -> design_follows (not same_architecture_retrained: the developer names a difference and does not say 'same')."}
{"child": "dolly-v1-6b", "parent": "gpt-j-6b", "relation": "fine_tuned_from", "evidence": "declared", "source": "http://web.archive.org/web/20230402044135/https://huggingface.co/databricks/dolly-v1-6b", "note": "Card: 'derived from EleutherAI's GPT-J (released June 2021) and fine-tuned on a ~52K record instruction corpus'; blog: '6 billion parameter model from EleutherAI'."}
{"child": "dolly-v1-6b", "parent": "alpaca-52k", "relation": "trained_on", "evidence": "declared", "source": "http://web.archive.org/web/20230402044135/https://huggingface.co/databricks/dolly-v1-6b", "note": "Card: 'fine-tuned on a ~52K record instruction corpus (Stanford Alpaca)'; blog: 'data from Alpaca'. The Alpaca set is text-davinci-003 output (see the `alpaca-52k` dataset edge), so the closed-model path runs through it."}
{"child": "gpt-jt-6b-v1", "parent": "gpt-j-6b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://www.together.ai/blog/releasing-v1-of-gpt-jt-powered-by-open-source-ai", "note": "Blog: 'A fork of GPT-J-6B, fine-tuned on 3.53 billion tokens'; card: 'a fork of EleutherAI's GPT-J (6B)'."}
{"child": "gpt-jt-6b-v1", "parent": "the-pile", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/togethercomputer/GPT-JT-6B-v1/blob/f34aa35f906895602c1f86f5685e598afdea8051/README.md", "note": "Card: 2.62B tokens with UL2 loss on the Pile, and 55% of the second-stage mix."}
{"child": "gpt4all-j", "parent": "gpt-j-6b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://static.nomic.ai/gpt4all/2023_GPT4All-J_Technical_Report_2.pdf", "note": "Report abstract: 'deriving its weights from the Apache-licensed GPT-J model rather than the GPL-licensed of LLaMA'; card: 'Finetuned From: GPT-J'."}
{"child": "gpt4all-j", "parent": "gpt4all-j-prompt-generations", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/nomic-ai/gpt4all-j/blob/5000faf803b3edcbeff9bdb6fdbbabd1a42addd0/README.md", "note": "Card metadata `datasets: nomic-ai/gpt4all-j-prompt-generations`, and the report names the GPT4All-J dataset as its training set. Nomic is both uploader and developer, so declared. Default revision v1.0."}
{"child": "gpt4all-j", "parent": "gpt4all", "relation": "trained_on", "evidence": "declared", "source": "https://static.nomic.ai/gpt4all/2023_GPT4All-J_Technical_Report_2.pdf", "note": "Report: 'we curated the GPT4All-J dataset by augmenting the original 400k GPT4All examples' and the 800k set is 'a superset of the original 400k points GPT4All dataset'. This is the distillation path: the GPT4All examples were generated with GPT-3.5-Turbo (see the `gpt4all` dataset record)."}
{"child": "guanaco-7b", "parent": "llama-7b", "relation": "adapter_on", "evidence": "declared", "source": "https://huggingface.co/timdettmers/guanaco-7b/blob/cad9de16fb306d5bb1feb901333fe0aa7bd700d8/README.md", "note": "Card: 'open-source finetuned chatbots obtained through 4-bit QLoRA tuning of LLaMA base models'; 'Lightweight checkpoints which only contain adapter weights'; usage loads huggyllama/llama-7b (a third-party re-upload of LLaMA-7B) with adapters timdettmers/guanaco-7b. adapter_config.json: LoRA r 64, targets q,k,v,o,gate,up,down proj; its base path is a local path (/gscratch/zlab/llama/7B). QLoRA trains the adapter against a frozen 4-bit quantized base; the adapter itself is not a quantization of LLaMA, so not quantized_from."}
{"child": "guanaco-7b", "parent": "oasst1", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/timdettmers/guanaco-7b/blob/cad9de16fb306d5bb1feb901333fe0aa7bd700d8/README.md", "note": "Card: 'on the OASST1 dataset'; paper abstract: Guanaco outperforms previous open models on the Vicuna benchmark. The card says evaluation used March 2023 ChatGPT/Bard outputs; that is evaluation, not training, so no influence edge."}
{"child": "mistral-7b-sft-beta", "parent": "mistral-7b-v0-1", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://huggingface.co/HuggingFaceH4/mistral-7b-sft-beta/blob/d7b9d226fb91df9de570ec0ee119f46e007f6142/README.md", "note": "Card: 'This model is a fine-tuned version of mistralai/Mistral-7B-v0.1 on the HuggingFaceH4/ultrachat_200k dataset.' Model description: 'Finetuned from model: mistralai/Mistral-7B-v0.1'. Uploader HuggingFaceH4 is the Zephyr developer."}
{"child": "mistral-7b-sft-beta", "parent": "ultrachat", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/HuggingFaceH4/mistral-7b-sft-beta/blob/d7b9d226fb91df9de570ec0ee119f46e007f6142/README.md", "note": "Card: fine-tuned on the ultrachat_200k dataset (H4's filtered subset of UltraChat; the `ultrachat` record covers the parent corpus, so this maps a filtered subset to its source). The distilled step: UltraChat text was written by ChatGPT (see the ultrachat record's distilled_from_outputs edge)."}
{"child": "mpt-7b-chat", "parent": "mpt-7b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://www.databricks.com/blog/mpt-7b", "note": "Blog: MPT-7B-Chat 'built by finetuning MPT-7B'; archived card: 'built by finetuning MPT-7B on ...'."}
{"child": "mpt-7b-chat", "parent": "sharegpt-vicuna", "relation": "trained_on", "evidence": "declared", "source": "http://web.archive.org/web/20230718105903/https://huggingface.co/mosaicml/mpt-7b-chat", "note": "Archived card names 'ShareGPT-Vicuna' (link jeffwan/sharegpt_vicuna, unreachable). The `sharegpt-vicuna` record is the LMSYS collection; mapping by name. Reviewer: confirm the id mapping."}
{"child": "mpt-7b-chat", "parent": "alpaca-52k", "relation": "trained_on", "evidence": "declared", "source": "http://web.archive.org/web/20230718105903/https://huggingface.co/mosaicml/mpt-7b-chat", "note": "Archived card lists Alpaca (link tatsu-lab/alpaca)."}
{"child": "openchat-3-5", "parent": "mistral-7b-v0-1", "relation": "fine_tuned_from", "evidence": "declared_by_uploader", "source": "https://huggingface.co/openchat/openchat_3.5/raw/0fc98e324280bc4bf5d2c30ecf7b97b84fb8a19b/config.json", "note": "No card sentence names the base. Evidence: config.json `_name_or_path: imone/Mistral_7B_with_EOT_token`, `model_type: mistral`, dims identical to Mistral 7B, card tag `mistral`; and the Starling-LM-7B-alpha card (a third party) says 'Openchat 3.5 (based on Mistral-7B-v0.1)'. Uploader imone is the OpenChat author, so this is the developer's own HF metadata, not a prose statement. The intermediate 'Mistral_7B_with_EOT_token' checkpoint (vocab +2) was not inspected."}
{"child": "openchat-3-5", "parent": "openorca", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/datasets/imone/OpenOrca_FLAN", "note": "Card lists imone/OpenOrca_FLAN; that dataset's card: 'the OpenOrca GPT4 subset with the original FLAN answers. Each even row contains the OpenOrca GPT4 answer, while each odd row contains the corresponding FLAN answer.' Mapped to the `openorca` record as a derived subset (a partial mapping; the GPT-4 half is the OpenOrca distillation of GPT-4)."}
{"child": "openorcaxopenchat-preview2-13b", "parent": "llama-2-13b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://huggingface.co/Open-Orca/OpenOrcaxOpenChat-Preview2-13B/blob/49d50e4348c4f454302e997fc39841f035738604/README.md", "note": "Card: 'We have used our own OpenOrca dataset to fine-tune Llama2-13B using OpenChat packing.'"}
{"child": "openorcaxopenchat-preview2-13b", "parent": "openorca", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/Open-Orca/OpenOrcaxOpenChat-Preview2-13B/blob/49d50e4348c4f454302e997fc39841f035738604/README.md", "note": "Card: 'trained on a curated filtered subset of most of our GPT-4 augmented data' from the OpenOrca dataset (its GPT-4 half distills GPT-4 outputs; see the openorca record). The subset, not the whole set, was used."}
{"child": "orca-2-7b", "parent": "llama-2-7b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://arxiv.org/abs/2311.11045", "note": "Paper 4.2: 'We start with LLaMA-2-7B or LLaMA-2-13B checkpoint and finetune it'; card: 'Orca 2 is a finetuned version of LLAMA-2'."}
{"child": "orca-2-7b", "parent": "flan-v2", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2311.11045", "note": "Orca 2 paper 4.2: 'We start with LLaMA-2-7B or LLaMA-2-13B checkpoint and finetune it on the train split of FLAN-v2 dataset for one epoch.'"}
{"child": "orca-2-7b", "parent": "orca-1-data", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2311.11045", "note": "Orca 2 paper 4.2: 'We then train on 5 million ChatGPT data from Orca 1 for 3 epochs. Then we train on the combination of 1 million GPT-4 data from Orca 1 and Orca 2's 817K data for 4 epochs.'"}
{"child": "orca-2-7b", "parent": "orca-2-data", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2311.11045", "note": "Same passage: 'Orca 2's 817K data'; section 4: 'we created a new dataset with ~817K training instances, which we will refer as Orca 2 dataset.'"}
{"child": "platypus2-13b", "parent": "llama-2-13b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://huggingface.co/garage-bAInd/Platypus2-13B/blob/dc1024c1b9df38f57f6436a02d31706cb0deaa01/README.md", "note": "Card: 'Platypus-13B is an instruction fine-tuned model based on the LLaMA2-13B transformer architecture'; 'instruction fine-tuned using LoRA on 1 A100 80GB'. Uploader garage-bAInd is the developer (card: trained by Cole Hunter & Ariel Lee). 'Based on ... architecture' alone would not show weight descent; the LoRA fine-tuning statement plus the Llama 2 tokenizer/config values do. The paper's abstract also describes 'fine-tuning and merging LoRA modules'; this card does not say whether a merge step produced this checkpoint."}
{"child": "platypus2-13b", "parent": "open-platypus", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/garage-bAInd/Platypus2-13B/blob/dc1024c1b9df38f57f6436a02d31706cb0deaa01/README.md", "note": "Card metadata and text: garage-bAInd/Open-Platypus. Uploader is the dataset's builder."}
{"child": "pythia-chat-base-7b-v0-16", "parent": "pythia-6-9b-deduped", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://github.com/togethercomputer/OpenChatKit/blob/main/README.md", "note": "OpenChatKit README: 'Pythia-Chat-Base-7B is a 7B-parameter fine-tuned variant of Pythia-6.9B-deduped from Eleuther AI'; the HF card says 'fine-tuned from EleutherAI's Pythia 7B' (no variant named). The README is what pins the deduped 6.9B."}
{"child": "pythia-chat-base-7b-v0-16", "parent": "oig", "relation": "trained_on", "evidence": "declared", "source": "https://github.com/togethercomputer/OpenChatKit/blob/main/README.md", "note": "README: 'The chat model was trained on the OIG dataset built by LAION, Together, and Ontocord.ai'; card: 43M-instruction collection (OIG-43M). The `oig` record covers the LAION OIG collection ('currently at 44M'); the 43M snapshot vs the current collection is not distinguished."}
{"child": "redpajama-incite-7b-instruct", "parent": "redpajama-incite-7b-base", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://huggingface.co/togethercomputer/RedPajama-INCITE-7B-Instruct/blob/7f36397b9985a3f981cdb618f8fec1c565ca5927/README.md", "note": "Card lists 'Base Model: RedPajama-INCITE-7B-Base' and says the model 'was fine-tuned for few-shot applications on the data of GPT-JT'. No `base_model` metadata; the relation rests on the card text."}
{"child": "redpajama-incite-7b-instruct", "parent": "redpajama-data-instruct", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/togethercomputer/RedPajama-INCITE-7B-Instruct/blob/7f36397b9985a3f981cdb618f8fec1c565ca5927/README.md", "note": "Card metadata lists togethercomputer/RedPajama-Data-Instruct; the dataset card describes it as P3 plus Natural Instruction, decontaminated against HELM, consistent with the model card's 'GPT-JT data minus HELM overlap'."}
{"child": "stablelm-tuned-alpha-7b", "parent": "stablelm-base-alpha-7b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://huggingface.co/stabilityai/stablelm-tuned-alpha-7b/blob/25071b093c15c0d1cb2b2876c6deb621b764fcf5/README.md", "note": "Card: 'a suite of 3B and 7B parameter decoder-only language models built on top of the StableLM-Base-Alpha models and further fine-tuned on various chat and instruction-following datasets.' Same layer/hidden/head/context values as the base record."}
{"child": "stablelm-tuned-alpha-7b", "parent": "alpaca-52k", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/stabilityai/stablelm-tuned-alpha-7b/blob/25071b093c15c0d1cb2b2876c6deb621b764fcf5/README.md", "note": "Card training-dataset section lists Alpaca first; uploader is the developer."}
{"child": "stablelm-tuned-alpha-7b", "parent": "gpt4all", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/stabilityai/stablelm-tuned-alpha-7b/blob/25071b093c15c0d1cb2b2876c6deb621b764fcf5/README.md", "note": "Card lists 'GPT4All Prompt Generations' and describes it as 'generated by GPT-4'. The GPT4All technical report (see the `gpt4all` record) says GPT-3.5-Turbo; the edge follows the dataset name, and the card's model attribution is flagged, not adopted."}
{"child": "stablelm-tuned-alpha-7b", "parent": "databricks-dolly-15k", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/stabilityai/stablelm-tuned-alpha-7b/blob/25071b093c15c0d1cb2b2876c6deb621b764fcf5/README.md", "note": "Card lists Databricks Dolly, 15k, and its metadata names HuggingFaceH4/databricks_dolly_15k (a re-host of databricks-dolly-15k)."}
{"child": "stablelm-tuned-alpha-7b", "parent": "sharegpt-vicuna", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/stabilityai/stablelm-tuned-alpha-7b/blob/25071b093c15c0d1cb2b2876c6deb621b764fcf5/README.md", "note": "Card: 'ShareGPT Vicuna (English subset)', metadata jeffwan/sharegpt_vicuna, which is now unreachable (HTTP 401), so identity with the `sharegpt-vicuna` record (LMSYS collection) is by name and could not be checked. Reviewer: confirm the id mapping."}
{"child": "starcoderbase", "parent": "starcoderdata", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/api/datasets/bigcode/starcoderdata", "note": "Dataset card (public API description): 'This is the dataset used for training StarCoder and StarCoderBase. It contains 783GB of code in 86 programming languages ... approximately 250 Billion tokens.' NOTE the 250B-token figure differs from the 1T tokens seen in training (multiple epochs per the paper); the dataset file itself is gated."}
{"child": "starling-rm-7b-alpha", "parent": "llama-2-7b-chat", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://huggingface.co/berkeley-nest/Starling-RM-7B-alpha/blob/6c6b4d5627834fe010d2c001632de2b94db81d66/README.md", "note": "Card: 'Starling-RM-7B-alpha is a reward model trained from Llama2-7B-Chat'; 'Finetuned from model: Llama2-7B-Chat'. Blog: 'Our reward model is fine-tuned from Llama2-7B-Chat'. The final layer is replaced by a scalar head, so weights below it descend."}
{"child": "starling-rm-7b-alpha", "parent": "nectar", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/berkeley-nest/Starling-RM-7B-alpha/blob/6c6b4d5627834fe010d2c001632de2b94db81d66/README.md", "note": "Card: 'We train the reward model with preference dataset berkeley-nest/Nectar'. Blog: 'trained with our K-wise loss on the Nectar dataset'."}
{"child": "tinyllama-1-1b-intermediate-step-1431k-3t", "parent": "llama-2-7b", "relation": "same_architecture_retrained", "evidence": "declared", "source": "https://arxiv.org/abs/2401.02385", "note": "Card: 'We adopted exactly the same architecture and tokenizer as Llama 2.' Paper: 'Following the same architecture and tokenizer as Llama 2, we name our model TinyLlama.' Relation chosen from those exact words (method.md: 'same' -> same_architecture_retrained). Parent id llama-2-7b stands for the Llama 2 design; the developer names Llama 2 generally, not a size. TENSION: the paper's own hyperparameters use grouped-query attention, which the llama-2-7b record notes is not used at 7B; and the size is 1.1B. If the reviewer reads 'exactly the same' as too strong, design_follows is the alternative."}
{"child": "tinyllama-1-1b-intermediate-step-1431k-3t", "parent": "slimpajama", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2401.02385", "note": "Card metadata `datasets: cerebras/SlimPajama-627B`; paper names SlimPajama as a primary source (GitHub subset removed). The `slimpajama` record has availability `removed` (HF API 401) — the corpus is recorded, not necessarily obtainable."}
{"child": "tinyllama-1-1b-intermediate-step-1431k-3t", "parent": "starcoderdata", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2401.02385", "note": "Card metadata `datasets: bigcode/starcoderdata`; paper: 'the training data of StarCoder', code-related samples only. `starcoderdata` is gated."}
{"child": "tulu-2-7b", "parent": "llama-2-7b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://huggingface.co/allenai/tulu-2-7b/blob/3c6e328ae91fabdd0daf09de16887de9615c1f66/README.md", "note": "Card: 'Tulu 2 7B is a fine-tuned version of Llama 2'; 'Finetuned from model: meta-llama/Llama-2-7b-hf'. Paper abstract: models 'finetuned on the V2 mixture'. Uploader allenai (Ai2) is the developer."}
{"child": "tulu-2-7b", "parent": "tulu-v2-sft-mixture", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/allenai/tulu-2-7b/blob/3c6e328ae91fabdd0daf09de16887de9615c1f66/README.md", "note": "Card metadata `datasets: allenai/tulu-v2-sft-mixture`; paper: finetuned on Tulu-V2-mix. Ai2 built the mixture. The distillation inside it (GPT4-Alpaca, OpenOrca GPT-4 subset) is recorded on the dataset."}
{"child": "openorca-platypus2-13b", "parent": "platypus2-13b", "relation": "merged_from", "evidence": "declared", "source": "https://huggingface.co/Open-Orca/OpenOrca-Platypus2-13B/blob/04e22880de5edcda7b86092242ac0834bf191190/README.md", "note": "Card: 'OpenOrca-Platypus2-13B is a merge of garage-bAInd/Platypus2-13B and Open-Orca/OpenOrcaxOpenChat-Preview2-13B.' The card does not state the merge method, weights or ratios. Parent 1 of 2. Trained by Cole Hunter & Ariel Lee per the card."}
{"child": "openorca-platypus2-13b", "parent": "openorcaxopenchat-preview2-13b", "relation": "merged_from", "evidence": "declared", "source": "https://huggingface.co/Open-Orca/OpenOrca-Platypus2-13B/blob/04e22880de5edcda7b86092242ac0834bf191190/README.md", "note": "Card: 'OpenOrca-Platypus2-13B is a merge of garage-bAInd/Platypus2-13B and Open-Orca/OpenOrcaxOpenChat-Preview2-13B.' The card does not state the merge method, weights or ratios. Parent 2 of 2. Trained by Open-Orca per the card. Preview2 was staged by the preparer so this edge can resolve."}
{"child": "starcoder", "parent": "starcoderbase", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://arxiv.org/abs/2305.06161", "note": "Paper abstract: 'We fine-tuned StarCoderBase on 35B Python tokens, resulting in the creation of StarCoder.' Release post: 'We fine-tuned StarCoderBase model for 35B Python tokens, resulting in a new model that we call StarCoder.' Developers' own words."}
{"child": "starcoder", "parent": "starcoderdata", "relation": "trained_on", "evidence": "declared", "source": "https://huggingface.co/api/datasets/bigcode/starcoderdata", "note": "Dataset card (public API description): 'the dataset used for training StarCoder and StarCoderBase'. StarCoder's own fine-tuning used the Python subset of that training data (paper sec. 5.6)."}
{"child": "starling-lm-7b-alpha", "parent": "openchat-3-5", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://huggingface.co/berkeley-nest/Starling-LM-7B-alpha/blob/1dddf3b95bc1391f6307299eb1c162c194bde9bd/README.md", "note": "Card: 'Finetuned from model: Openchat 3.5 (based on Mistral-7B-v0.1)'; 'a language model trained from Openchat 3.5 with reward model ... and policy optimization method APA'. Developers' own card. PARENT openchat-3-5 IS STAGED (promote it first)."}
{"child": "starling-lm-7b-alpha", "parent": "starling-rm-7b-alpha", "relation": "feedback_from", "evidence": "declared", "source": "https://huggingface.co/berkeley-nest/Starling-LM-7B-alpha/blob/1dddf3b95bc1391f6307299eb1c162c194bde9bd/README.md", "note": "Card: 'a language model trained from Openchat 3.5 with reward model berkeley-nest/Starling-RM-7B-alpha and policy optimization method APA'. Blog: 'we fine-tuned the Openchat 3.5 language model using the learned reward model.' Developers' own card and blog; the uploader IS the developer. Parent is a model whose scores, not text, shaped the policy. PARENT starling-rm-7b-alpha IS STAGED (promote it first; it in turn needs nothing unpromoted except openchat-3-5 for this record)."}
{"child": "tulu-2-dpo-7b", "parent": "tulu-2-7b", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://arxiv.org/abs/2311.10702", "note": "Paper abstract: 'a LLaMA-2 70B model finetuned on Tulu-V2-mix and further trained using direct preference optimization (DPO)'; Table 3 compares Tulu V2 models 'with and without DPO finetuning' at 7B, 13B, 70B. The paper states this for the family, not in a 7B-specific sentence, so the 7B DPO checkpoint's start point is read from that framing. Card metadata instead names only Llama-2-7b-hf as base_model (the original base, not the SFT stage)."}
{"child": "tulu-2-dpo-7b", "parent": "ultrafeedback", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2311.10702", "note": "Card metadata `datasets: HuggingFaceH4/ultrafeedback_binarized` (a filtered, binarized UltraFeedback); paper: DPO on 'a filtered and binarized form of UltraFeedback'. Recorded like Zephyr: trained_on the UltraFeedback record, which itself carries feedback_from gpt-4. Tulu's own paper notes UltraFeedback used TruthfulQA prompts (contamination caveat for evaluation)."}
{"child": "wizardcoder-15b-v1-0", "parent": "starcoder", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://arxiv.org/abs/2306.08568", "note": "Card fine-tuning section: 'We fine-tune StarCoder-15B with the following hyperparameters' and the reproduce command `--model_name_or_path \"bigcode/starcoder\"`; paper: 'we fine-tune StarCoder'."}
{"child": "wizardcoder-15b-v1-0", "parent": "code-evol-instruct", "relation": "trained_on", "evidence": "declared", "source": "https://arxiv.org/abs/2306.08568", "note": "Card: 'trained with 78k evolved code instructions'; paper: Code Alpaca evolved with Code Evol-Instruct, then StarCoder fine-tuned. The `code-evol-instruct` record has availability `unknown` (only third-party reproductions found); this edge is to that record, not to a released file."}
{"child": "zephyr-7b-beta", "parent": "mistral-7b-sft-beta", "relation": "fine_tuned_from", "evidence": "declared", "source": "https://huggingface.co/HuggingFaceH4/mistral-7b-sft-beta/blob/d7b9d226fb91df9de570ec0ee119f46e007f6142/README.md", "note": "Developer's card for the SFT model: 'It is the SFT model that was used to train Zephyr-7B-β with Direct Preference Optimization.' The Zephyr card's own 'fine-tuned version of Mistral-7B-v0.1' edge stays: this edge records the intermediate SFT stage, so Zephyr's path is Mistral-7B-v0.1 -> mistral-7b-sft-beta -> zephyr-7b-beta."}
