Records
96 models, 33 datasets, and 154 edges. Each record page shows every field with its value, status, and source, then its parents and children. Every edge carries an evidence tag:
Definitions are in the method. The same data as JSON: each record at /data/models/<id>.json or /data/datasets/<id>.json, and every edge in edges.jsonl.
Models
- Alpaca 7B
- alpaca-lora-7b
- Baize v2 7B
- bloom
- BLOOM 7B1
- bloom-560m
- bloomz-7b1
- Cerebras-GPT 13B
- ChatGPT (Nov. 2022 launch model)
- Code Llama 7B
- CodeLlama-7b-Instruct-hf
- CodeLlama-7b-Python-hf
- Dolly v1 6B
- dolly-v2-12b
- Falcon 7B
- falcon-40b
- falcon-40b-instruct
- falcon-7b-instruct
- GPT-2 XL (1.5B)
- GPT-3 (175B, 2020 paper)
- GPT-4 (Mar. 2023)
- GPT-J 6B
- GPT-JT-6B-v1
- GPT-Neo 2.7B
- gpt-neo-1.3B
- gpt-neo-125m
- GPT-NeoX-20B
- gpt2
- gpt2-large
- gpt2-medium
- GPT4All-J
- guanaco-7b
- Llama 2 7B
- Llama 2-Chat 7B
- LLaMA 7B
- Llama-2-13b-hf
- Llama-2-70b-hf
- Mistral 7B (v0.1)
- Mistral 7B SFT β
- Mistral-7B-Instruct-v0.1
- Mistral-7B-Instruct-v0.1-GPTQ
- Mistral-7B-OpenOrca
- Mixtral-8x7B-Instruct-v0.1
- Mixtral-8x7B-v0.1
- MPT-7B
- MPT-7B-Chat
- mpt-7b-instruct
- MythoLogic-L2-13b
- MythoMax-L2-13b
- Nous-Hermes-llama-2-7b
- open_llama_7b
- OpenChat 3.5
- OpenHermes-2.5-Mistral-7B
- OpenOrca-Platypus2-13B
- OpenOrcaxOpenChat-Preview2-13B
- OPT 6.7B
- opt-1.3b
- opt-125m
- opt-13b
- opt-30b
- opt-66b
- Orca 2 7B
- phi-2
- Platypus2-13B
- Pythia 6.9B
- pythia-1.4b
- pythia-12b
- pythia-160m
- pythia-1b
- pythia-2.8b
- pythia-410m
- pythia-6.9b-deduped
- pythia-70m
- Pythia-Chat-Base-7B-v0.16
- Qwen-7B
- RedPajama-INCITE-7B-Base
- RedPajama-INCITE-7B-Instruct
- SOLAR-10.7B-v1.0
- stablelm-base-alpha-7b
- StableLM-Tuned-Alpha-7B
- StarCoder
- StarCoderBase
- Starling-LM-7B-alpha
- Starling-RM-7B-alpha
- text-davinci-003
- TinyLlama-1.1B (intermediate-step-1431k-3T)
- Tulu 2 7B
- Tulu 2 DPO 7B
- Vicuna 7B v1.3
- vicuna-13b-v1.5
- vicuna-7b-v1.5
- WizardCoder-15B-V1.0
- WizardLM-13B-V1.2
- xgen-7b-8k-base
- Zephyr 7B β
- zephyr-7b-alpha
Datasets
- Alpaca instruction data (52K)
- Alpaca-Cleaned
- Baize SDF data (ChatGPT-ranked self-generations)
- Baize self-chat data
- Code Evol-Instruct (WizardCoder)
- databricks-dolly-15k
- FLAN-v2 Collection (Flan 2022)
- GPT4All prompt generations
- GPT4All-J prompt generations
- Nectar
- Open Instruction Generalist (OIG)
- Open-Platypus
- OpenAssistant Conversations (OASST1)
- OpenHermes 2.5
- OpenOrca
- Orca 1 explanation-tuning data (FLAN-5M / FLAN-1M)
- Orca 2 dataset (~817K)
- RedPajama-Data-1T
- RedPajama-Data-Instruct
- RefinedWeb
- ROOTS corpus
- ShareGPT conversations (LMSYS collection)
- SlimPajama
- StarCoderData
- The Pile
- The Pile (deduplicated)
- The Stack (near-deduplicated)
- Tulu V2 SFT mixture
- UltraChat
- UltraFeedback
- WebText
- WizardLM Evol-Instruct 70k
- xP3