Dataset · gpt4all-j-prompt-generations
GPT4All-J prompt generations
Builder: Nomic AI
Availability: available · checked 2026-09-24 · source
Raw record: /data/datasets/gpt4all-j-prompt-generations.json
Fields
- id
- gpt4all-j-prompt-generations
- identifiers
- huggingface
- nomic-ai/gpt4all-j-prompt-generations
- builder
- Nomic AI
- release_date
- 2023-04-10partial · sourcenote: HF dataset repo creation date, not the builder's own release announcement.
- availability
- available · checked 2026-09-24 · source
- content
- Card: dataset used to train GPT4All-J and GPT4All-J-LoRA; versions v1.0 (original), v1.1-breezy (instances of 'AI language model' removed), v1.2-jazzy (also 'I'm sorry, I can't answer...' removed), v1.3-groovy (v1.2 with ShareGPT and Dolly added, ~8% semantic duplicates removed). Metadata lists 808,812 examples. GPT4All-J technical report: 'the 800k point GPT4All-J dataset that is a superset of the original 400k points GPT4All dataset', built also from 'custom-generated creative questions'; 'we have spent about $800 in OpenAI API credits so far to generate the training samples'; 'The assistant data was gathered from OpenAI's GPT3.5-Turbo'. The dataset card itself does not name a generator.recorded · sourcenote: The v1.3 additions (ShareGPT, Dolly) are not GPT-3.5-Turbo output as the report describes it; the edge below applies to the assistant data the report describes. Per-row provenance is in the `source` column, not read.
- primary_sources
- record_history
- date:2026-09-24 · change:created from primary sources (tranche 2 dataset records) · by:wilson-pruitt + claude ·date:2026-09-24 · change:content corrected from the GPT4All-J technical report (superset of gpt4all; GPT-3.5-Turbo assistant data) · by:wilson-pruitt + claude ·
Parents
Influence without weights
- distilled_from_outputs → ChatGPT (Nov. 2022 launch model) declared source Builders' report: 'The assistant data was gathered from OpenAI's GPT3.5-Turbo'; the dataset 'is a superset of the original 400k points GPT4All dataset'. v1.3 also adds ShareGPT and Dolly rows, which this edge does not cover. Snapshot not_recorded.
Children
Training data
- ← trained_on GPT4All-J declared source Card metadata `datasets: nomic-ai/gpt4all-j-prompt-generations`, and the report names the GPT4All-J dataset as its training set. Nomic is both uploader and developer, so declared. Default revision v1.0.
Read in
No station on the reading path has touched this record yet.