Model · gpt-j-6b
GPT-J 6B
Developer: EleutherAI (Ben Wang, Aran Komatsuzaki)
Availability: available · checked 2026-09-24 · source
Raw record: /data/models/gpt-j-6b.json
Fields
- id
- gpt-j-6b
- identifiers
- huggingface
- EleutherAI/gpt-j-6bEleutherAI/gpt-j-6B
- developer
- EleutherAI (Ben Wang, Aran Komatsuzaki)
- release_date
- 2021-06partial · sourcenote: Repo commit 'gpt-j-6b release' is 2021-06-09 UTC; the linked announcement post's URL is dated 2021-06-04. Month is certain, day is not.
- weights_status
- open
- availability
- available · checked 2026-09-24 · source
- license
- Apache 2.0recorded · sourcenote: Repo README: 'The weights of GPT-J-6B are licensed under version 2.0 of the Apache License.'
- architecture
- family
- decoder_only
- note
- Card: 28 layers, d_model 4096, d_ff 16384, 16 heads of 256 dims; RoPE applied to 64 dimensions of each head. Trained with Mesh Transformer JAX.
- n_layers
- 28recorded · source
- hidden_size
- 4096recorded · source
- n_heads
- 16recorded · source
- vocab_size
- 50400recorded · sourcenote: Card: embedding matrix is 50400 but only 50257 entries are used by the GPT-2 tokenizer.
- positional_encoding
- rotary (RoPE), partial: 64 of 256 dims per headrecorded · source
- context_length
- 2048recorded · source
- training_data
- The Pile (EleutherAI's ~800GB curated English corpus); 402B tokens over 383,500 steps.recorded · source
- techniques
- transformer-decoderrotary-position-embeddingbyte-pair-encoding
- primary_sources
- record_history
- date:2026-09-24 · change:created from primary sources (Phase 1 seed, batch 2) · by:wilson-pruitt + claude ·date:2026-09-24 · change:availability checked and recorded · by:wilson-pruitt + claude ·
Parents
Training data
Children
Weights descend
- ← fine_tuned_from Dolly v1 6B declared source Card: 'derived from EleutherAI's GPT-J (released June 2021) and fine-tuned on a ~52K record instruction corpus'; blog: '6 billion parameter model from EleutherAI'.
- ← fine_tuned_from GPT-JT-6B-v1 declared source Blog: 'A fork of GPT-J-6B, fine-tuned on 3.53 billion tokens'; card: 'a fork of EleutherAI's GPT-J (6B)'.
- ← fine_tuned_from GPT4All-J declared source Report abstract: 'deriving its weights from the Apache-licensed GPT-J model rather than the GPL-licensed of LLaMA'; card: 'Finetuned From: GPT-J'.
Design
- ← same_architecture_retrained GPT-NeoX-20B declared source Paper sec. 2.1: 'Our model architecture is almost identical to that of GPT-J', though it uses GPT-3 as the stated reference 'because there is no canonical published reference on the design of GPT-J.' No weights passed.
Read in
No station on the reading path has touched this record yet.