Model · gpt-3
GPT-3 (175B, 2020 paper)
Developer: OpenAI
Availability: never_released · checked 2026-09-24 · source
Weights never released; API access status not tracked by this field.
Raw record: /data/models/gpt-3.json
Fields
- id
- gpt-3
- developer
- OpenAI
- release_date
- 2020-05-28partial · sourcenote: arXiv v1 date of the paper.
- weights_status
- api_only
- availability
- never_released · checked 2026-09-24 · sourcenote: Weights never released; API access status not tracked by this field.
- license
- Proprietary; no weights released.partial · source
- architecture
- family
- decoder_only
- note
- Stub specimen, recorded because five open models name GPT-3's paper as their design reference. Paper sec. 2.1: 'We use the same model and architecture as GPT-2 ... with the exception that we use alternating dense and locally banded sparse attention patterns.' The paper describes 8 sizes; values are for 175B (Table 2.1).
- n_layers
- 96recorded · source
- hidden_size
- 12288recorded · source
- n_heads
- 96recorded · source
- vocab_size
- nullnot_recorded
- positional_encoding
- nullnot_recorded
- context_length
- 2048recorded · source
- training_data
- 300B tokens from a weighted mix of filtered Common Crawl, WebText2, Books1, Books2 and English Wikipedia.recorded · source
- techniques
- transformer-decoderbyte-pair-encoding
- primary_sources
- record_history
- date:2026-09-24 · change:created as a stub specimen to resolve design_follows edges · by:wilson-pruitt + claude ·
Parents
Design
- same_architecture_retrained → GPT-2 XL (1.5B) declared source Paper: 'the same model and architecture as GPT-2', except alternating dense and locally banded sparse attention. gpt2-xl stands in for the GPT-2 family.
Children
Design
- ← same_architecture_retrained GPT-Neo 2.7B declared source Card: 'designed using EleutherAI's replication of the GPT-3 architecture.' Like GPT-3, alternates dense (global) and local attention layers.
- ← design_follows GPT-NeoX-20B declared source Paper: architecture 'largely follows that of GPT-3'.
- ← design_follows Pythia 6.9B declared source Paper: 'Our model architecture and hyperparameters largely follow Brown et al. (2020)'.
- ← design_follows OPT 6.7B declared source Paper: trained 'to roughly match the performance and sizes of the GPT-3 class of models'; hyperparameters 'largely follow Brown et al. (2020)'.
- ← design_follows Falcon 7B declared source Card: 'broadly adapted from the GPT-3 paper'.
- ← design_follows falcon-40b declared_by_uploader source Card verified this session: 'The architecture is broadly adapted from the GPT-3 paper (Brown et al., 2020), with the following differences...' -- same wording as falcon-7b.
- ← design_follows Cerebras-GPT 13B declared source Paper: 'Cerebras-GPT models have a GPT-3-like architecture, an autoregressive transformer decoder model. The main difference is that unlike GPT-3, which uses alternating dense and sparse-banded attention, we use dense attention in all decoder blocks.' 'GPT-3-like' with a stated exception -> design_follows (not same_architecture_retrained: the developer names a difference and does not say 'same').
Read in
No station on the reading path has touched this record yet.