Pretraining
Dated sources are over 18 months old; other dates are unknown.
Publication dates and source age
Sources counted: 2
Newest dated source: 2024-07-31
Assessed at this edition's evidence cutoff: 2026-09-15. 18-month boundary: 2025-03-15.
Some publication dates are unknown; the newest dated source may not be the newest source overall.
Publication age does not tell us whether a claim is still valid. Reading an old source again does not make its publication date newer. An update date does not establish that the passage we used was updated.
Open in the glossary Reading notes · Structured record
In plain language
An initial training stage that provides a starting model for later use or adaptation. [1]
Reference this explanation or suggest a correction
Link to this explanation · Suggest a correction · How corrections work
Limits & distinctions
A pretrained language model is not automatically an instruction-following assistant. Pretraining and later adaptation serve different objectives. [2]
Reference this explanation or suggest a correction
Link to this explanation · Suggest a correction · How corrections work
A fuller explanation
For a language model, this often means adjusting weights on large text collections to predict missing or next text pieces. Later training can change how the model handles specific tasks. [1]
Reference this explanation or suggest a correction
Link to this explanation · Suggest a correction · How corrections work
How it relates to the map
Use this distinction when a proposal refers specifically to training new base models. [1]
Reference this explanation or suggest a correction
Link to this explanation · Suggest a correction · How corrections work
Share this page
https://theaiatlas.org/ideas/pretraining/
Download a share image · Vector image
Image previews are summaries. Keep the page link so readers can check the evidence.
Sources and what we read
1. How do Transformers work?
Publication dates and source age
Sources counted: 1
Publication dates are unavailable.
Assessed at this edition's evidence cutoff: 2026-09-15. 18-month boundary: 2025-03-15.
Some publication dates are unknown; the newest dated source may not be the newest source overall.
Publication age does not tell us whether a claim is still valid. Reading an old source again does not make its publication date newer. An update date does not establish that the passage we used was updated.
Read self-supervised language modeling and transfer learning. Used for the pretraining distinction, without adopting historical dates, performance comparisons or claims of understanding from this teaching page.
2. The Llama 3 Herd of Models
Source is over 18 months old.
Publication dates and source age
Sources counted: 1
Newest dated source: 2024-07-31
Assessed at this edition's evidence cutoff: 2026-09-15. 18-month boundary: 2025-03-15.
Publication age does not tell us whether a claim is still valid. Reading an old source again does not make its publication date newer. An update date does not establish that the passage we used was updated.
Read the two training stages, post-training methods and section 5.4.8 limitations in version 2. Dates follow arXiv submission history. One documented approach, not a universal training recipe.
Edition and machine-readable evidence
Content version 0.20.0. Evidence cutoff 2026-09-15; this does not mean every source was read on that day.
Pinned complete dataset · Complete evidence page · Agent consumption guide