A field guide to AI positions

Pretraining

Dated sources are over 18 months old; other dates are unknown.

Publication dates and source age

Sources counted: 2

Newest dated source: 2024-07-31

Assessed at this edition's evidence cutoff: 2026-09-15. 18-month boundary: 2025-03-15.

Some publication dates are unknown; the newest dated source may not be the newest source overall.

Publication age does not tell us whether a claim is still valid. Reading an old source again does not make its publication date newer. An update date does not establish that the passage we used was updated.

Open in the glossary Reading notes · Structured record

In plain language

An initial training stage that provides a starting model for later use or adaptation. [1]

Reference this explanation or suggest a correction

Link to this explanation · Suggest a correction · How corrections work

Limits & distinctions

A pretrained language model is not automatically an instruction-following assistant. Pretraining and later adaptation serve different objectives. [2]

Reference this explanation or suggest a correction

Link to this explanation · Suggest a correction · How corrections work

A fuller explanation

For a language model, this often means adjusting weights on large text collections to predict missing or next text pieces. Later training can change how the model handles specific tasks. [1]

Reference this explanation or suggest a correction

Link to this explanation · Suggest a correction · How corrections work

How it relates to the map

Use this distinction when a proposal refers specifically to training new base models. [1]

Reference this explanation or suggest a correction

Link to this explanation · Suggest a correction · How corrections work

Share this page

https://theaiatlas.org/ideas/pretraining/

Download a share image · Vector image

Image previews are summaries. Keep the page link so readers can check the evidence.

Sources and what we read

  1. 1. How do Transformers work?

    Publication dates and source age

    Sources counted: 1

    Publication dates are unavailable.

    Assessed at this edition's evidence cutoff: 2026-09-15. 18-month boundary: 2025-03-15.

    Some publication dates are unknown; the newest dated source may not be the newest source overall.

    Publication age does not tell us whether a claim is still valid. Reading an old source again does not make its publication date newer. An update date does not establish that the passage we used was updated.

    Read self-supervised language modeling and transfer learning. Used for the pretraining distinction, without adopting historical dates, performance comparisons or claims of understanding from this teaching page.

  2. 2. The Llama 3 Herd of Models

    Source is over 18 months old.

    Publication dates and source age

    Sources counted: 1

    Newest dated source: 2024-07-31

    Assessed at this edition's evidence cutoff: 2026-09-15. 18-month boundary: 2025-03-15.

    Publication age does not tell us whether a claim is still valid. Reading an old source again does not make its publication date newer. An update date does not establish that the passage we used was updated.

    Read the two training stages, post-training methods and section 5.4.8 limitations in version 2. Dates follow arXiv submission history. One documented approach, not a universal training recipe.

Edition and machine-readable evidence

Content version 0.20.0. Evidence cutoff 2026-09-15; this does not mean every source was read on that day.

Pinned complete dataset · Complete evidence page · Agent consumption guide