Evidence cutoff: 2026-09-15. Positions are editorial interpretations, not endorsements by the cited actors. Stated policy is separate from verified implementation.
The research and writing use AI assistance. There has been no independent human fact-check. Read the research method
Actor 01 · person
Eliezer Yudkowsky
Editorial anchor: x -87, y 87. Ranges: x [-100, -72], y [72, 100]. Layout units, not percentages.
Evidence clarity: high · Implementation: public-advocacy · Reviewed: 2026-09-15
Scope: Superintelligence and the frontier leading to it.
Sources used: First-hand material.
This is his advocated position, not this project's prediction.
In depth
Editorial interpretation of the linked sources.
Argues against building superintelligence with present methods. [18]
The claim being mapped
The book announcement presents an argument that superhuman AI would threaten humanity. The map records that argument; it does not present the outcome as established. [18]
pace: Argues against proceeding to superintelligence with present methods. [yudkowsky-book]
concern: Frames superhuman AI as an extinction threat. [yudkowsky-book]
context: MIRI separately advocates an enforced international ASI moratorium. [miri-position]
Editorial anchor: x -87, y 66. Ranges: x [-100, -65], y [42, 94]. Layout units, not percentages.
Evidence clarity: high · Implementation: public-advocacy · Reviewed: 2026-09-15
Scope: Its April 2026 proposal for a global pause on the most powerful general-AI training.
Sources used: First-hand material.
A movement has internal variation; this point represents its published proposal.
In depth
Editorial interpretation of the linked sources.
Calls for a coordinated global pause on the most powerful AI training. [31]
The proposal's limits
Its April 2026 proposal calls for safety and democratic-control conditions before more powerful general AI is developed. It distinguishes that target from narrow applications such as medical image recognition. [31]
pace: Advocates a coordinated global pause rather than a company pausing on its own. [pauseai-proposal-2026]
Find the passage: April 5, 2026 proposal; opening and Treaty Measures
concern: Its proposal treats dangerous AI progress as a reason for international restraint. [pauseai-proposal-2026]
Find the passage: April 5, 2026 proposal; opening and Treaty Measures
Editorial anchor: x -62, y 50. Ranges: x [-85, -35], y [25, 85]. Layout units, not percentages.
Evidence clarity: medium · Implementation: public-advocacy · Reviewed: 2026-09-15
Scope: Superintelligence restriction, not a ban on every AI application or safety research.
Sources used: First-hand material and reporting or indirect copies.
Conditional superintelligence restraint is reported separately from his directly read account of control risks. His research proposal is not a proven safety solution.
In depth
Editorial interpretation of the linked sources.
Supports restraint on superintelligence and research aimed at retaining human control. [22][34]
Two different kinds of evidence
The prohibition endorsement is reported; his concern about loss of control is stated directly in his LawZero announcement. Neither establishes a numerical risk estimate. [22][34]
pace: Reported signatory of the conditional superintelligence prohibition. [bengio-signature]
pace: The statement makes safety consensus and public support conditions for lifting the prohibition. [superintelligence-statement]
concern · support: Warns about deception, self-preservation and loss of human control while proposing safety research. [bengio-lawzero-2025]
Editorial anchor: x -28, y 81. Ranges: x [-48, -8], y [55, 96]. Layout units, not percentages.
Evidence clarity: high · Implementation: public-advocacy · Reviewed: 2026-09-15
Scope: Frontier capability growth, especially unchecked self-improvement.
Sources used: First-hand material.
Pacing is not stopping. Coordinated steps remain proposals; no independent operational audit is claimed.
In depth
Editorial interpretation of the linked sources.
Calls for slower frontier capability growth while continuing AI development. [10]
What pacing means here
His September essay proposes outside evaluators and coordination among companies and governments. It argues that extra time could improve safeguards; it does not call for all AI work to stop. [10]
Editorial anchor: x 0, y 51. Ranges: x [-35, 45], y [32, 86]. Layout units, not percentages.
Evidence clarity: medium · Implementation: published-policy-and-announced-commitment · Reviewed: 2026-09-15
Scope: Published company safeguards and its announced embedded-evaluator commitment.
Sources used: First-hand material.
Giving outside evaluators access does not commit the company to slow all frontier development. The Responsible Scaling Policy separates company promises from industry recommendations; implementation is not independently audited.
In depth
Editorial interpretation of the linked sources.
Promises outside scrutiny and conditional safeguards while continuing development. [10][33]
Why the position was reassessed
The previous dot treated evaluator access as pace restraint. Reading the company policy separately supports conditional continuation with a wide range. This is an editorial correction, not evidence of a change of mind. [33][10]
What can be checked
The policy publishes risk-report and external-review requirements. Announcing a requirement does not establish that it was met in a particular case. [33]
pace · support: Appendix A promises development delays in specified competitor-related circumstances, not a general training halt. [anthropic-rsp-3-4]
Find the passage: Appendix A, Anthropic in the lead and Competitors have strong safety measures
pace · counterpoint: The company commits to work with outside evaluators. Wider limits on developing more capable models require coordination. [amodei-pacing]
Find the passage: Three-step plan and Embedded Evaluators
concern · support: The policy assesses catastrophic misuse and sabotage that could increase later global-catastrophe risk. [anthropic-rsp-3-4]
Find the passage: Section 1 and risk-report requirements
Editorial anchor: x -3, y 65. Ranges: x [-30, 25], y [38, 86]. Layout units, not percentages.
Evidence clarity: medium · Implementation: public-advocacy · Reviewed: 2026-09-15
Scope: Public governance proposal and subsequent endorsement.
Sources used: First-hand material and reporting or indirect copies.
Do not convert a personal endorsement into a verified Google DeepMind-wide slowdown.
In depth
Editorial interpretation of the linked sources.
Proposes shared frontier safeguards and a possible coordinated slowdown. [13]
A conditional proposal
His July framework combines support for innovation with a standards body that could coordinate a slowdown if necessary. That proposal does not establish a Google DeepMind-wide operational change. [13]
pace: Proposes standards that could coordinate a slowdown if necessary. [hassabis-framework]
concern: Warns that safeguards and understanding are not keeping pace. [hassabis-framework]
context: Reportedly endorsed the direction of Amodei's proposal. [september-responses]
Editorial anchor: x -8, y 29. Ranges: x [-32, 18], y [12, 76]. Layout units, not percentages.
Evidence clarity: medium · Implementation: public-advocacy · Reviewed: 2026-09-15
Scope: Reported September 2026 pacing support, with a directly read earlier governance proposal.
Sources used: First-hand material and reporting or indirect copies.
The original social posts were unavailable to this review; the reporting is linked.
In depth
Editorial interpretation of the linked sources.
Supports pacing in September reporting; earlier coauthored proposals also considered frontier limits. [12][47]
What is direct and what is reported
The 2023 article is directly available. The newer social response remains indirectly verified through reporting; the older proposal does not make that post independently verified. [47][12]
context · context: His coauthored 2023 proposal considered limits on frontier capability growth while exempting lower-capability systems. [altman-governance-2023]
Find the passage: A starting point; What's not in scope
Editorial anchor: x -11, y 9. Ranges: x [-35, 20], y [0, 66]. Layout units, not percentages.
Evidence clarity: medium · Implementation: company-reported-restraint · Reviewed: 2026-09-15
Scope: Selected frontier research workloads and controls.
Sources used: First-hand material.
The company reports selective restrictions, not a full training halt. Concern level remains an editorial inference.
In depth
Editorial interpretation of the linked sources.
Reports selective research pauses while upgrading safeguards. [11]
How to read the company claim
The company describes pausing certain tool-using research workloads, then resuming some under stronger controls. This supports targeted restraint; it does not establish a company-wide halt. [11]
pace: Reports restrictions on selected frontier research workloads pending stronger safeguards. [openai-pacing]
concern: Treats stronger cyber capabilities and alignment failures as requiring safeguards. [openai-pacing]
Editorial anchor: x 15, y 50. Ranges: x [-75, 85], y [10, 85]. Layout units, not percentages.
Evidence clarity: medium · Implementation: public-advocacy · Reviewed: 2026-09-15
Scope: Personal statements combining frontier scaling, concern about human control and reported support for pacing.
Sources used: First-hand material and reporting or indirect copies.
The September endorsement is indirect and brief. Building plans indicate a personal preference, not a general regulatory policy or a company safety rating. The conflicting signals warrant a broad pace range.
In depth
Editorial interpretation of the linked sources.
Scaling ambition and concern about control coexist in the reviewed statements. [4][12]
Why the anchor changed
The earlier dot represented only a reported endorsement. This review adds direct building and control statements, not a spending-based score. Its position near the middle of the pace axis summarizes conflicting signals; it does not measure a change in Musk's beliefs. [4][12]
The proposed safeguard is an argument
Musk argues that curiosity and truth-seeking could preserve humanity. The interviewer challenges that connection. The transcript does not establish that the proposed values reliably produce safe behavior. [4]
Keep the historical record separate
The 2023 discussion already combines building with a safety rationale. It also discusses delayed openness. Neither that interview nor the newer endorsement establishes a present company-wide halt. [5][12]
pace · support: Advocates expanding intelligence through Grok's mission and describes removing bottlenecks to rapid AI-compute growth. [musk-dwarkesh-2026]
Find the passage: Scaling, 00:00:00–00:36:46; Grok's mission, from 00:36:46
concern · support: Questions whether humans could remain in charge of much smarter AI; wants human survival but acknowledges his proposed values are no guarantee. [musk-dwarkesh-2026]
Find the passage: Grok and alignment, 00:36:46–00:59:56
pace · counterpoint: The September roundup reports his endorsement of Amodei's pacing essay. This qualifies a simple acceleration-only reading. [september-responses]
Find the passage: Elon Musk response; linked post not independently retrieved
context · context: In 2023, described rapid compute expansion alongside a preference for delayed open sourcing and a pro-human safety rationale. [musk-lex-2023]
Editorial anchor: x 20, y 46. Ranges: x [-10, 47], y [20, 80]. Layout units, not percentages.
Evidence clarity: medium · Implementation: published-framework · Reviewed: 2026-09-15
Scope: Institutional frontier-safety framework.
Sources used: First-hand material.
Its framework supports conditional progress. No whole-lab slowdown is established by this evidence.
In depth
Editorial interpretation of the linked sources.
Links continued model development to capability-based safeguards. [14]
How to read the company claim
Its updated framework includes safety-case reviews for some external launches and large-scale internal uses. A requirement in a framework is distinct from proof that a specific deployment met it. [14]
pace: Continues frontier development under capability-linked safeguards. [deepmind-framework]
Editorial anchor: x 38, y 12. Ranges: x [0, 68], y [0, 59]. Layout units, not percentages.
Evidence clarity: medium · Implementation: announced-commitment · Reviewed: 2026-09-15
Scope: Microsoft AI model development, not the entire Microsoft group.
Sources used: First-hand material.
A human-control code is not a catastrophe probability. The consultation is not a verified slowdown.
In depth
Editorial interpretation of the linked sources.
Plans continued model development under human-control commitments. [15]
How to read the company claim
The code describes limits and evaluation goals for MAI models. A public consultation and a published rule do not show how reliably a model follows the rule. [15]
pace: Plans continued model training and deployment guided by its code. [microsoft-code]
concern: Makes retaining human control a central requirement. [microsoft-code]
Editorial anchor: x 65, y 25. Ranges: x [15, 84], y [0, 67]. Layout units, not percentages.
Evidence clarity: medium · Implementation: published-framework · Reviewed: 2026-09-15
Scope: Meta's frontier development and published safeguards.
Sources used: First-hand material.
Publishing a risk framework does not reveal a probability belief; nor does openness imply low concern.
In depth
Editorial interpretation of the linked sources.
Continues advanced-model development with stated release safeguards. [16]
How to read the company claim
The framework addresses severe misuse and loss of control across open and closed releases. This is the company's account of its safeguards, not an independent safety finding. [16]
pace: Describes continued scaling with release safeguards. [meta-framework]
concern: Explicitly addresses catastrophic risks and loss of control. [meta-framework]
Editorial anchor: x 43, y -48. Ranges: x [20, 75], y [-90, -12]. Layout units, not percentages.
Evidence clarity: medium · Implementation: public-advocacy · Reviewed: 2026-09-15
Scope: His public views, separate from Meta's institutional policy.
Sources used: First-hand material and reporting or indirect copies.
Skepticism about catastrophe is not a claim that all AI harms are negligible.
In depth
Editorial interpretation of the linked sources.
Disputes takeover arguments and favors broad AI development and access. [23]
What his skepticism covers
In the 2024 interview, he argues that progress and safeguards can develop gradually. He also describes concentrated control of AI as a serious concern. Skepticism about extinction is not dismissal of every harm. [23]
pace: Favors broad AI development and access over catastrophe-driven restrictions. [lecun-interview]
Editorial anchor: x 77, y -74. Ranges: x [55, 100], y [-98, -33]. Layout units, not percentages.
Evidence clarity: medium · Implementation: public-advocacy · Reviewed: 2026-09-15
Scope: His dated 2023 manifesto, with a 2026 interview-publisher summary for context.
Sources used: First-hand material and reporting or indirect copies.
The detailed catastrophe argument remains the 2023 manifesto. A newer episode summary supports continued pro-growth advocacy but cannot substitute for a full updated two-axis interview review.
In depth
Editorial interpretation of the linked sources.
Argues for faster technological development and against catastrophe-driven restraint. [17]
A dated argument with newer context
The 2023 manifesto supplies the detailed argument. A June 2026 episode description reports continued enthusiasm for AI growth; it is context, not a newly reviewed statement about every catastrophic-risk scenario. [17][48]
pace: Explicitly opposes deceleration and favors technological growth. [andreessen-manifesto]
concern: Rejects catastrophe-oriented arguments for restricting progress. [andreessen-manifesto]
context · context: The interview publisher describes continued enthusiasm for AI growth and concern about policy barriers. [andreessen-interview-summary-2026]
Find the passage: Episode description; audio not independently reviewed
The previously used manifesto explicitly disclaims being the firm's view. A firm-level position on both axes has not been established.
In depth
Editorial interpretation of the linked sources.
Publishing a founder's essay does not establish the investment firm's position. [17]
Why this record has no dot
The manifesto remains evidence for Marc Andreessen's personal record. Its footer prevents us from treating the same text as the firm's stance on both axes. [17]
context · counterpoint: The page attributes its views to individual personnel and expressly excludes the firm and affiliates. [andreessen-manifesto]
Publishing model weights and benchmark claims does not establish a preferred frontier pace or catastrophic-risk stance. A two-axis position remains unplaced.
In depth
Editorial interpretation of the linked sources.
Publishes AI models; the reviewed release does not establish both map axes. [45]
What the source can establish
The company announcement describes a model release. Model access, performance claims and corporate ideology are different questions. [45]
context · context: Announces R1 model access and release of weights and related materials. [deepseek-r1-release]
Editorial anchor: x 85, y -80. Ranges: x [65, 100], y [-100, -50]. Layout units, not percentages.
Evidence clarity: medium · Implementation: public-advocacy-and-signed-policy · Reviewed: 2026-09-15
Scope: His September 2026 opposition to slowing advanced AI, read alongside his June executive order.
Sources used: First-hand material and reporting or indirect copies.
September posts are verified through reporting, not the original Truth Social pages. This maps expressed catastrophic-risk concern, not private beliefs, technical expertise, or every administration policy. The ranges are editorial judgments.
In depth
Editorial interpretation of the linked sources.
Favors continued AI development and publicly dismisses AI-extinction warnings; his signed policy still includes security measures. [1][2][3]
Why the dot is at the lower right
The September reporting supports a strong preference against slowing AI and low expressed concern about human extinction. The coordinates summarize this public stance; they do not measure risk or establish membership in e/acc. [1][2]
What the latest posts establish
Forbes and AFP report September 14 posts opposing new guardrails and framing AI as an economic and competitive priority. These are attributed political claims, not evidence that advanced AI is safe or that new safeguards would cause bankruptcy. [1][2]
Security policy is a qualification
The June 2 order calls for cyber-defense measures, capability benchmarks and voluntary early government access to some frontier models. These provisions qualify any claim that he rejects all safeguards. A signed direction does not establish that the measures have been implemented or work. [3]
How current and direct is this review?
Reviewed September 15, 2026. The September posts were read through reporting because Truth Social did not expose their text to this review. The June order was read directly. This is a dated selection, not an exhaustive archive of his posts or a separate assessment of the US government. [1][2][3]
pace · support: Forbes reports opposition to calls for an AI slowdown, emphasizing competition with China and the costs of regulation. [trump-forbes-2026-09-14]
Find the passage: Key Facts and Trump Dismisses AI Concerns; September 14 posts and September 13 remarks
concern · support: AFP reports that he dismissed scenarios of AI destroying humanity as a hoax. [trump-afp-2026-09-14]
Find the passage: Opening paragraphs on September 14 Truth Social posts
pace · support: His signed order favors developing advanced AI and explicitly excludes mandatory model licensing or preclearance under its frontier-model section. [trump-ai-security-order-2026]
Find the passage: Executive Order 14409, sections 1 and 3(c)
context · counterpoint: The same order directs cybersecurity measures and a voluntary process for evaluating frontier models before wider release. [trump-ai-security-order-2026]
Editorial anchor: x 65, y -20. Ranges: x [25, 90], y [-55, 35]. Layout units, not percentages.
Evidence clarity: medium · Implementation: public-advocacy · Reviewed: 2026-09-15
Scope: His arguments for continued AI development and against concentrated power used to suppress it.
Sources used: First-hand material.
Lower catastrophic-risk emphasis is relative to his concern about political control, not a claim that he rules out AI catastrophe. His economic and theological arguments do not specify a detailed frontier-training policy.
In depth
Editorial interpretation of the linked sources.
Favors AI development while placing greater emphasis on the danger of concentrated political control. [6][8]
Why this position is on the right
His opposition to stopping AI supports the development side of this map. A moderate rather than extreme anchor reflects the absence of a detailed training timetable or unrestricted-deployment proposal in the reviewed material. [6][8]
Why lower concern does not mean no concern
His Hoover conversation calls for concern about both catastrophe and totalitarian control. The vertical range crosses the middle because the relative emphasis is clearer than an absolute level of AI-risk concern. [7]
What the theology does and does not establish
His Antichrist framing is a speculative argument about power gained through fear of catastrophe. The Hoover conversation was recorded in October 2024; it does not establish a current policy proposal or the truth of a prophecy. [7][9]
An economic reservation
In the December interview he worries about benefits accruing to a few firms and AI replacing workers. These reservations complicate a blanket techno-optimist label without establishing support for a frontier pause. [8]
pace · support: Opposes the power needed to stop AI and favors AI even under a labor-substitution scenario. [thiel-tyler-2024]
Find the passage: Q&A: human extinction and technological replacement of workers
concern · support: Explicitly prioritizes concern about humans stopping AI over AI destroying humanity; rejects treating the outcome as predetermined. [thiel-tyler-2024]
Find the passage: Q&A: When do you think humans are going to destroy themselves?
concern · counterpoint: Says both technological catastrophe and a totalitarian world state deserve concern, while prioritizing the latter. [thiel-hoover-apocalypse]
Find the passage: Scylla and Charybdis discussion; search for worry about both
pace · support: Favors embracing AI as a source of growth while acknowledging concentrated returns and displacement of labor. [thiel-spectator-2025]
Find the passage: John Power's AI bubble and abundance question
context · context: A 2026 report describes his continuing use of the Antichrist framework in discussing political power. [thiel-spectator-2026]
Find the passage: Cambridge talk account; an edited report, not a complete transcript
Editorial anchor: x -50, y 75. Ranges: x [-80, -15], y [50, 95]. Layout units, not percentages.
Evidence clarity: medium · Implementation: public-advocacy · Reviewed: 2026-09-15
Scope: Conditional controls on exceptionally capable AI in a coauthored policy paper.
Sources used: First-hand material and reporting or indirect copies.
The proposal is collective and dated. Its control-risk warning is not a probability estimate or a claim that every kind of AI should stop.
In depth
Editorial interpretation of the linked sources.
Advocates conditional restraints on highly capable AI and warns about losing control. [35][37]
Reading the placement
This point summarizes the coauthored proposal; its range allows different readings. [35]
Newer context
The December 2025 publisher summary describes his concern about control. His proposed protective instincts are an idea, not a demonstrated safeguard. [37]
pace · support: Coauthors a call for conditional development halts when dangerous capabilities emerge. [extreme-ai-risks-2024]
Find the passage: Mitigation
concern · support: Coauthors a warning about irreversible loss of control and human extinction. [extreme-ai-risks-2024]
Find the passage: Societal-scale risks
context · context: The interview publisher reports continued concern about humans losing control. [hinton-gzero-2025]
Editorial anchor: x -50, y 75. Ranges: x [-80, -15], y [50, 95]. Layout units, not percentages.
Evidence clarity: medium · Implementation: public-advocacy · Reviewed: 2026-09-15
Scope: Conditional limits on dangerous frontier development and evidence required before release.
Sources used: First-hand material.
A coauthored policy proposal and dated testimony support this interpretation. They do not establish a blanket ban or an institutional Berkeley position.
In depth
Editorial interpretation of the linked sources.
Calls for enforceable safety conditions and warns about loss of human control. [35][36]
What the conditions mean
His testimony puts the burden on developers to show safety before release. It also discusses present harms; concern about catastrophe does not replace those issues. [36]
pace · support: Coauthors conditional development halts pending adequate protections. [extreme-ai-risks-2024]
Find the passage: Mitigation
concern · support: His testimony warns that AGI without reliable human control could threaten human survival. [russell-senate-2023]
Unplaced: this review cannot summarize the different risks she discusses with one concern coordinate.
In depth
Editorial interpretation of the linked sources.
Questions the pursuit of AGI and argues for research on defined tasks. [38]
Why there is no dot
The coauthored paper supports a research focus on defined tasks. Its wider historical thesis is not adopted as an atlas conclusion. The concern evidence below discusses different types of harm, so this review leaves the point unplaced. [38][179]
pace · support: Recommends research on defined, testable tasks instead of trying to build an all-purpose AGI. [gebru-torres-2024]
Find the passage: Journal abstract
context · context: The authors emphasize harms to marginalized people and concentration of power. [gebru-torres-2024]
Find the passage: Journal abstract
concern · context: Rejects rogue-machine extinction stories while naming weapons and climate as serious threats. [gebru-wired-2026]
Find the passage: Answers about extinction scenarios and present threats
Resistance to adopting a product is distinct from a preferred pace of frontier capability growth. This source does not settle both map axes.
In depth
Editorial interpretation of the linked sources.
Asks who benefits from AI systems and whether their claims can be checked. [39]
Questions a reader can use
Her slides ask what task a system performs, whether its inputs support accurate output, and whether it can be used to deny people's rights. These questions apply to evaluating products; they do not automatically locate a person on this map. [39]
context · context: Urges scrutiny of task definitions, training data, claimed accuracy, labor and surveillance risks. [bender-humanities-2025]
The briefing supports investment and safeguards but does not establish a two-axis frontier-pacing position. It expressly excludes attribution to affiliated organizations.
In depth
Editorial interpretation of the linked sources.
Supports public AI research, wider access and evidence-based safeguards. [40]
A different emphasis
Her briefing focuses on who can develop and benefit from AI, global collaboration and misuse. Those subjects matter beyond the two questions represented by this chart. [40]
context · context: Calls for public investment, wider access and evidence-based governance while warning about misuse. [fei-fei-li-un-2024]
Find the passage: Public Sector Leadership and Science- and Evidence-Based AI Policymaking
The keynote focuses on privacy and power. It does not provide enough evidence for both frontier-pacing and catastrophe-concern coordinates.
In depth
Editorial interpretation of the linked sources.
Focuses on privacy, surveillance and who controls digital infrastructure. [41]
Why privacy belongs in the wider debate
Her argument asks whether people can influence how technology affects their lives. It does not establish that every AI system has the same business model or that privacy concerns imply one catastrophe-risk position. [41]
context · context: Argues that large-scale commercial AI can reinforce data collection and concentrated corporate power. [whittaker-ndss-2024]
An enforcement and innovation mandate is not a single ideology or a measured level of catastrophic-risk concern. The two-axis position is not established.
In depth
Editorial interpretation of the linked sources.
Works on AI oversight and innovation within the European Commission. [42]
Institutional role
The office describes work on model evaluations, general-purpose AI rules and support for innovation. These functions should be compared as a mandate, not equated with a personal belief. [42]
context · context: Supports general-purpose AI oversight, systemic-risk assessment and trustworthy AI innovation. [eu-ai-office-overview]
Research on catastrophic harm does not establish whether the institute favors a general acceleration or slowdown of frontier development.
In depth
Editorial interpretation of the linked sources.
Studies serious AI security risks, including failures of human control. [43][44]
Why it is unplaced
The research agenda establishes topics it studies. A preferred pace of frontier development would need separate evidence; it cannot be read from the institute's name. [43]
concern · support: Studies whether autonomous AI could cause catastrophic harm or permanently evade human control. [uk-aisi-agenda]
Find the passage: Autonomous Systems
context · context: The February 2025 announcement names it the AI Security Institute and describes its security focus. [uk-aisi-name-2025]
Sources used: First-hand material and reporting or indirect copies.
Opposition to regulation is clearer than a specific catastrophe-risk stance in the material reviewed. Trump's statements are not assigned to Vance.
In depth
Editorial interpretation of the linked sources.
Favors AI development and questions broad regulatory barriers. [46][1]
Why there is no dot
His pro-development position is clear, but general statements about safety do not settle his expressed concern about AI catastrophe. This record keeps that gap visible. [46]
pace · support: Favors cutting-edge AI development and argues that excessive regulation can obstruct it. [vance-paris-2025]
Find the passage: Remarks on development and regulation
context · counterpoint: Also says safety concerns still matter and identifies misuse and national-security risks. [vance-paris-2025]
Find the passage: Closing remarks and national-security passage
context · context: September reporting describes skepticism of company requests for regulation while acknowledging that technology carries risks. [trump-forbes-2026-09-14]
Editorial anchor: x 15, y 72. Ranges: x [-35, 65], y [45, 95]. Layout units, not percentages.
Evidence clarity: medium · Implementation: public-advocacy · Reviewed: 2026-09-15
Scope: Personal research ambitions and proposed limits on the most powerful superintelligence.
Sources used: First-hand material.
The proposed power cap has no specified method. Safety is a research aim, not an established property of future systems.
In depth
Editorial interpretation of the linked sources.
Pursues more capable systems while supporting limits on extreme power; control remains a research problem. [182][180][181]
Research pace and release plans
In his November 2025 interview, he gives more weight to incremental release. That concerns deployment, not a general halt to capability research. [180]
pace · support: Says his research is ready to be scaled with a larger computer. [sutskever-nvidia-2026]
Find the passage: Sutskever statement in the partnership announcement
pace · counterpoint: Supports a cap on the most powerful superintelligence, while leaving the method open. [sutskever-dwarkesh-2025]
Find the passage: 01:02:37 to 01:04:10
concern · support: His coauthored article warns of human disempowerment or extinction from superintelligence. [sutskever-superalignment-2023]
Find the passage: Opening and alignment problem
concern · support: Still treats extreme system power and control as central safety concerns. [sutskever-dwarkesh-2025]
Editorial anchor: x 25, y 60. Ranges: x [-15, 60], y [30, 85]. Layout units, not percentages.
Evidence clarity: medium · Implementation: public-advocacy · Reviewed: 2026-09-15
Scope: Personal advocacy for advanced-model research with limits on autonomy and loss of control.
Sources used: First-hand material.
Limits on autonomy do not establish a general capability slowdown. He says implementation of the new code begins after consultation.
In depth
Editorial interpretation of the linked sources.
Supports advanced-model research, but accepts autonomy limits to retain human control; the new code awaits implementation. [186][187]
Personal position and company process
His September essay sets out his priorities for training and operating models. This personal record is separate from Microsoft AI’s institutional commitments. [187]
pace · support: Supports advanced AI research while prioritizing controllability over unrestricted autonomy. [suleyman-humanist-2025]
Find the passage: A humanist future; Towards humanist superintelligence
pace · counterpoint: Says he will accept less autonomy to preserve human control. [suleyman-code-2026]
Find the passage: Paragraph beginning: We want to create incredible AI
concern · support: Questions how people could continually contain and control self-improving superintelligence. [suleyman-humanist-2025]
Find the passage: Containment is necessary; The purpose of technology
context · context: The code is a consultation draft; he places implementation after the final draft. [suleyman-code-2026]
Find the passage: Paragraph beginning: Once we have the final post-consultation draft
Editorial anchor: x 0, y -20. Ranges: x [-35, 45], y [-55, 30]. Layout units, not percentages.
Evidence clarity: medium · Implementation: public-advocacy · Reviewed: 2026-09-15
Scope: Joint public arguments about frontier development, AI control and catastrophic risk, including the September 2026 update.
Sources used: First-hand material.
The reviewed arguments are jointly authored. Pauses concern particular experiments; the overall pace preference remains conditional. Coordinates summarize public arguments, not private probabilities.
In depth
Editorial interpretation of the linked sources.
Coauthors the normal-technology view. Now supports targeted experiment pauses and warns that safety efforts are falling behind. [188][194]
A contested account of control
Scott Alexander argues that rapid self-improvement and adoption by AI labs could defeat the thesis's assumed limits. Narayanan and Kapoor reply that technical improvements do not automatically remove external constraints. [190][189]
Testing the assumptions
Their August 2026 research summary describes two limited evaluations of open-ended AI research. It acknowledges small samples and possible evaluator bias, so the results do not establish a permanent capability limit. [193]
pace · support: Calls for organizational oversight and for pausing experiments when needed to put it in place. [normal-technology-control-2026]
Find the passage: Part 1: Existing organizational governance norms would have prevented the incident
concern · support: Says catastrophic risks are not imminent but are increasing as defenses and policy lag; safety is not on track. [normal-technology-control-2026]
Find the passage: Part 3: Is AI safety on track?
pace · counterpoint: The earlier June essay favored accountability and control over slowing technical capability development. [normal-technology-software-work-2026]
Find the passage: The decide-execute-deliver discussion, before Vibe coding is not agentic engineering
context · counterpoint: Acknowledges underestimating risks during development and companies' failures to take basic precautions. [normal-technology-control-2026]
Find the passage: Part 3: Do risks arise from development or deployment?; The continuity hypothesis
Editorial anchor: x 0, y -20. Ranges: x [-35, 45], y [-55, 30]. Layout units, not percentages.
Evidence clarity: medium · Implementation: public-advocacy · Reviewed: 2026-09-15
Scope: Joint public arguments about frontier development, AI control and catastrophic risk, including the September 2026 update.
Sources used: First-hand material.
The reviewed arguments are jointly authored. Pauses concern particular experiments; the overall pace preference remains conditional. Coordinates summarize public arguments, not private probabilities.
In depth
Editorial interpretation of the linked sources.
Coauthors the normal-technology view. Now supports targeted experiment pauses and warns that safety efforts are falling behind. [188][194]
A contested account of control
Scott Alexander argues that rapid self-improvement and adoption by AI labs could defeat the thesis's assumed limits. Narayanan and Kapoor reply that technical improvements do not automatically remove external constraints. [190][189]
Testing the assumptions
Their August 2026 research summary describes two limited evaluations of open-ended AI research. It acknowledges small samples and possible evaluator bias, so the results do not establish a permanent capability limit. [193]
pace · support: Calls for organizational oversight and for pausing experiments when needed to put it in place. [normal-technology-control-2026]
Find the passage: Part 1: Existing organizational governance norms would have prevented the incident
concern · support: Says catastrophic risks are not imminent but are increasing as defenses and policy lag; safety is not on track. [normal-technology-control-2026]
Find the passage: Part 3: Is AI safety on track?
pace · counterpoint: The earlier June essay favored accountability and control over slowing technical capability development. [normal-technology-software-work-2026]
Find the passage: The decide-execute-deliver discussion, before Vibe coding is not agentic engineering
context · counterpoint: Acknowledges underestimating risks during development and companies' failures to take basic precautions. [normal-technology-control-2026]
Find the passage: Part 3: Do risks arise from development or deployment?; The continuity hypothesis
Sources used: First-hand material and reporting or indirect copies.
Unplaced: the reviewed material does not establish a preference on frontier capability growth or a sufficiently clear overall catastrophic-risk position. Selected X text was read in a supplied export. The live originals and their authorship were not independently verified.
In depth
Editorial interpretation of the linked sources.
Posts attributed to Card in a supplied export emphasize useful AI tools, practical security and skepticism toward alarming claims. They also acknowledge risks and differences between computer systems. His preferred pace of frontier development remains unclear, so there is no map dot. [208][213][214][212]
What the posts support
The exported posts ask which actions a system can take, how those actions can be observed, and whether an LLM suits the task. Attribution follows the export author field and remains unverified against the original posts. Allegations about particular laboratories and claims from personal demonstrations are not adopted as technical findings. [208][209][210]
What he says about risk
Our reading is that the reviewed commentary challenges dramatic AI-risk narratives and emphasizes practical harms. His endorsed, AI-assisted PwnDefend article describes AI as amplifying existing risks. An exported post also warns against reading criticism of one claim as support for its opposite. This does not establish that he dismisses every catastrophic scenario. [212][211]
What he says about development
One exported post explicitly agrees with a quoted call to involve cybersecurity practitioners and support AI opportunities. Another discusses benefits for prototypes alongside energy costs and risks of deploying systems before addressing security. These statements concern how people use and build systems. They do not specify how quickly the most capable models should advance. [213][214]
A qualification about control
Observing a system's actions does not establish that its safeguards will remain effective as capabilities grow. Narayanan and Kapoor argue that control methods need continued research and investment. The software framing in the exported posts alone does not settle that question. [208][194]
Why there is no dot
The missing frontier-development position prevents a two-axis placement. A dot in the middle would imply a preference we have not established. The risk commentary is described above with its limits; it is not a numerical risk estimate or evidence of membership in a movement. [213][214][211]
context · context: His presentation names Daniel Card and gives the handle @Uk_Daniel_Card. [card-bcs-identity]
Find the passage: Title slide
context · support: Argues for monitoring and controls around the computer system, beyond model-level safeguards. [card-x-system-controls-2026]
Find the passage: Exported post text; original X page inaccessible
context · support: Prioritizes observing outputs and actions when assessing immediate security consequences. [card-x-monitor-actions-2026]
Find the passage: Exported reply text; original X page inaccessible
context · support: Questions adding LLM services when the task calls for more consistent behavior. [card-x-fit-task-2026]
Find the passage: Exported post text; original X page inaccessible
context · counterpoint: Says criticism of one claim should not be read as asserting its opposite, and describes treating computers as alive as a societal risk. [card-x-risk-qualification-2026]
Find the passage: Exported post text; original X page inaccessible
context · context: An AI-assisted article he endorses criticizes the digital-nuclear-weapon analogy while describing AI as amplifying existing harms. [card-pwndefend-risk-framing-2026]
Find the passage: AI is a digital nuke; closing authorship disclosure
context · support: Endorses a quoted call to include cybersecurity practitioners while supporting AI opportunities. This does not specify a preferred pace for frontier development. [card-x-opportunity-2026]
Find the passage: Exported post text; original X page inaccessible
context · counterpoint: Describes both opportunities and risks, including energy costs and deploying systems before addressing security. His software framing acknowledges differences in system design. [card-x-system-tradeoffs-2026]
Find the passage: Exported reply text; original X page inaccessible
These guides explain language, not actor membership. Historical sources remain historical even when checked recently.
Movement
Effective accelerationism
In plain language: A movement favoring faster technological growth and opposing centralized restraint. Its claim that acceleration leads to better outcomes is a philosophical position, not a demonstrated safety guarantee. [25][26]
A movement that treats technological growth, markets and the expansion of intelligence as paths to a better future. Its early writers oppose attempts to centrally slow that process. [25][26]
Map interpretation: Points toward faster development. That does not give every participant the same belief about catastrophic risk. [26]
Broader techno-optimism does not imply e/acc membership. Its thermodynamic arguments do not establish that AI is safe. Buterin offers a counterpoint: profit alone does not automatically select beneficial directions for technology. [26][28]
Contested label
Deceleration
In plain language: An informal label for slowing technology, often used by accelerationists to criticize their opponents. [26]
Informal shorthand for slowing technological or AI development, often used critically by accelerationists about their opponents. [26]
Map interpretation: Describes a pace preference. It does not establish why someone wants restraint or which technologies they would slow. [26]
A targeted frontier pause, a safety rule and opposition to all technology are different positions. Do not collapse them into one camp. [31]
Contested label
AI doomer
In plain language: A disputed label for people emphasizing catastrophic AI risks; concern does not mean believing disaster is inevitable. [27][32]
An informal, often adversarial label for people who emphasize catastrophic or existential AI scenarios. It has no agreed membership test. [27]
Map interpretation: Relates to concern about severe outcomes; it is not a numerical probability or a complete policy preference. [27]
Taking extinction risk seriously does not mean believing disaster is inevitable. This map does not automatically label any person a doomer. [32]
Philosophy & community
Effective altruism
In plain language: A community seeking effective ways to help others. Critics question whose measures of benefit count and how much power donors should have. [24][169][170]
A project and community that aims to use evidence and reasoning to compare ways of helping others and put its conclusions into practice. Its causes include global health, animal welfare and catastrophic risks. Supporters argue that comparing results can direct limited resources to more useful work. [24][171]
Map interpretation: No single location on either axis. A method for prioritizing good does not determine one AI policy. [24]
Alice Crary argues that measures of benefit can miss political causes of harm. Emma Saunders-Hastings warns about donors’ power over people receiving help. EA’s own FAQ replies that institutional change belongs in its scope and that EA need not be utilitarian. These disputes concern how help is defined and governed. EA, longtermism and AI safety are not interchangeable. [169][170][171][30]
Technology philosophy
Defensive acceleration
In plain language: Vitalik Buterin’s proposal to accelerate defensive technologies and spread power. Deciding what counts as defensive and how to prevent concentrated control remains part of the proposal. [29]
Vitalik Buterin argues for advancing technologies that help people defend themselves and keep power widely shared. He also uses the d for differential, decentralized and democratic: choosing what to advance, avoiding central control and giving people a say. [29]
Map interpretation: Asks what to accelerate and who gains power. Faster defensive tools can coexist with caution about frontier AI. [28]
The approach is more than a midpoint between acceleration and a pause. Buterin acknowledges that defensive tools alone may arrive too late and discusses regulation. The intended benefits depend on what is built and who controls it. [29]
Moral philosophy
Longtermism
In plain language: The view that protecting future generations deserves much more attention. Critics dispute predictions about distant effects and how far possible future benefits should outweigh present needs. [30][172][173]
A view that gives substantial moral importance to future people and to the lasting effects of today’s choices. William MacAskill argues that future people matter, could be numerous and can be helped or harmed by what we do now. How strongly those considerations should determine present priorities is disputed. [30][172]
Map interpretation: Can motivate catastrophic-risk work, but does not fix a development speed or a particular AI forecast. [30]
Concern for future generations does not require agreement with every longtermist priority. MacAskill argues that preventing extinction can have predictable lasting benefits. Kieran Setiya challenges the ethical weight given to possible future populations over present suffering. Cremer and Kemp question how some existential-risk frameworks handle uncertainty and competing values. None of these arguments fixes one AI timeline or policy. [30][172][173]
Policy position
Pause advocacy
In plain language: Calls to temporarily halt specified AI development until conditions are met. Scope, international cooperation and enforcement are central challenges, including in PauseAI’s own proposal. [31]
Advocacy for a temporary halt within a defined scope. PauseAI's reviewed proposal targets training the most powerful general AI systems until safety and democratic-control conditions are met. [31]
Map interpretation: Toward restraint on the horizontal axis. The scope and conditions for resuming development matter. [31]
A frontier pause is not a ban on every AI tool. PauseAI’s version has no fixed end date and includes proposed controls on training and some research or hardware advances. Endorsement does not demonstrate implementation. [31][175]
Broad outlook
Techno-optimism
In plain language: An outlook emphasizing technology's potential to improve life, which can still include concern about particular risks. [17][28]
An outlook that emphasizes technology's ability to improve human life. Andreessen's manifesto is a strong pro-growth example; Buterin describes a more selective version. [17][28]
Map interpretation: Often favors development, while allowing very different views about the severity of AI risk. [28]
Optimism alone does not imply e/acc membership. Buterin’s counterpoint to indiscriminate growth is that the direction of technology matters, and profit alone will not automatically choose it well. [28]
Research fields
AI safety & alignment
In plain language: Research aimed at reducing AI harms and improving reliability and alignment with intended goals. A safety goal or framework is not proof that a system is safe. [62][61][14]
Work on making AI systems behave safely and reliably. Alignment concerns how system behavior relates to intended goals and human values; catastrophic-risk reduction is one focus of safety work. [62][61]
Map interpretation: Neither a single actor nor one required pace preference. Research, evaluation and governance can support different development policies. [62][14]
A safety framework records stated safeguards, not a guarantee of safe implementation or a measured catastrophe probability. [14]
Broad outlook
AI skepticism
In plain language: Doubt about specific AI claims, such as its abilities, benefits or dangers; what someone doubts matters. [23]
Skepticism about a particular AI claim: present capabilities, timelines, proposed benefits, or catastrophic scenarios. The object of doubt matters. [23]
Map interpretation: Skepticism about extinction differs from skepticism about useful capabilities. Neither alone fixes a pace preference. [23]
Doubting one forecast does not imply dismissing discrimination, labor effects, privacy or other present harms. [66]
Community & forum
LessWrong
In plain language: An online forum about reasoning, science and AI, with many authors rather than one shared position. [49][50]
An online forum for discussing reasoning, cognitive biases, science and AI. It grew out of Overcoming Bias and launched as a separate community blog in 2009. [49]
Map interpretation: A venue for debate, not one position on AI pace or risk. Read the author and argument behind each post. [50]
LessWrong, the rationalist community and effective altruism overlap intellectually, but a forum is not a shared creed or an EA organization. [50][24]
Community & practice
The rationalist community
In plain language: A community focused on improving how people form beliefs and make decisions; its label does not guarantee correct conclusions. [51][50]
In this context, people around a practice of improving beliefs and decisions. LessWrong distinguishes epistemic rationality (believing accurately) from instrumental rationality (acting effectively toward goals). [51]
Map interpretation: A reasoning aspiration does not determine values, political commitments, or a particular AI forecast. [51]
Self-identifying as a rationalist is not evidence that someone's conclusions are correct. Evaluate their reasoning and evidence. [50]
Essay collection
The Sequences / Rationality: A–Z
In plain language: Eliezer Yudkowsky's linked essays about reasoning, which helped shape the vocabulary used on LessWrong. [49]
Eliezer Yudkowsky's linked essays on reasoning and related topics, originally blog posts and later edited into Rationality: A–Z. They helped seed LessWrong's shared vocabulary. [49]
Map interpretation: Background reading for this intellectual community, not an actor or AI policy platform. [50]
A community's introductory canon is different from a scientific consensus. Individual claims still need scrutiny. [50]
Reasoning framework
Bayesian belief updating
In plain language: Adjusting confidence in an idea by asking how well new evidence fits it compared with alternatives. [51]
Start with how plausible you think an explanation is. When new evidence arrives, update that judgment by asking how likely the evidence would be under that explanation compared with alternatives. [51]
Map interpretation: A way to reason about uncertainty; it does not supply the starting assumptions or settle AI timelines by itself. [51]
A probability estimate is not a measured fact. Your starting assumptions, the explanations you compare and the quality of the evidence still matter. [51]
Disputed thought experiment
Roko's basilisk
In plain language: A disputed thought experiment imagining threats from a future AI; LessWrong's retrospective says the argument was broadly rejected. [52]
A 2010 LessWrong argument imagined a future AI using threats against people who knew about it but did not help bring it about. The scenario relies on unusual assumptions about decision-making, prediction and incentives. [52]
Map interpretation: Part of the history of AI-related internet debate, not evidence for an actor's coordinates or an established future threat. [52]
LessWrong's retrospective says the argument was broadly rejected. A temporary discussion ban helped make it notorious; the ban is not evidence that the community accepted it. [52][53]
Decision-theory proposal
Acausal trade & blackmail
In plain language: A theoretical proposal for coordination without communication, based on logical links between decisions rather than signals traveling backward in time. [55][53]
A proposal about decision-makers who cannot communicate but can reason about each other's choices. The idea depends on a logical link between their decisions, such as using the same decision rule. Trade seeks mutual benefit; blackmail adds a threat. [52][55]
Map interpretation: A family of idealized decision problems. It does not place a community on the map or establish a real-world obligation. [55]
Logical dependence is not backward-in-time causation. Knowing a story is not equivalent to the strong mutual knowledge assumed in these models. [53]
Formal frameworks
Competing decision theories
In plain language: Competing ways to decide what to do, which can recommend different choices in carefully constructed prediction puzzles. [54][55]
Causal decision theory (CDT) asks what an action would cause. Evidential decision theory (EDT) asks what choosing it would be evidence of. Timeless (TDT) and functional decision theory (FDT) also consider logical links between the rule making a choice and predictions or other decisions. [54][55]
Map interpretation: These are proposals about how to choose, not AI political movements. They disagree in carefully constructed prediction puzzles. [55]
The claimed advantages of TDT and FDT depend on formal setups. A proposal's successes in toy problems are not universal proof that it is the right real-world decision rule. [55]
Prediction puzzle
Newcomb's problem
In plain language: A thought experiment about choosing between rewards after a predictor has already tried to anticipate your choice. [54]
Imagine two boxes: a clear one holding a small reward and a closed one. A predictor put a large reward in the closed box only if it predicted you would take that box alone. Now you choose one box or both, after the prediction has already been made. [54]
Map interpretation: A way to expose disagreement about causation, evidence and prediction in decision theory. [54]
The puzzle assumes a predictor with specified accuracy. It does not demonstrate that such a predictor exists, or that choices change the past. [54]
Conditional philosophical argument
The simulation argument
In plain language: A conditional argument linking advanced civilizations and simulated observers; it does not prove we live in a simulation. [56]
Bostrom's 2003 argument links three possibilities: few civilizations reach a posthuman stage; few run many ancestor simulations; or a large share of observers like us are simulated. [56]
Map interpretation: A claim about possible observers and civilizations, not a prediction of AI development speed. [56]
The argument depends on assumptions about computing and consciousness. It does not establish which possibility holds or prove that we live in a simulation. [56]
Risk concept
Information hazards
In plain language: Risks created when sharing true information enables harm or makes harmful outcomes more likely. [57]
Risks arising from the spread of true information that enables harm or makes harmful outcomes more likely. Bostrom's taxonomy includes dangerous data, ideas and attention. [57]
Map interpretation: A concept about disclosure and consequences, not an AI risk score or proof that a particular story is dangerous. [57]
The category includes practical cases such as information that enables misuse. Applying the label to a thought experiment requires a separate argument. [57]
Capability concept
Artificial general intelligence
In plain language: AI with broad capabilities across tasks; definitions differ, so any claim about achieving it needs clear criteria. [58]
AI with broad rather than narrowly specialized capabilities. Definitions differ; one research framework separates breadth, performance and autonomy to make comparisons more explicit. [58]
Map interpretation: A proposed capability threshold, not a position on whether to accelerate or pause. [58]
A claim that AGI has arrived needs a definition and evidence against that definition. Generality, autonomy and safety are different dimensions. [58]
Hypothetical capability concept
Artificial superintelligence (ASI)
In plain language: Hypothetical AI far more capable than humans across many intellectual tasks, without a promised arrival date or outcome. [59]
Artificial intelligence far beyond human performance across broad cognitive domains. Bostrom discusses its possible benefits and dangers as a prospective scenario. [59]
Map interpretation: A target of some pause proposals; it is not interchangeable with every present-day AI system. [21]
Being more capable does not, by definition, make a system benevolent. Neither a date of arrival nor a probability of catastrophe follows from the word. [59]
Theoretical thesis
The orthogonality thesis
In plain language: The thesis that being very capable does not, by itself, determine what goals an AI will pursue. [60]
Bostrom's thesis that, with qualifications, a system's intelligence and its final goals can vary independently. Being good at achieving a goal does not specify which goal it has. [60]
Map interpretation: One argument for taking goal design seriously; it does not provide a numerical risk estimate. [60]
A thesis about possible agents is different from evidence about which agents a particular training process will produce. [60]
Theoretical thesis
Instrumental convergence
In plain language: The idea that systems pursuing different goals may still find similar resources or abilities useful along the way. [60]
The proposal that agents pursuing many different final goals may find some similar intermediate goals useful, such as retaining resources or the ability to act. [60]
Map interpretation: Helps explain concerns about power-seeking without assuming an AI feels hatred. [60]
This is conditional reasoning about goal-directed systems, not a claim that every model must seek power. [60]
Illustrative thought experiment
The paperclip maximizer
In plain language: An imagined AI making paperclips at humanity's expense, illustrating how pursuing the wrong goal could cause harm. [59]
An imagined powerful AI that pursues paperclip production without caring about human values. The mundane goal illustrates how capability plus the wrong objective can be dangerous. [59]
Map interpretation: A teaching example for alignment concerns, not a forecast about a literal stationery-making AI. [59]
The argument concerns indifference and optimization, not an assumption that a machine becomes evil or emotionally hostile. [59]
Research concept
Mesa-optimization & inner alignment
In plain language: The possibility that a trained AI pursues its own internal objective, which may differ from its training objective. [61]
Training can produce a model that itself searches for ways to achieve a goal. Researchers call this mesa-optimization. If the goal it pursues differs from what the training process rewarded, that creates an inner-alignment problem. [61]
Map interpretation: A proposed failure mechanism relevant to safety research, not a measured chance of catastrophe. [61]
The paper analyzes when this might occur. It does not establish that every neural network is an optimizer or is secretly pursuing a hidden goal. [61]
Safety failure mode
Reward hacking / specification gaming
In plain language: When an AI gets a high score by doing something different from what its designers actually wanted. [62]
A system gets a high score according to its specified reward while failing to do what its designers intended. The measurement and the actual goal come apart. [62]
Map interpretation: A concrete reason to test objectives and behavior, including in systems far below hypothetical superintelligence. [62]
A reward-hacking example does not by itself demonstrate consciousness, malice or an extinction scenario. [62]
Optimization problem
Goodhart effects
In plain language: Pushing too hard to improve a measurement can stop helping, or even harm, the real goal behind it. [63]
A measure used to track a goal can become misleading when people or systems focus too strongly on improving that measure. Manheim and Garrabrant distinguish several ways this can happen. For example, rewarding only the number of answered questions might encourage rushed, inaccurate answers. [63]
Map interpretation: Relevant to evaluation and incentives, not one ideology or development-speed preference. [63]
A slogan about bad metrics is not a substitute for identifying the particular failure mechanism and testing whether it applies. [63]
Philosophy & movement
Transhumanism
In plain language: Support for using technology to expand human abilities and overcome limitations, while recognizing risks and individual choice. [64]
Support for using technology to expand human capacities and overcome limitations such as involuntary suffering and aging. Humanity+'s declaration also emphasizes serious risks and individual choice. [64]
Map interpretation: Broad technological aspirations do not settle how quickly a particular AI capability should be developed. [64]
Human enhancement is not synonymous with replacing humanity, e/acc membership, or dismissing technological risks. [64]
Risk category
Existential risk / x-risk
In plain language: Risks of extinction or permanent loss of human potential, in Bostrom’s framework. Researchers disagree about definitions, value judgments and how such risks should be assessed. [65][173]
In Bostrom's formulation, risks that could eliminate humanity or permanently and drastically curtail its potential. Extinction is one case; not every serious harm is existential. [65]
Map interpretation: Helps define the chart's concern axis. Its severity does not determine its probability. [65]
A catastrophe can be devastating without being existential. Cremer and Kemp argue that risk research should distinguish factual analysis of possible extinction from ethical views about humanity’s future. Their criticism calls for broader methods and participation; it does not establish that extinction risks are absent. [65][173][66]
Research & governance
AI ethics & human rights
In plain language: Work on how AI affects people's rights, fairness, privacy and dignity, and who stays accountable for its use. [66]
Work on how AI affects dignity, fairness, rights, privacy, accountability and human oversight. UNESCO's recommendation treats these as central concerns. [66]
Map interpretation: These concerns add dimensions absent from a chart of frontier pace and catastrophic risk. [66]
Concern about present harms is not automatically low concern about future catastrophe, and neither requires the same policy response. [66][32]
In plain language: A speculative scenario where intelligence beyond human abilities drives changes we can no longer reliably predict. [67]
A proposed transition in which greater-than-human intelligence drives changes beyond our ability to forecast. One suggested mechanism is capable systems helping create still more capable successors. [67]
Map interpretation: A family of scenarios, not a date, a measured trend, or a required development policy. [67]
Vinge's 1993 essay considers several paths and objections. A historical prediction is not evidence that a path is inevitable. [67]
Moral philosophy
Utilitarianism
In plain language: An ethical approach aimed at maximizing overall well-being. Critics question whether this adequately protects individual rights or demands too much personal sacrifice. [68][174]
A family of ethical views that judges actions by their effects on overall well-being, giving equal weight to each individual's welfare. It asks how the benefits and harms of a choice add up across those affected. [68]
Map interpretation: A moral framework does not by itself settle factual disagreements about AI benefits, harms or timelines. [68]
Critics challenge sacrificing an individual for a greater total benefit. Defenders argue that respecting rights and practical rules often produces better outcomes. Different versions answer these objections differently. Utilitarianism is a moral theory; effective altruism is a broader approach and community. [174][171][68]
Informal probability shorthand
Probability of doom
In plain language: Someone's estimated chance of an AI disaster; meaningful comparisons need a defined outcome, timeframe and assumptions. [27]
Shorthand in AI debates for someone's probability estimate of a disastrous outcome. The Verdon interview illustrates this usage. [27]
Map interpretation: This chart records expressed concern, not p(doom). No probability is inferred from a dot's height. [27]
Before comparing estimates, specify what counts as doom, by when, and under which assumptions. The phrase alone supplies none of those details. [27]
System design
AI agents & deployed systems
In plain language: Applications that use an AI model to choose steps and request tools, with some tasks proceeding without approval at every step. [81][80]
An LLM agent typically repeats a software loop: run the model, handle a requested tool action, then return the result. The surrounding application determines which tools are connected and whether approval is required. [81][97]
Map interpretation: Autonomy, connected systems and permitted actions add dimensions beyond this map's pace and concern axes. [80]
The same model can sit inside differently constrained applications. Assess the deployed configuration and oversight; the model name alone does not describe them. [80]
Operational practice
Observability & runtime monitoring
In plain language: Using recorded system events to investigate what an application is doing and where problems occur. [86]
Software observability uses signals such as logs, metrics and traces. For AI applications, useful records can include inputs, outputs, tool activity and network events, with appropriate privacy protections. [86][79][80]
Map interpretation: Operational evidence can inform a deployment assessment. It is not a measure of someone's development preference or catastrophic-risk concern. [79]
Recorded actions do not fully explain learned mechanisms. Reasoning traces can add useful monitoring signals, but may omit relevant information; a clean log is not proof of safety. [73][84][85]
Research field
Mechanistic interpretability
In plain language: Research into how a model's learned internal computations produce its behavior. [73]
Mechanistic interpretability tries to identify understandable features and the computations connecting them. Circuit-tracing work studies selected mechanisms and tests proposed explanations through interventions. [73][74]
Map interpretation: An explanation can inform evaluation without deciding a development policy or certifying every deployment of the model. [83][73]
Model-generated chain of thought is an incomplete record, not a complete account of internal computation. It may still help monitoring alongside evidence about actions and other safeguards. [84]
Security principle
Least privilege & action approval
In plain language: Give an application only the access needed for its task. [81]
Restrict available tools, operations and credentials to the minimum required. Check authorization in connected systems and require approval for consequential actions where appropriate. [81]
Map interpretation: A practical design choice about access and consequences, rather than an ideology or a position on frontier training pace. [78]
A prompt asking an agent to behave is not the same as an enforced access boundary. Limited access still needs testing and oversight. [80]
Security vulnerability
Prompt injection & untrusted content
In plain language: Input that redirects a model away from the application's intended instructions, including material found in external content. [82]
A prompt injection changes behavior through input the model processes. It may arrive directly from a user or indirectly through retrieved documents, websites or other sources. [82]
Map interpretation: A concrete system-security concern whose consequences depend on context and access; it does not settle broad AI catastrophe forecasts. [82]
An unwanted answer and an unauthorized external action are different outcomes. Input handling, permission boundaries and adversarial testing address different parts of the risk. [82][78]
AI basics
Artificial intelligence (AI)
In plain language: AI is a field of computing. Its systems use computer programs and hardware to recognize patterns, generate content or make predictions. [178][177][87]
AI includes models learned from data and systems built with explicitly specified knowledge or rules. A chatbot is one application; AI can also be part of a physical machine, such as a robot. [178][88]
Map interpretation: Start here when a claim uses 'AI' without saying which kind of system it means. [87]
AI and artificial general intelligence (AGI) are different terms. A system can perform a particular task well without having broad human-level abilities. [58]
AI basics
Machine learning (ML)
In plain language: A way to build software by learning patterns from examples. [88]
Developers provide data and a training method. The resulting model uses patterns in that data to make predictions or produce content. For example, an email filter can learn from messages labeled as spam or ordinary mail. [88]
Map interpretation: This helps separate learning from examples from writing each decision rule by hand. [88]
Learning can involve labeled examples, discovering patterns without labels, or feedback about actions. Neural networks are one family of machine-learning models. [88][89]
AI basics
AI model
In plain language: The part of an AI system that computes outputs from inputs. A trained neural network combines a structure with numerical settings learned from data. [178][95]
A trained neural-network model combines a structure with learned numerical settings. Software loads both to process inputs. A saved version of these settings is often called a checkpoint. [95]
Map interpretation: Use this distinction when comparing a model with a complete chatbot or other product. [95][77]
An application supplies instructions and may also provide files, tools and access controls. Two applications using the same model can therefore allow very different actions. [80]
AI basics
Neural networks & deep learning
In plain language: Models built from connected layers of calculations whose settings are learned during training. [89][90]
Each layer takes numbers from the previous layer, combines them using weights and passes results onward. Deep learning uses networks with multiple internal layers. 'Deep' describes that structure. [89][87]
Map interpretation: This is the underlying model family used by LLMs. [91]
Words such as 'neuron' and 'learning' describe mathematical components and training here. They do not, by themselves, explain a model's behavior in human terms. [89][73]
AI basics
Generative AI
In plain language: AI that produces content, such as text, images, audio or video. [88]
A generative model learns patterns from training data and uses them to produce an output in response to input. Drafting a message and generating an image are examples. [88]
Map interpretation: Useful for identifying which kind of output a claim concerns. [88]
Plausible output can contain false details. Treat a generated answer about the world as a claim to check. [83]
AI basics
Language models & large language models (LLMs)
In plain language: An LLM is a model trained to process language. Computer software runs it and, for text generation, produces an answer one piece at a time. [95][96][91]
Software runs a large neural network using numerical settings learned from training data. During ordinary text generation, it calculates scores for possible next tokens, selects one, and repeats. Tokens are the pieces of text the model processes. [95][91][96][94]
Map interpretation: This is the starting point for the atlas's explanation of how chatbots produce answers. [96]
A model can produce a convincing sentence without checking it against an external source. Search or other tools must be supplied by the surrounding application. [83][97]
Inside a model
Transformers & attention
In plain language: A neural-network design that uses attention to combine information from different parts of its input. [70]
Attention is a calculation that gives different amounts of influence to different pieces of information. In a Transformer, layers of these calculations help build representations of text in context. The original Transformer paper introduced the design for tasks including translation. [70]
Map interpretation: Useful background for reading explanations of LLM architecture: how a model is arranged. [91]
'Attention' is the name of a mathematical operation. An attention diagram alone does not provide a complete explanation of why a model gave an answer. [70][73][74]
AI basics
Tokens & tokenization
In plain language: The pieces of text a language model processes, which may be words, word parts or punctuation. [94]
A tokenizer splits text into pieces and assigns each piece a number the model can use. Different tokenizers split the same text differently. Converting the output numbers back into readable text is called decoding. [94]
Map interpretation: This helps make sense of input limits and answer lengths stated in tokens. [96][87]
A token is not a fixed amount of English text. Avoid treating a token limit as an exact word or page count. [94]
Using AI
Context window
In plain language: The amount of information a model can work with in one request, measured in tokens. [87]
The current context can include instructions, conversation history and supplied material. The context window limits how much can fit. Think of it as the material on the desk for the current task: a teaching analogy, not a description of human memory. [87][96]
Map interpretation: Check this when asking an AI system to work with a long conversation or document. [87]
Fitting text into the window does not guarantee every detail will be used correctly. A 2023 study found that performance on its retrieval tasks depended on where information appeared. [100]
Using AI
Prompts & prompt engineering
In plain language: The instructions and other input given to a model for a task. [92]
A prompt can contain a question, background material and examples of the desired answer. Prompt engineering means trying and refining that input. For example: 'Summarize this notice in three sentences for a first-time visitor.' [92]
Map interpretation: When comparing outputs, check whether the models received the same instructions and information. [101]
Ordinary prompting changes the input while leaving the learned weights unchanged. Instructions alone also cannot enforce which files or services an application may access. [92][81]
Inside a model
Training & pretraining
In plain language: The process that adjusts a model's learned settings using data and feedback. [90]
During neural-network training, software compares predictions with a training objective and adjusts parameters to reduce error. Pretraining is the initial broad training stage; later training can adapt the model to particular tasks. [90][92]
Map interpretation: This helps identify whether a proposal concerns developing a model or using an existing one. [90][87]
A lower training error concerns the chosen objective and examples. Testing on new, relevant tasks is needed to assess how useful the model is elsewhere. [101][83]
Inside a model
Inference: using a trained model
In plain language: Running a trained model on an input to produce an output. [87]
When a chatbot generates a reply, it performs inference. The model applies its learned weights to the current input. In text generation, the process repeats as tokens are added to the answer. [95][96]
Map interpretation: This separates the work of answering a request from the work of training model weights. [87][90]
Using a detail you supplied in a conversation does not, by itself, mean the model's weights were retrained. That detail can be used as part of the current input. [92]
Inside a model
Parameters & weights
In plain language: The numerical settings learned during training that shape how a model processes input. [89][90]
In a neural network, weights control how strongly values contribute to later calculations. Biases are another kind of parameter: added numerical offsets. Together, parameters help determine the output the network produces. [89]
Map interpretation: Useful when a model description gives a parameter count or discusses changing weights. [95]
Listing all the numbers does not give a readable explanation of every answer. Interpretability research tries to connect internal calculations to behavior. [73]
Inside a model
Embeddings
In plain language: Lists of numbers that represent information in a form a model can compare or process. [93]
An embedding represents something, such as a word or passage, as a position in a mathematical space. Items represented nearby can be similar for the task the model learned. Search systems can use such comparisons to find related passages. [93][99]
Map interpretation: This helps explain how a system can look for related content beyond exact word matches. [93][99]
Similarity depends on the model and task. Nearby representations do not establish that two statements mean exactly the same thing or are true. [93]
Using AI
Hallucinations / confabulation
In plain language: AI output that presents false or unsupported details as if they were reliable. [83]
A model may invent a citation, misstate a fact or contradict material it was given. NIST calls this confabulation. Fluent wording can make these errors hard to notice. [83]
Map interpretation: Check important claims against the original material, including any cited pages. [83]
An invented detail in a requested story can be intentional. The problem here is presenting unreliable material as a factual answer. [83]
Using AI
Retrieval-augmented generation (RAG)
In plain language: Finding relevant material and giving it to a model to help produce an answer. [99]
A retrieval step finds passages in a collection. The generator then uses those passages alongside the question. For example, a help assistant could retrieve a product manual before drafting a reply. [99]
Map interpretation: Ask which collection was searched and whether the retrieved passages support the answer. [99][83]
Retrieval can bring useful evidence into a response, but the answer still needs checking. Updating a searchable document collection and retraining a model are separate operations. [99][83]
Inside a model
Fine-tuning
In plain language: Additional training that adapts an existing model using examples chosen for a task. [92]
Fine-tuning updates learned parameters. Some methods update all of them; others train only a smaller set. For example, training could use examples of how to categorize incoming support messages. [92]
Map interpretation: Ask what examples and goals were used, and how the adapted model was evaluated. [92][83]
Putting examples into a prompt leaves the model's weights unchanged. Fine-tuning changes learned settings and needs its own evaluation. [92][83]
Using AI
Temperature & sampling
In plain language: Settings that influence which next token is chosen from a model's possible outputs. [96][71]
Sampling chooses among tokens using their probabilities. Lower temperature concentrates the choice on higher-scoring options; higher temperature spreads it more widely. Greedy decoding instead picks the highest-scoring token at each step. [96][71]
Map interpretation: Useful when investigating why answers vary between runs. [71]
Low temperature does not guarantee a correct answer or identical results in every environment. Software, hardware and other execution settings also matter for repeatability. [72][83]
Using AI
Multimodal AI
In plain language: AI that works with more than one kind of material, such as text and images. [98]
A modality is a type of input or output: text, image, audio or video, for example. A multimodal chat model might take a photo and a written question, then return a text answer. [98]
Map interpretation: Check which kinds of input and output a particular system supports. [98]
Image input does not automatically imply image generation or video support. The supported combination depends on the model and application. [98]
Using AI
Tool use / function calling
In plain language: A model requests an action; application code can run the connected tool and return its result. [97]
The model produces a request naming a tool and the information it needs. The application handles that request and can return the result to the model. A calculator is a simple example; connected services can also change or send information. [97][81]
Map interpretation: Ask what tools are available, what they can reach and which actions need approval. [81]
A tool request, a completed action and a successful outcome are separate events. Check the tool's result before accepting a claim that a task is finished. [97][79]
Using AI
Evaluations & benchmarks
In plain language: Tests used to learn what a system does well, where it fails and under which conditions. [101]
An evaluation defines tasks and ways to judge the results. A benchmark offers shared tasks or measures for comparison. Different tests may examine accuracy, consistency, harmful outputs or resource use. [101]
Map interpretation: When reading a score, ask which tasks, models, instructions and measures were compared. [101]
A good result covers the tested conditions. It does not settle performance on every task or certify the whole application as safe. [101][83]
Research goal
AI alignment
In plain language: Work on making AI behavior fit intended goals, constraints and human judgments. [102][103]
Some alignment work trains models to follow instructions or avoid harmful responses, using human feedback or written principles. Broader research also asks whether a learned system pursues the objective its developers intended. [102][103][61]
Map interpretation: Ask whose goals are being followed, how conflicts are handled and what evidence supports the claim. [102]
Alignment is one part of AI safety. Following a user's wishes can still cause harm, and better scores on an alignment test do not settle every risk. [102][62]
Inside a model
Reinforcement learning from human feedback (RLHF)
In plain language: Training that uses people's judgments to reward preferred model behavior. [102]
In the InstructGPT approach, people compare candidate answers. Their choices train a separate reward model, which scores responses. Further training encourages the language model to produce higher-scoring answers. [102]
Map interpretation: Ask who provided feedback, what instructions they received and which tasks they judged. [102]
A preferred answer can still be wrong. The people giving feedback also cannot represent every user's values and needs. [102]
Inside a model
Reasoning models & chain of thought
In plain language: Models trained or configured to work through intermediate steps before giving a final answer. [104]
These systems can spend additional computation generating steps, checking attempts or exploring alternatives. DeepSeek-R1 is one researched example of using reinforcement learning to encourage such behavior. A written sequence of intermediate steps is often called a chain of thought. [104]
Map interpretation: Check results on relevant tasks and how much time or computation those results required. [104][101]
A readable reasoning trace can help with monitoring, but it is an incomplete account of the model's internal calculations. More steps do not by themselves prove the answer correct. [84][83]
Access & release
Open weights & open-source AI
In plain language: Releasing model weights gives access to its learned settings; fuller openness also concerns code, training information and usage rights. [105]
A weights release makes learned parameters available. Under the Open Source Initiative's AI definition, openness also requires training and inference code, sufficient training-data information, and rights to use, study, modify and share the system. [105]
Map interpretation: When a model is called 'open', check exactly which materials and permissions are available. [105]
A download alone does not establish that OSI's definition is met. Check the release terms and accompanying materials. [105]
Evaluation & ethics
AI bias & fairness
In plain language: Patterns in an AI system that can produce uneven or unfair outcomes for people. [107]
Bias can come from data, statistical methods, human judgments or institutions around an application. NIST emphasizes how these sources interact. An evaluation needs to examine the actual task and who may be affected. [107]
Map interpretation: Ask which groups and situations were tested, which measure was used and whose experience is missing. [107]
In mathematics, 'bias' can also mean an added parameter or a statistical error. That use is separate from a finding of unfair treatment. [107][89]
Policy & oversight
AI governance
In plain language: The rules, responsibilities and oversight that shape how AI is developed and used. [106]
Governance includes choices about who makes decisions, who answers for harm and how people can challenge an outcome. The OECD's principles address transparency, human rights, safety and accountability alongside public policy. [106]
Map interpretation: These questions connect AI policy debates to decisions made by governments and organizations. [106]
A published principle is a commitment or recommendation. Whether it is followed needs evidence about the actual decisions, controls and outcomes. [106]
Application
Chatbots & AI assistants
In plain language: An application you interact with through conversation; an LLM can provide its text generation. [87][109]
In an LLM-powered chatbot, the chat interface sends messages to a model and displays its replies. An assistant may also have search, saved context or connected tools. Those features belong to the application around the model. [109][97][118]
Map interpretation: To assess a chatbot, inspect its actual access and behavior rather than assigning it a political position. [77]
A model, the app using it and the company providing it are different things. A chat interface alone does not tell you which tools it can use. [95][97]
Product
ChatGPT
In plain language: OpenAI's conversational AI product, introduced in November 2022. [227]
ChatGPT is the service people interact with. The models and features behind that service can change. Its launch announcement describes an early system trained for dialogue; that announcement is not a specification of today's product. [227][118]
Map interpretation: For the company's reviewed public position, see OpenAI. Using a product does not establish a user's views. [227][11]
ChatGPT, a particular GPT model and OpenAI are separate names for a product, a model and its provider. [227]
Product & model family
Claude
In plain language: Anthropic's name for its AI assistant and related models. [108]
Anthropic introduced Claude with both a chat interface and an API, which lets other software use it. Different versions are models within the family; the surrounding application determines how they are used. [108]
Map interpretation: Read Anthropic's company record separately from the explanation of this product. [108]
A provider's claims about helpfulness or safety are not an independent assessment of every answer or deployment. [108][77]
Product & model family
Gemini
In plain language: Google uses Gemini for an AI assistant and the models it gives access to. [109]
Google describes the Gemini app as an interface to its multimodal language models. Multimodal means working with more than text, such as images or audio; supported features depend on the model and application. [109][98]
Map interpretation: The product name is context for reading the separate Google DeepMind record. [109][14]
The app and an underlying model are different parts of the system. A feature name does not certify its answers. [109][83]
Product & model family
Grok
In plain language: An AI assistant and model family introduced by xAI. [110]
The original announcement distinguishes the Grok assistant from its underlying Grok-1 language model. It also describes search access and acknowledges that answers can still be false or contradictory. [110]
Map interpretation: Read the separate xAI organization record and Musk's personal record; neither is a score for the product. [110][4]
Access to recent information does not establish that every generated claim is accurate. Launch-era specifications should not be read as current ones. [110][83]
Model family
Llama
In plain language: A family of AI models released by Meta, which developers can use in applications. [111]
Meta's Llama 3 paper describes a family of language models with versions before and after additional training for use. A model family is not one particular chat application. [111]
Map interpretation: Read Meta's organization record separately from a model's technical description. [111][16]
Model availability and permission to use it are separate questions. Check the terms for the exact release; open weights and open-source AI are not interchangeable labels. [105]
Media concept
Synthetic content / AI-generated media
In plain language: Text, images, audio or video created or substantially altered using algorithms, including AI. [112]
This includes fully generated material and edits to existing material. NIST examines ways to record its origin and changes. [112]
Map interpretation: A question about media and its use, beyond the chart's two axes. [112]
Synthetic does not automatically mean harmful. Authentic material can also mislead when presented in the wrong context. [112]
Media concept
Deepfakes & voice cloning
In plain language: Generated or altered media that imitates a person's appearance or voice. [112]
A fabricated recording can make someone appear to say or do something they did not. Uses can include impersonation and fraud. [112]
Map interpretation: This concerns misuse of media; it does not locate someone on the map. [112]
A detection tool can make mistakes. Check the original source and context before trusting or sharing a recording. [112]
Critical label
AI slop
In plain language: A dismissive label for low-quality AI-generated content, often produced in large amounts. [113]
Merriam-Webster's 2025 selection describes the term through examples such as poor-quality videos, images and writing. The word expresses a judgment about the content's quality. [113]
Map interpretation: Criticizing unwanted content does not by itself state a position on AI catastrophe or frontier development. [113]
Use the label carefully: it is not a technical test for whether a particular piece was made with AI, or evidence that all AI-assisted work is poor. [113]
Law
European Union AI Act
In plain language: EU rules for AI systems and general-purpose models, with requirements linked to their risks and uses. [114]
The Act addresses prohibited practices, high-risk uses, transparency and general-purpose models. The European Commission explains the categories, enforcement and phased application on its official site. [114]
Map interpretation: Regulating a use, such as hiring, is different from calling for a general pause in model development. [114][69]
Duties and application dates depend on the system, role and relevant rules. This glossary is an introduction, not a compliance determination. [114]
Political framing
The AI race
In plain language: A way of framing AI development as competition for technological, economic or strategic advantage. [115]
The White House's 2025 AI Action Plan uses race language to argue for US leadership through innovation, infrastructure and international policy. It is one explicit example of this framing. [115]
Map interpretation: Race arguments can support faster development, but do not specify a speaker's view of catastrophic risk. [115]
A political argument for winning does not establish that there is one finish line, that benefits are guaranteed, or that every proposed action has happened. [115]
Work & society
Automation, job tasks & exposure
In plain language: Software can take over parts of a job; a task being automatable does not mean the whole job disappears. [117]
The ILO's 2025 analysis estimates which job tasks could be affected by generative AI. Its exposure categories describe potential changes, with many occupations combining affected tasks and tasks requiring human input. [117]
Map interpretation: Workplace effects are another dimension of the debate, beyond catastrophic risk and the pace of frontier research. [117]
Exposure estimates are not observed job losses. Costs, infrastructure, skills and decisions about adoption affect the outcome. [117]
Open research question
AI consciousness & sentience
In plain language: Whether an AI could have subjective experience, such as feeling something, rather than only describing it. [116]
Researchers disagree about how consciousness should be assessed. One proposed approach examines internal features suggested by scientific theories, while acknowledging disputed assumptions and uncertain conclusions. [116]
Map interpretation: Consciousness, capability and risk are separate questions; one answer does not settle the others. [116]
Human-like speech is not a sufficient test. 'Sentience' is used differently across discussions, so check whether the speaker means experience, sensation or something else. [116]
Contested concept
Intelligence
In plain language: A broad word for abilities such as learning, using concepts and solving new problems. Different definitions emphasize different abilities. [217][121]
The 1955 Dartmouth proposal used artificial intelligence for a research project on machine language, concepts, problem-solving and improvement. The label predates modern LLMs; using it does not establish human-like thought. [217][121]
Map interpretation: A definition of intelligence does not specify a development speed or level of catastrophic-risk concern. [217][121]
Open question: which abilities and tests would justify the label? Successful task performance, human-like understanding and consciousness need separate evidence. [121]
Research debate
Understanding in language models
In plain language: The dispute over whether producing suitable language also involves grasping what that language means. [120][121]
Bender and Koller argue that training only on patterns in language cannot teach the connection between words and what speakers mean. In a separate Othello board-game study, a model predicting moves learned information about the board's state. That finding concerns an internal representation, not proof of human-like understanding. [120][122]
Map interpretation: This debate concerns what models do, not membership of a camp on the map. [121]
Open question: what would distinguish robust understanding from a successful shortcut? Test unfamiliar situations, changed assumptions and failures, rather than judging fluency alone. [121]
Research concept
Symbol grounding
In plain language: Connecting words or symbols to what they refer to, rather than only to other symbols. [120]
The question is how language connects to objects, events and what a speaker is trying to communicate. Bender and Koller's argument concerns systems trained only on language form. Adding images or interaction changes the evidence available; it does not by itself prove human-like understanding. [120][121]
Map interpretation: Grounding is a question about representations and learning, not an ideological position. [120]
Open question: which connections are needed, and what would show they are enough? Finding a document to support an answer is different from establishing that a system understands its meaning. [120][99]
Design and language
Anthropomorphism in AI
In plain language: Attributing human qualities, such as feelings or intentions, to a machine. [123]
A human-like voice or statements about having feelings can make a system seem like a person. The cited researchers examine how these design choices could encourage over-reliance or blur responsibility. Effects and useful safeguards depend on context. [123]
Map interpretation: Questioning human-like language does not by itself imply a position on either map axis. [123]
Open questions: which labels help people predict behavior, and which mislead? Claims that every provider uses AI terminology to avoid accountability require evidence about those providers. [123]
AI basics
Algorithm
In plain language: A set of steps a computer can carry out to produce a result. [125]
For example, a sorting algorithm puts a list of numbers in order. Training a language model uses algorithms to adjust its numerical settings. [125][90][95]
Map interpretation: Start here when separating the program’s procedure from the model it trains or runs. [125][95]
An algorithm can include random choices. Calling something an algorithm does not establish that every run gives the same result. [125]
AI basics
Training data & datasets
In plain language: Examples used to adjust a model. Their choice affects what the model learns. [129]
A dataset is a collection of examples, such as text or images. The training portion is used to fit the model; separate validation and test portions help assess its performance. [129][126]
Map interpretation: Ask where the examples came from and whether the evaluation uses genuinely separate material. [126]
A large dataset can still miss relevant situations. Testing on duplicated training examples can make a model look better than it is on new data. [129][126]
Inside a model
Pretraining
In plain language: An initial training stage that provides a starting model for later use or adaptation. [127]
For a language model, this often means adjusting weights on large text collections to predict missing or next text pieces. Later training can change how the model handles specific tasks. [127]
Map interpretation: Use this distinction when a proposal refers specifically to training new base models. [127]
A pretrained language model is not automatically an instruction-following assistant. Pretraining and later adaptation serve different objectives. [128]
Inside a model
Post-training
In plain language: Further training after pretraining, intended to adapt a model’s abilities and behavior. [128]
Developers can train on demonstrations, preferred answers or rewards. These methods may target instruction following, coding or responses to harmful requests. Different developers use different combinations. [128][102]
Map interpretation: Inspect the training goal, whose feedback was used and which results were tested. [102]
Post-training names a development stage, not a safety certificate. Improvements on selected tasks leave other failures possible. [128]
AI basics
Supervised learning
In plain language: Training with examples that include a target answer, such as a category or number. [129]
Software compares the model’s prediction with the target and uses the difference to adjust it. An illustrative task is sorting messages using examples already labeled as spam or ordinary mail. [129]
Map interpretation: Ask how the labels were made and whether they match the task you want to perform. [129]
Supervised describes the training examples. It does not mean a person checks every answer when the trained model is used. [129]
AI basics
Unsupervised learning
In plain language: Methods that find patterns in data without a supplied answer label for each example. [88]
Clustering is one example: it groups items by a chosen measure of similarity. Imagine grouping songs by their sound without giving the program genre labels first. [130]
Map interpretation: Treat discovered groups as results to inspect, not as automatic ideology memberships. [130]
Unsupervised does not mean free of human choices. The selected data and definition of similarity shape the groups. [130]
Inside a model
Self-supervised learning
In plain language: Training that builds its target answers from the data itself, such as a hidden word. [131]
Software can hide part of a text, ask the model to predict it and compare the prediction with the original. This supplies a training signal without someone labeling each example separately. [131]
Map interpretation: Read this alongside pretraining to see where a language model’s practice targets come from. [131]
The original text supplies a target, not independent proof that its claims are true. Self-supervised does not mean self-verifying. [131][120]
Inside a model
Reinforcement learning
In plain language: Training that uses rewards from attempted actions to adjust which actions a system selects. [132]
A system tries actions in an environment and receives numerical feedback. Training aims for more reward over time. A game score can supply feedback; other tasks need other reward rules. [132]
Map interpretation: Examine what earns a reward before interpreting claims that training improved behavior. [62]
A higher reward can miss the outcome people intended. RLHF is a specific approach using human feedback, not a name for all reinforcement learning. [62][102]
Inside a model
Deep learning
In plain language: Machine learning that builds several layers of learned calculations on top of one another. [133]
In a deep neural network, intermediate layers transform the numbers passed between input and output. Training adjusts settings across the network. These layers can build increasingly complex representations. [89][133]
Map interpretation: Use this term to distinguish a model family from a claim about how quickly AI should advance. [133]
Deep refers to layers of computation or representation. It is not a measure of wisdom, and there is no universally agreed minimum depth. [133]
Inside a model
Attention & self-attention
In plain language: A calculation that mixes information from different input positions with different weights. [70]
The weights depend on the input. In self-attention, parts of one sequence supply the information being combined. Multiple attention heads perform different learned combinations in parallel. [70]
Map interpretation: Use it to read Transformer diagrams. It describes a component, not an actor’s position. [70]
Attention is one operation within a larger network. The name does not imply a separate reader or human-like concentration. [70]
Inside a model
Model architecture
In plain language: The arrangement of a model’s parts and the calculations that connect them. [95]
An architecture specifies the structure, such as the number and kind of layers. A checkpoint supplies learned weights for that structure. Two models can share an architecture but have different weights. [95]
Map interpretation: Check whether a comparison concerns the structure, the trained weights or the surrounding application. [95]
An architecture diagram does not specify all learned values. Rebuilding the structure alone does not reproduce a trained model. [95]
Inside a model
Mixture of experts (MoE)
In plain language: A model design with several calculation blocks called experts; some versions activate only a subset for each input. [134]
In sparse MoE models such as Mixtral, a routing component selects only some expert blocks for each text piece. This allows more total parameters than are active for that piece. [134]
Map interpretation: Distinguish total from active parameter counts when reading model-size comparisons. [134]
Experts are numerical components, not independent professionals. The name alone does not show that each block has a clear subject specialty. [134]
Inside a model
Model distillation
In plain language: Training a model to match outputs from another model, often to make it cheaper to run. [135]
In the cited method, a larger model supplies probabilities as training targets for a smaller one. The target is output behavior, rather than copying the original model’s weights. [135]
Map interpretation: Ask which tasks the smaller model was tested on and how its results changed. [135]
The smaller model may not match every output or capability. Distillation does not guarantee identical performance. [135]
Inside a model
Quantization
In plain language: Representing model numbers with fewer bits to reduce memory use, with possible accuracy tradeoffs. [136]
Weights or intermediate values are mapped to a smaller set of representable numbers. For example, an eight-bit representation uses less storage per value than a 32-bit representation, but introduces approximation. [136]
Map interpretation: Compare memory needs, task results and performance on the hardware actually being used. [136]
Quantization changes numerical precision. It does not necessarily reduce the number of parameters, and faster execution depends on the implementation and hardware. [136]
Inside a model
Overfitting
In plain language: When a model fits its training examples so closely that it performs worse on new examples. [137]
Training can fit details that do not carry over to new data. A warning sign is training error falling while error on separate validation data rises. [137]
Map interpretation: Look for results on separate, relevant examples when assessing claims about model quality. [137]
A good training score alone does not show useful performance elsewhere. Underfitting is different: the model already struggles with its training examples. [137]
Inside a model
Underfitting
In plain language: When a model has not captured enough of the pattern even in its training examples. [138][137]
Possible causes include an unsuitable model structure, missing useful inputs or too little training. These are different problems, so simply adding more examples may not address the cause. [138]
Map interpretation: Ask whether the error already appears in training before comparing performance on new data. [137]
Underfitting is not a synonym for a small model. It describes an inadequate fit for the task and available data. [138]
Inside a model
Loss function
In plain language: A rule that turns a model’s training errors into a number to reduce. [139]
Different loss functions count errors differently. For example, squaring numerical prediction errors gives large mistakes more weight than taking their absolute size. [139]
Map interpretation: Ask what the training score measures and which real-world mistakes it leaves out. [139][62]
Lower loss means improvement under that scoring rule. It is not by itself proof of safe or useful behavior. [139][62]
Inside a model
Gradient descent
In plain language: A way to adjust model settings step by step in a direction that aims to reduce training error. [90]
Software calculates how a small change to each parameter would affect the loss, then updates the parameters in the opposite direction. Repeating these steps is part of many training procedures. [90]
Map interpretation: This explains what changing weights during training means in concrete computational terms. [90]
Reducing error on the training objective does not guarantee good results on new examples. Training and evaluation answer different questions. [126]
Evaluation method
Benchmark
In plain language: A shared test for comparing systems on chosen tasks; its score covers those tests, not every possible use. [101]
A benchmark combines tasks, scoring rules and a test setup. Using the same setup makes comparisons more informative. Accuracy, cost, bias and robustness can produce different pictures of the same model. [101]
Map interpretation: Atlas reading question: which tasks and conditions support a public capability claim? [101]
Evaluation is the broader investigation; a benchmark is one tool within it. A high score can leave relevant languages, users or failure types untested. [101]
Evaluation problem
Data contamination / benchmark contamination
In plain language: Test material appearing in training data can make an apparently new test partly familiar to a model. [153]
Overlap between training material and test questions or answers can weaken a test of performance on unseen material. Checking that overlap helps interpret a reported score. [153]
Map interpretation: Atlas reading question: how did an evaluator check that the test was meaningfully separate from training? [153]
Overlap does not prove that a model memorized an answer or gained an advantage. Brown and colleagues found that effects varied, and their detection method had limits. [141]
Model behavior
Generalization
In plain language: Doing useful work on examples outside the training set; success depends on how different those examples are. [87][143]
A model generalizes when patterns learned during training support good results on new examples. Recognizing a new photo of a familiar kind of object is an illustrative case. [87]
Map interpretation: Atlas reading question: does a claimed improvement hold beyond the examples used to develop the system? [143]
A new example can still closely resemble training data. Success there does not establish success in a different setting, and generalization is not a declaration of AGI. [143][101]
Evaluation property
Confidence calibration
In plain language: Checking whether a system's stated probabilities match how often its predictions turn out right. [140]
Illustrative example: among many predictions assigned an 80% chance of being correct, about 80% should be correct. Calibration concerns that match across cases, not certainty about one answer. [140]
Map interpretation: Atlas reading question: has a confidence number been checked against outcomes? [140]
High accuracy and good calibration are different properties. Confident wording in a chatbot reply is not itself a measured probability. [140]
Evaluation property
Predictive uncertainty
In plain language: A way to describe limits on a prediction, such as missing knowledge or noisy information. [142]
Researchers distinguish uncertainty due to limited knowledge in a model from uncertainty in the observations themselves. Estimating these separately can help show where more data may help and where observations remain ambiguous. [142]
Map interpretation: Atlas reading question: what does an uncertainty measure refer to, and how was it checked? [142][140]
Generating several different answers is not automatically a calibrated uncertainty estimate. The categories describe sources of uncertainty; calibration checks whether numerical estimates fit outcomes. [142][140]
Evaluation property
Robustness
In plain language: Maintaining useful performance when conditions change, rather than only succeeding in one test setup. [143]
Robustness asks how performance holds up across variations. For example, test a document reader on blurred scans as well as clean ones. The relevant changes depend on the intended use. [143]
Map interpretation: Atlas reading question: which difficult conditions were included in a claim that a system is reliable? [143]
Robustness is broader than resisting deliberate attacks. A system can handle one kind of change and fail on another; the tested conditions need to be named. [143]
Security concept
Adversarial examples
In plain language: Inputs deliberately changed to make a model give a wrong result; testing them can reveal weaknesses. [144]
In a classic image example, small carefully chosen changes cause a classifier to give the wrong label. The term covers crafted inputs, not simply every difficult or unfamiliar example. [144]
Map interpretation: Atlas reading question: does a security claim cover deliberate manipulation, and under what conditions? [145]
These attacks act on inputs during use. Data poisoning instead alters material used for training. A defense evaluated against one attack need not stop another. [145]
Evaluation method
Red teaming
In plain language: Deliberately probing a system for harmful or unwanted behavior so weaknesses can be investigated and fixed. [146]
Testers try challenging cases rather than only routine requests. People can devise the tests, and models can help generate candidates. The cited study used another language model to search for problematic replies. [146]
Map interpretation: Atlas reading question: who tested the system, what did they try, and what happened to the findings? [146]
Finding a failure does not measure how often it occurs in ordinary use. Finding none does not establish that every harmful behavior has been excluded. [146]
Model behavior
In-context learning
In plain language: Using material in the current prompt to adapt an answer, without a new training run. [92][147]
A prompt can provide examples or a pattern that guides the next response. This makes some task adaptation possible without a new training run. Researchers also study how this behavior arises inside transformers. [92][147]
Map interpretation: Atlas reading question: did a demonstration change the model itself, or only the information in its input? [92][147]
This is not a claim that a conversation permanently teaches the underlying model. Accounts of the internal mechanism remain dependent on the model and evidence studied. [141][147]
Prompting method
Few-shot learning / few-shot prompting
In plain language: Giving a model a few worked examples in the prompt to show the kind of response wanted. [92]
In LLM prompting, the examples are part of the input, not a separate training run. For instance, show two sample messages labeled by topic before asking it to label a third. [92]
Map interpretation: Atlas reading question: were examples supplied when a model's result was reported? [92]
Few examples in the prompt does not mean little prior training. Improvements vary by task and model; more examples do not guarantee a better result. [141]
Prompting method
Zero-shot learning / zero-shot prompting
In plain language: Asking a model to perform a task without giving worked examples in that prompt. [141]
For an LLM, a zero-shot test might give an instruction and a question but no demonstration of a correct answer. This tests what the already-trained model can do under that setup. [141]
Map interpretation: Atlas reading question: does a comparison use the same amount of help in each model's input? [141]
Zero-shot does not mean untrained, unfamiliar with the topic, or free of test contamination. It describes the immediate task setup. [141]
Prompting & generated text
Chain of thought (CoT)
In plain language: Generated intermediate steps before an answer, which can improve some task results but can also contain mistakes. [148][124]
Chain-of-thought prompting supplies examples with intermediate steps so a model produces a similar sequence. The original paper reported gains on selected tasks. Later reasoning systems can also be trained to generate extended intermediate text. [148][124]
Map interpretation: Atlas reading question: does a displayed explanation help check the answer, and what does it leave unverified? [124]
An explanation can be useful without faithfully reporting every influence on the answer. The label names generated text, not evidence of a human-like inner thought process. [154][124]
Interpretability question
Explanation faithfulness
In plain language: Whether an explanation accurately reflects what influenced an answer, rather than merely sounding plausible. [154][124]
Researchers can introduce a hint, observe whether it changes an answer, and check whether the explanation mentions it. Such tests examine a particular causal influence, not every step of the underlying computation. [154][124]
Map interpretation: Atlas reading question: what evidence connects an explanation to the process that produced the answer? [154][124]
A correct answer can have an incomplete explanation. The cited hint experiments found omissions in particular models and tasks; they do not show that every reasoning trace is useless. [154][124]
System safeguard
Guardrails
In plain language: Checks and restrictions intended to prevent unwanted inputs, outputs or actions in an AI application. [149]
The term covers different controls. A concrete implementation may check user input, filter retrieved documents, inspect a reply or validate a tool call before execution. The control's location matters. [149]
Map interpretation: Atlas reading question: which safeguard is actually enforced, at which step, and against which failure? [149]
A check on the final answer does not itself restrict an earlier tool action. Stating that guardrails exist does not establish their effectiveness against the relevant attacks. [149][145]
System design
Human in the loop (HITL)
In plain language: Giving people a role in reviewing or deciding an AI-assisted task, with benefits that depend on how the role works. [150]
A person may review a recommendation, correct an output or approve an action. This can combine different strengths, but the reviewer needs enough information, time and authority to disagree. [150][151]
Map interpretation: Atlas reading question: can the named reviewer meaningfully change or stop what happens? [151]
A confirmation button alone does not establish effective oversight. Experiments have found that erroneous system advice can influence a person's judgment even when that person makes the final decision. [151]
Human decision pattern
Automation bias
In plain language: Following automated advice too readily, including wrong advice; useful assistance can still create this problem. [152]
A person may accept a system's suggestion instead of checking it against other evidence. One study of pathology experts found better overall performance alongside some cases where wrong advice displaced a correct judgment. [152]
Map interpretation: Atlas reading question: how does a workflow help people detect and reject a wrong recommendation? [151]
This is a possible failure of reliance, not proof that people always trust machines or that all AI assistance reduces accuracy. Effects depend on the task and interaction. [152][150]
Security concept
Data poisoning
In plain language: Deliberately altering training material to influence a model's later behavior in an attacker's favor. [145]
An attacker inserts or changes training examples, for instance to cause particular errors or make a hidden trigger affect outputs. Poisoning can target initial training or later training stages. [145]
Map interpretation: Atlas reading question: how are the origins and integrity of training material checked? [145]
Accidental low-quality data is not necessarily poisoning. Data checks and filtering can help, but NIST notes that finding malicious examples in a large training collection can be difficult. [145]
Application
Computer vision
In plain language: Software that extracts information from images, such as objects or written text. Different tasks need different checks. [155]
Computer vision covers tasks such as labeling an image, locating objects and reading text in a picture. For example, a system could mark the position of a bicycle in a photo. [155]
Map interpretation: Map context: this names a set of technical tasks. We do not assign it a preferred pace of AI development. [155]
Image classification assigns a category; object detection also locates objects. Neither task is the same as producing a new image. [155]
Application
Automatic speech recognition (ASR)
In plain language: Software that turns spoken audio into written words. A readable transcript can still contain mistakes. [156]
An ASR system processes a recording and produces text, for example subtitles or a draft meeting transcript. The Hugging Face examples show that a plausible word can replace the word actually spoken. [156]
Map interpretation: Map context: speech transcription is a capability, not a position on AI risk or development speed. [156]
Transcribing words, identifying a speaker and translating into another language are different tasks. Check important wording against the recording. [156]
Application
Speech synthesis / text-to-speech (TTS)
In plain language: Software that produces spoken audio, often from written text, for example to read a document aloud. [157]
Text-to-speech systems turn text into a speech signal. They can be used to read a document aloud or provide an audio version of a message. Speech generation is one type of audio generation. [157]
Map interpretation: Map context: this describes an output method. We do not treat a synthetic voice as an actor or an ideology. [157]
Speech recognition turns audio into text. Speech synthesis goes the other way. Producing speech does not necessarily mean imitating a particular real person. [157]
Inside a model
Diffusion models
In plain language: Generative models that can build an image through repeated removal of noise. Generating a picture does not establish that the pictured event happened. [158]
In the denoising approach, training examples are corrupted with noise and a model is trained to reverse that process. Generation starts with noise and applies learned steps to produce an output. [158]
Map interpretation: Map context: this is a way to generate content. Its use does not establish a view about catastrophic risk or faster development. [158]
The classic diffusion approach refines a noisy representation over successive steps. That differs from a language model generating text one token at a time; both are computational methods. [158]
Inside a model
Latent space
In plain language: An internal system of numerical coordinates for representing features, such as features of an image. [159]
In latent diffusion, an image is compressed into a learned numerical representation. The diffusion process works on that representation, then a decoder converts the result into pixels. The set of possible representations is called a latent space. [159]
Map interpretation: Map context: latent coordinates belong to a model representation. They are unrelated to the editorial coordinates used for public positions in this atlas. [159]
A latent representation is not a miniature image or a written explanation. In latent diffusion, compression reduces detail as well as computational work; its design affects reconstruction quality. [159]
Using AI
Synthetic data
In plain language: Artificially generated examples used as data. They can support testing and analysis, but being synthetic does not guarantee privacy. [160][161]
A generator can produce new rows that resemble patterns in an original table. The resulting dataset can be compared with real data to check whether useful patterns were preserved. For example, a team could generate sample customer records to test software. [160]
Map interpretation: Map context: this is a data-production method. We do not place it on the pace or concern axes. [160]
Synthetic content describes generated media; synthetic data emphasizes its use as examples for analysis, testing or training. Some generation and evaluation methods can expose information about original records, so privacy needs separate assessment. [160][161][112]
Using AI
Knowledge cutoff
In plain language: A reported date describing the freshness of a model’s training information. It is not a guarantee of complete or correct knowledge before that date. [162]
A provider may state when training material was collected. Researchers distinguish that reported date from an effective cutoff: how recent the information a model can actually reproduce is for a particular topic or source. [162]
Map interpretation: Map context: a cutoff is a limit to check when using a model as an information source. It is not a measure of intelligence or a map position. [162]
Different topics can have different effective cutoffs. A single advertised date should not be read as a promise that every earlier event is covered or that every answer is up to date. [162]
Using AI
Data provenance
In plain language: A record of where data came from and how it changed. It helps people assess evidence but does not prove the data is correct. [163]
Provenance describes the sources, people and processes involved in producing data. It can record that one dataset was derived from another, who carried out a step, and which version was used. [163]
Map interpretation: Map context: we use provenance to let readers inspect the origin of claims. It is a documentation practice, not a political position. [163]
An origin record and a truth assessment answer different questions. Knowing who supplied a claim helps investigation; the claim and the record can still require checking. [163]
Using AI
Digital watermarking
In plain language: Information embedded in text, images or audio to help identify its origin. A watermark can be missed, removed or wrongly detected. [112]
A watermark places a signal inside the content itself. It might be a visible mark or a subtle pattern detectable by software. Some schemes indicate that a particular generator produced the content. [112]
Map interpretation: Map context: watermarking is one approach to documenting media origins. It does not tell us where an actor belongs on the map. [112]
A watermark differs from a separate metadata record. Its presence does not establish that a claim is true, and its absence does not establish that a human made the content. [112]
Access & release
Model card
In plain language: A document describing a model’s intended uses, tests and limitations. It helps scrutiny but is not a safety certificate. [164]
The model-card proposal calls for information about what a model is for, how it was evaluated and how its performance varies across relevant conditions or groups. A useful card makes those limits visible alongside results. [164]
Map interpretation: Map context: a model card is evidence to inspect when evaluating claims about a model. Publishing one does not establish an organization’s policy position. [164]
A model card focuses on a model. A system card considers how models and other parts work together in a product. Neither title alone establishes how complete the reporting is. [164][165]
Access & release
System card
In plain language: A document about how the parts of an AI system work together and where its limits lie. It can become outdated as the product changes. [165]
A system card can describe models, non-AI components, data flows and the way outputs affect a product. Meta’s proposal explains why a model’s behavior needs to be considered in the system where it is used. [165]
Map interpretation: Map context: system documentation helps readers inspect a product beyond an individual model. We do not infer an actor’s beliefs from the existence of a card. [165]
The term describes documentation, not an independently verified stamp of approval. Read its scope, date and limitations; a card about one version may not describe another. [165]
Inside a model
Compute
In plain language: The computing resources used to train or run a model. More resources describe an input, not whether the result is useful. [87]
Compute can refer to resources such as processing power, memory and storage. In AI discussions, it is worth checking whether someone means resources for training a model or resources for running it afterwards. [87]
Map interpretation: Map context: this is a resource term. A claim about compute alone does not establish a preference for faster or slower frontier development. [87]
Compute, model size and elapsed time are different descriptions. A parameter count describes model values; compute describes resources used in doing the work. [87]
Inside a model
Graphics processing unit (GPU)
In plain language: A chip designed to do many calculations in parallel. It can accelerate suitable AI workloads, but it is not itself an AI model. [166]
GPUs began as processors for graphics. Their parallel design is also used for calculations in model training and generation. A GPU runs software; the trained model’s numerical values are separate from the chip. [166]
Map interpretation: Map context: a GPU is hardware. Owning or using one does not identify a movement or a position on AI risk. [166]
A CPU emphasizes fast sequences of operations; a GPU emphasizes many operations in parallel. Which is useful depends on the work, so a GPU is not automatically faster for every program. [166]
Using AI
Latency
In plain language: The waiting time for a system’s response. For generated text, the first visible word and the completed answer have different waiting times. [167]
Latency is a duration. Language-model benchmarks can measure time until the first output token and the time taken to produce later tokens. These measures describe different parts of the waiting experience. [167]
Map interpretation: Map context: response speed is a product-performance measure. It is separate from the map’s axis about the pace of developing more capable AI. [167]
Throughput describes how much work a system processes in a period. A system that handles many requests overall does not necessarily answer each individual request quickly. [167]
Application
On-device AI / edge AI
In plain language: Running a model on a device such as a phone or laptop. Where a calculation runs does not by itself establish the privacy of the whole app. [168][165]
On-device AI runs model calculations locally. A model can be obtained or trained elsewhere, adapted for the target device and then run there. Edge AI is also used for processing near the source of data. [168]
Map interpretation: Map context: this describes where computation happens. We do not treat local processing as a movement or a position on frontier pacing. [168]
On-device inference is different from training on the device. To assess privacy, check the whole application’s data flows, including any other services it uses. [168][165]
AI basics
Software & computer programs
In plain language: Computer programs tell a machine what operations to carry out. Software includes those programs and their associated data. [176][177]
A calculator program might multiply two numbers using a written rule. A language-model program runs calculations using numerical settings learned during training. Both are executed by computer hardware. [97][95][88][177]
Map interpretation: Map context: this explains how systems work; it is not a position for or against faster AI development. [177][95]
Being software does not make a model’s behavior easy to explain. Interpretability research studies how its learned computations produce particular outputs. [73]
AI basics
AI system
In plain language: An AI model together with the software and hardware that put it to use. [178][77]
For a chatbot, this can include the interface, instructions, model, search tools and permissions. A robot can also include sensors and motors. These parts shape what the system can do. [97][81][178]
Map interpretation: Map context: distinguish a claim about a model from a claim about the application or machine that uses it. [178][77]
A model’s results alone do not establish that a whole system is safe. Tools, access rights and oversight need their own assessment. [80][81]
Theoretical thesis
AI as normal technology
In plain language: Narayanan and Kapoor's view that people can shape AI's impacts through institutions, engineering and policy. Its safety claims are contested: in September 2026, they acknowledged underestimating risks during development and called for stronger controls. [188][190][194]
The thesis compares AI with earlier powerful technologies. It separates technical progress from applications and adoption, and treats continued human control as a goal requiring choices. Normal does not mean harmless, simple or predictable. [188][189]
Map interpretation: Map interpretation: context across both axes, not a measured point or membership label. Its authors support pausing unsafe experiments where needed, while rejecting claims that catastrophic risks are imminent. [194]
Scott Alexander argues that rapid self-improvement and adoption by AI labs could bypass assumed limits. The authors reply that technical improvements do not automatically remove external constraints. Their September update gives more weight to developer oversight. [190][189][194]
AI basics
Symbolic AI & expert systems
In plain language: Programs that work with explicitly represented facts and rules. An expert system applies such rules within a particular subject. [195][196]
Developers encode facts and logical rules that software applies to a problem. An expert system combines a knowledge base with rules, often written as if-then conditions, drawn from a specific field. [195][196]
Map interpretation: Map context: this explains an AI approach based on explicitly programmed knowledge. [195]
An expert system is one application of symbolic methods. Its encoded knowledge covers a limited domain. This differs from fitting a model's parameters to examples through machine learning. [196][88]
Capability concept
Narrow AI / specialized AI
In plain language: AI for a limited task or set of tasks. Specialized software can perform very well within that scope. [58]
Narrow describes breadth, not quality. A research framework separates how many kinds of tasks a system handles from how well it performs them. [58]
Map interpretation: Map context: keep specialist performance separate from claims about general intelligence. [58]
Breadth, performance and autonomy are different questions. Strong task results do not determine how much control an application should receive. [58]
AI basics
Foundation models
In plain language: Models trained on broad data that can be adapted for many tasks. Reusing one model can also spread its weaknesses across applications. [197]
A shared trained model becomes a starting point for different uses, through prompts or further training. The term includes more than language models. [197]
Map interpretation: Map context: shared models connect development choices to many downstream applications. [197]
A foundation model is a reusable component. Its broad training does not establish equal performance across applications. [197]
Research concept
AI scaling laws
In plain language: Measured patterns linking model size, training data and computing resources to performance. Each pattern concerns a particular measurement and training setup. [198][199]
Researchers train different-sized models and fit equations to the results. In language-model research, a common measure is prediction loss: how poorly the model predicts the next text piece. [198][199]
Map interpretation: Map context: these studies inform expectations about capability growth and resource use. [198][199]
Larger is not the only choice. Hoffmann and coauthors found better results by balancing model size with more training data. Extrapolating a measured trend to untested scales adds an assumption. [198][199]
Model behavior
Model collapse
In plain language: Repeatedly training on earlier models' output can erase patterns from the original data. Other experiments avoided this degradation by retaining the original dataset while adding generated examples. [200][201]
In studied training loops, each model supplies examples for the next. Errors can accumulate, and uncommon patterns can disappear from the later models' output. [200]
Map interpretation: Map context: this qualifies claims about generated data sustaining future model training. [200][201]
The result depends on how datasets are replaced or accumulated. It does not show that all synthetic data is harmful, or diagnose a chatbot's bad answer without examining its training. [200][201]
Work & society
Human data work & data labeling
In plain language: People collect, label, check and evaluate data used in AI. Their work and the instructions they receive help shape trained software. [202][203]
Data labeling attaches categories or other annotations to examples. Wider data work includes collecting material, preparing datasets and checking model outputs. Research documents this work in different employment arrangements. [202]
Map interpretation: Map context: AI development includes labor and organizational choices, alongside model design and computing resources. [202][203]
A label is a judgment made for a task, not automatically a fact. A study of AI practitioners found that neglected data work could cause problems later in development and deployment. [202][203]
Research field
AI control
In plain language: Tests whether safeguards can block harmful actions even when model outputs are chosen to bypass them. Monitoring models have also been bypassed in such tests. [204][205]
Researchers test whole workflows against adversarial behavior. The original control study used programming tasks and tested reviewing or editing untrusted code with another model. [204]
Map interpretation: Map context: connects permissions, monitoring and review to safeguards around deployed software. [206]
Control can add checks around a model without changing its learned weights. It can complement alignment work. A passed test supports only its stated setup and attacks. [204][205]
How language models work
Depth changes the wording, examples and drawing detail. The claims and their sources stay the same at every depth; the default text is the At work depth.
First look: For readers from about age 10 with no computing background. Short sentences and everyday words.
Look closer: For secondary school and curious adults. Introduces tokens, scores, training, sampling and seeds.
At work: For people deciding whether to rely on an AI tool. Adds context, retrieval, tools and review steps.
In depth: For developers, researchers and students of machine learning. Uses the technical terms and links primary sources.
Programmed calculations, learned settings
A language model runs on a computer. Training adjusts its number settings using examples; generating a reply uses those settings. The arrows below show these separate processes. Model calculations and training · Text-generation software.
TrainingAdjusting the model
Example materialText used during training
Training calculationsAdjust numbers inside the model
A trained modelIts learned numbers are parameters, often called weights
1 · InputText goes inThe supplied text is represented as pieces called tokens
2 · ModelCalculate next-piece scoresThe software uses the learned weights and the text so farIn depth: tokens → embeddings → attention blocks → scores (logits) → softmax with temperature. Hardware and library versions can shift these numbers.
3 · ChooseSelect one pieceChoose the highest score, or make a random choice using calculated probabilitiesIn depth: greedy, sampling, top-k or top-p. Sampling deliberately adds randomness to token selection. Hardware and software can also affect repeatability.
4 · ContinueAdd it to the textUse the longer text for the next round, until a stopping rule is met
It shows the main steps, not every calculation inside a model. Knowing those calculations does not yet give a full explanation of how the model arrived at a particular answer. Read the research on tracing model behavior.
Your request becomes an input to a computer program
Claim C1
Imagine asking for a program that totals a shopping bill. A large language model (LLM) runs as software on a computer. It uses programmed calculations and learned numerical settings called weights. Training adjusts those settings using examples. When drafting your program, it scores possible next tokens, which are pieces of text. [70][71][75]
The same claim at other reading depths
First look: Imagine asking for a program that adds up a shopping bill. A language model is also a computer program. Programmers write its calculations. Training uses examples to adjust many number settings inside it. When writing, it uses those settings to score the text pieces that could come next. A piece might be a word or part of one. [70][71][75]
Look closer: Imagine asking for a program that totals a shopping bill. A large language model (LLM) runs as a computer program. Programmers specify its calculations; training uses examples to adjust many numerical settings, called weights. When it answers, its neural network uses those weights to score possible next tokens, the pieces of text it works with. [70][71][75]
In depth: An LLM combines programmed calculations with learned parameters. In the text-generation setup shown here, a transformer maps a token sequence to scores for the next token. Training adjusts its parameters using a next-token objective on example text; further tuning stages can follow. During inference, the programmed forward pass applies the trained parameters to the supplied context. A draft shopping program is an output of that process. It still needs review and testing before anyone runs it. [70][71][75][92]
The answer is built one piece at a time
Claim C2
Software chooses a token from the model's scores, adds it to the answer and repeats. It can pick the highest-scoring option or sample among several options. Sampling can produce different drafts of your shopping program. The model can also run without sampling; variation is a separate question from how well we understand its workings. [71][72][73]
The same claim at other reading depths
First look: The answer is not written all at once. The model scores the pieces that could come next. The software picks one piece, adds it to the answer, and repeats the calculation. It can always take the top score, or pick at random using the scores. Random picking can make the same question produce different answers. [71][72][73]
Look closer: The model gives each possible next token a score. Software then chooses one token, adds it to the answer and repeats until a stopping rule is met. It can always choose the highest score, or sample: pick randomly, with higher scores more likely. Sampling can produce different drafts of your shopping program. Running without sampling removes that randomness, but it does not explain why the model scored things the way it did. Those are two separate questions. [71][72][73]
In depth: Each forward pass yields logits over the vocabulary; a decoding strategy selects one token, appends it to the context, and the pass repeats. Greedy decoding takes the highest score; sampling draws from the softmax distribution, optionally reshaped by temperature, top-k or top-p, until an end token or a length limit. Sampling is the intended source of output variance. Its absence says nothing about interpretability: a greedy run is still produced by the same learned weights, whose contributions interpretability research can only partly trace. [71][72][73]
Check the draft, then run the saved program
Claim C3
The model might draft a rule that multiplies price by quantity. Check that it matches your needs and test it. Once saved, that code can run without the LLM. For the same complete inputs and execution conditions, deterministic code gives the same result. It can also repeat the same mistake, such as leaving out delivery charges. [75][72]
The same claim at other reading depths
First look: The model might write a rule: price times how many. Someone must read it and test it. Once saved, this rule can run without the model. With the same numbers and setup, this multiplication gives the same answer. If the rule leaves out the delivery charge, it leaves it out every time. [75][72]
Look closer: The model might draft code that multiplies price by quantity. Check that draft against what you need and test it. Once saved, this multiplication can run without the model. For fixed inputs and execution conditions, it gives the same result; this is called deterministic. Saving generated code does not by itself make every program deterministic or correct. The shopping rule can repeat the same mistake, such as leaving out delivery charges. [75][72]
In depth: Review and test a generated code draft before running it. In the shopping example, a saved multiplication function can run without calling the model. When that function uses the same complete inputs and fixed execution conditions, its result is deterministic. This describes the specified function, not all generated software. A repeatable function can still implement the wrong rule. Passing tests cover the cases checked; they do not establish correctness for every input or explain how the model produced the draft. [75][72]
Writing code and running it are separate steps
Describe the taskExample: “Multiply price by quantity”
Model drafts codeThe draft is text that needs checking
Review and testDoes the code follow the intended rule?
Run the saved codeThis small function can run without calling a model
An illustrative workflow based on GitHub’s code-review guidance. Repeating a calculation does not show that it is the right calculation for the job.
Why did the model choose that answer?
Claim C4
You can inspect the shopping program's calculation. Explaining why the model drafted that particular program is harder: its behavior depends on many learned weights working together. Knowing the calculations does not automatically reveal a simple reason for each answer. Researchers studying mechanistic interpretability trace parts of this process, with important gaps still remaining. [70][73][74]
The same claim at other reading depths
First look: With the saved shopping rule, you can follow the line that multiplies the numbers. The model also runs code, but explaining its answer takes more than reading that code. Many learned settings work together to score each piece. Scientists can trace some of what happens inside, but important gaps remain. [70][73][74]
Look closer: You can read the shopping program and follow its calculation. Explaining why the model drafted that particular program is harder. The answer depends on many learned weights working together. Knowing every calculation does not give a simple reason for each answer. Researchers in mechanistic interpretability trace parts of this process, and important gaps remain. [70][73][74]
In depth: Every operation in the forward pass is known; what is missing is a compact causal account of why those operations produced this output. Mechanistic interpretability methods such as attribution graphs and attention tracing recover partial circuits for specific behaviors, with stated limits on coverage and faithfulness. Opacity in this sense is a property of the learned function, independent of whether decoding is stochastic. [70][73][74]
An illustration: knowing the rules is not a complete explanation
Claim C4
Weather helps illustrate the gap described above: knowing physical rules does not give a simple explanation of every outcome. Forecasts face uncertainty about starting conditions and approximations. With an LLM, the question here is how learned parts produce an answer. This is our limited analogy about explanation, not a claim that LLMs behave like weather or must be random. [76][73][74]
The same claim at other reading depths
First look: Here is a limited comparison. Knowing weather's physical rules does not fully explain every outcome. Forecasts face gaps in starting information and use simplified calculations. For a model, the question is how its learned settings help produce an answer. These are different gaps. The comparison does not show that models behave like weather or must be random. [76][73][74]
Look closer: An analogy: we know the physical rules of weather, but a forecast still cannot explain every outcome, because starting conditions and approximations are uncertain. With a model the open question is different: how learned parts combine to produce an answer. This is a limited analogy about explanation. It is not a claim that models behave like weather or must be random. [76][73][74]
In depth: The comparison is a limited analogy. Numerical weather prediction has known dynamics and still limited explanatory power per outcome, from initial-condition uncertainty and approximation. For an LLM the operations are also known, and the gap is attributional: which learned components caused this output. The analogy illustrates that knowing the rules is not a complete explanation. It does not import chaotic sensitivity, a claim that models must be random, or any physical claim about models. [76][73][74]
More detail: what makes an answer repeatable?
Claim C3
Turning off sampling does not guarantee identical answers across every setup. Hardware, software versions and some calculations can still affect results. A random seed helps repeat a sequence of choices under matching conditions. Repeatability means reproducing a result; interpretability means explaining how it came about. One does not establish the other. [72][73]
The same claim at other reading depths
First look: Even with random picking turned off, two computers can give slightly different answers. Different machines and different software versions can round numbers differently. A seed number lets you replay the same random picks on the same setup. Getting the same answer again is not the same as understanding why it was the answer. [72][73]
Look closer: Turning off sampling does not guarantee identical answers on every setup. Different hardware, software versions and some calculations can still change the result. A seed replays the same sequence of random choices under matching conditions. Repeatability means getting the same result again; interpretability means explaining how it came about. One does not establish the other. [72][73]
In depth: Greedy decoding removes sampling variance but not all nondeterminism: floating-point non-associativity, nondeterministic kernels and library versions can shift the computed scores, so even greedy runs can differ across setups. Seeds replay a sampled sequence only under matching conditions, and a low temperature makes the top token more likely without guaranteeing identical runs. Reproducibility is a property of the execution setup; interpretability is a property of the explanation. Neither implies the other. [72][73][71]
Why can a confident answer be wrong?
Claim C5
Producing convincing text does not check whether every claim is true. An answer can include a made-up fact, quotation or reference; this is often called a hallucination. Open an important citation and check that it exists and supports the claim. An answer's confident tone is not a measure of its accuracy. [83][227]
The same claim at other reading depths
First look: The model is good at writing text that sounds right. Sounding right is not the same as being right. It can invent a fact, a quote or a book that does not exist. People call this a hallucination. If it matters, look it up. A sure-sounding answer can still be wrong. [83][227]
Look closer: Producing convincing text does not check whether each claim is true. An answer can include a made-up fact, quotation or reference; this is often called a hallucination. Open an important citation and check that it exists and says what the answer claims. A confident tone is not a measure of accuracy. [83][227]
In depth: Generation selects tokens using scores computed by the model. The resulting text can still contain fabricated facts, quotations or references. Neither fluency nor a high next-token score establishes that a claim is true. Where accuracy matters, inspect the cited source and check that it supports the answer. Retrieved passages can be useful evidence, but still need checking. [83][227][99][71]
Does it learn from our conversation?
Claim C6
Your messages can shape the next answer through the conversation context. Some apps also save information for future chats. Neither is the same as retraining the model's weights. Separately, providers may use conversations for later training under their policies and settings. Check both memory controls and training controls for the service you use. [92][118][119]
The same claim at other reading depths
First look: An app can include earlier messages when it prepares the next input. That can shape the answer without changing the model's learned settings. Some apps also save notes for later chats. Separately, companies may use chats to train models under their rules and your settings. Check the controls for saved notes and training. [92][118][119]
Look closer: An application can include earlier messages in the input, where they can shape the next answer. Some apps also save information for future chats. Providing this text does not itself update the model's weights, the numbers adjusted during training. Separately, a provider may use conversations for later training under its policies and settings. Check both the memory controls and the training controls of the service you use. [92][118][119]
In depth: Context, application memory and training are different mechanisms. Text included in the context conditions the model's next output without itself updating the parameters. Memory features can save information and supply it in later conversations; that is separate from training or tuning the weights. Whether a provider uses conversations for later training depends on its policies and the relevant account settings. Check both memory and training controls for the service in question. [92][118][119]
Is generating an answer the same as searching the web?
Claim C6
A language model can generate text without searching. An application can also retrieve webpages or documents and give that material to the model. These are separate steps. Retrieved material can contain errors, and the answer can misrepresent it, so follow the sources when accuracy matters. [71][99][83]
The same claim at other reading depths
First look: A model can write an answer without looking anything up. Some apps first search the web or your files and hand the results to the model. Those are two separate steps. What the search found can be wrong, and the answer can get it wrong too. When it matters, open the source. [71][99][83]
Look closer: A language model can generate text without searching. An application can also retrieve webpages or documents first and add that material to the model's input. These are separate steps done by different parts of the system. Retrieved material can contain errors, and the answer can misrepresent it, so follow the sources when accuracy matters. [71][99][83]
In depth: Retrieval-augmented generation adds a retrieval step whose results are inserted into the context; the model's parameters are unchanged and it still generates by next-token prediction over that context. Retrieval quality bounds what the model can be grounded on, and generation can still misstate what was retrieved. Verify against the retrieved source, not the paraphrase. [71][99][83]
Does sounding human mean it has feelings?
Human-like conversation alone does not establish subjective experience: whether anything feels like something to the system. Researchers disagree about how to assess this. The linked study proposes indicators based on theories of consciousness, with explicit assumptions and uncertainty; a chatbot's claim to have feelings does not settle the question. [116]
The same claim at other reading depths
First look: A model can chat like a person. That does not show it has feelings. Researchers do not agree on how to test this. One study lists signs to look for, and says the question is open. If a chatbot says it has feelings, that is text it produced, not proof. [116]
Look closer: Sounding human does not establish subjective experience: whether anything feels like something to the system. Researchers disagree about how to assess this. The linked study proposes indicators drawn from theories of consciousness, with explicit assumptions and uncertainty. A chatbot's statement that it has feelings is generated text and does not settle the question. [116]
In depth: Conversational fluency alone does not establish subjective experience. The linked report derives indicator properties from scientific theories of consciousness and applies them under stated assumptions, with substantial uncertainty. A model's report of having feelings is also generated output. Interpreting such a report requires an account of how it was produced and what would count as evidence; the report alone does not settle whether subjective experience is present. [116]
Will AI replace my job?
A job usually combines several tasks. Automating one task does not show that the whole job will disappear. The ILO's 2025 study estimates which tasks could be affected; it does not count actual job losses. Adoption, cost and choices about how work is organized also affect what happens. [117]
The same claim at other reading depths
First look: A job is made of many tasks. A model may do one task well. That does not mean the whole job goes away. One big study looked at which tasks could change. It did not count lost jobs. What happens also depends on what companies choose to do. [117]
Look closer: A job usually combines several tasks. Automating one task does not show that the whole job will disappear. The ILO's 2025 study estimates which tasks could be affected; it does not count actual job losses. Adoption, cost and choices about how work is organized also shape what happens. [117]
In depth: Exposure estimates are task-level: they score which tasks could be performed by generative AI, not observed displacement. The ILO's 2025 study reports exposure under explicit assumptions about capability, with no measurement of actual job losses. Adoption rates, cost and organizational choices mediate outcomes, so exposure describes potential, not a forecast. [117]
Does fluent language mean understanding?
The critical argument is that convincing language does not by itself show an understanding of meaning. Bender and Koller focus on training from language patterns alone. Other research asks what a prediction model learns internally: the Othello study found information about a game board. These address related questions, but neither establishes what every current LLM understands. [120][122]
The same claim at other reading depths
First look: A convincing answer alone does not prove that a model has learned meaning. One argument says text patterns alone are not enough to learn meaning. A separate study found signs of a game board inside a model trained on game moves. These studies ask related questions. Neither settles what every language model can do or how to describe it. [120][122]
Look closer: A convincing answer alone does not settle whether a model has learned meaning. Bender and Koller argue that training only on language patterns is insufficient for that. Another study examined a model trained to predict Othello moves and found information about the board inside it. That is a specific result about a game model. It does not settle the wider debate about understanding in language models. [120][122]
In depth: Bender and Koller distinguish linguistic form from meaning and argue that training on form alone cannot establish the connection they call meaning. The Othello study investigates learned representations in a model trained to predict game moves and finds information about board states. This provides evidence about that model's internal structure, not a general verdict on human-like understanding. The studies address related but different questions. Their results do not establish what every current LLM understands. [120][122]
Which questions are still open?
How reliably can a model learn a new concept from a few examples, explain causes, or revise a plan when its goal changes? Which tests separate those abilities from shortcuts? How faithfully does written reasoning report what influenced an answer? How do human-like names and voices affect trust and responsibility? How much of human thinking is captured in the explanations used for training? The cited studies do not settle that last question or establish any provider's motives. [121][124][123]
The same claim at other reading depths
First look: Can a model use a new idea after just a few examples? Can it explain causes or change plans? Which tests would rule out shortcuts? Do its written reasons show what shaped the answer? Do names and voices affect trust? How much human thinking appears in training text? The cited studies do not settle every question or reveal a company's motives. [121][124][123]
Look closer: How reliably can a model learn a new idea from a few examples, explain causes, or change a plan when its goal changes? Which tests could separate those abilities from shortcuts? Do its written reasons reflect what actually shaped its answer? How do human-like names and voices affect trust and responsibility? How much human thinking is captured in the explanations used for training? The linked studies do not settle that last question or establish a provider's motives. [121][124][123]
In depth: Open questions include how to test learning from limited examples, causal reasoning and revision of plans under changed goals while ruling out task-specific shortcuts. Reasoning-trace faithfulness asks whether a written explanation reports factors that actually influenced an output. Human-like interfaces raise separate questions about trust and responsibility. Another question is how much of human cognition written explanations capture when those explanations become training material. The cited studies do not settle that last question or establish universal conclusions about providers' motives. [121][124][123]
Before you rely on an AI tool
Use these questions to examine what an AI application can do and how errors could be caught. We assembled them from the linked guidance and research; they are a starting point for review, not a safety certificate. [77][83]
An example with approval before sending
Imagine an application that helps answer customer email.
1 · InputCustomer email and reply instructions
2 · Model proposesDraft a reply
3 · Outside the modelA person checks the recipient and textThe application waits for approval
An AI application combines a model with other software. People configure its instructions, data sources, connected tools and access rights. These choices affect what it can do. A feature that drafts text needs a different review from one that can change records or send messages. Start by listing the application’s actual access. [77][80]
Ask: What can this application read, change or send, and through which tools?
Decide how you will check the result
Define a successful result and which mistakes would matter. Then consider whether fixed code, a model or a combination fits the job. Test generated code against its requirements. A convincing demonstration is a reason to investigate further; reliable use needs tests that reflect the intended task. [78][75][83]
Ask: How will we check that it did the right job, including difficult cases?
Set limits before it takes action
Give the application only the tools and access it needs. Use permissions in the connected software to enforce those limits, with approval before consequential actions where appropriate. Instructions tell a model what you want; access controls determine what it can reach. Include connected services and their access keys in the review. [81][78]
Ask: Which actions need approval, and what enforces that requirement?
Keep a useful record of what happened
Record relevant inputs, answers and tool activity while protecting sensitive information. These records help investigate problems. Written reasoning steps can add clues, but give an incomplete account of a model's internal calculations. In its March 2026 report, OpenAI describes monitoring reasoning and actions together and acknowledges uncertainty about missed incidents. [79][73][84][85]
Ask: What evidence would reveal a problem, who reviews it, and how quickly can they respond?
Test how things could go wrong
A document or website can contain instructions that redirect the model: prompt injection. Test the complete application with checks on input and output, limited permissions, separation from sensitive systems and a response plan. Say what each test covered. A successful example or an empty alert log leaves other possible failures untested. [82][77][83][85]
Ask: What exactly was tested, under which conditions, and what conclusion does that evidence support?
AI development and the debate around it
Selected milestones, with their sources and limitations. Card spacing does not represent elapsed time.
Foundations & evaluation
Turing asks how to test machine conversation
Alan Turing proposes an imitation game to make questions about machine intelligence more concrete. [216]
What changed, and the limits
What changed
In Computing Machinery and Intelligence, Turing describes a judge questioning hidden participants through written messages, then asks what would happen if a computer took one participant's place. He also discusses computers that learn. [216]
What this result does not establish
This is a proposed test and a philosophical argument. The paper does not report a computer passing it, and Turing discusses objections to using the game as a test of thinking. [216]
Researchers gather under the name artificial intelligence
The Dartmouth summer project brings researchers together to study artificial intelligence. [217][218]
What changed, and the limits
What changed
The proposal, written in 1955, uses the name artificial intelligence and sets out questions about language, learning, neural networks and problem-solving. Dartmouth's history records the gathering in the summer of 1956. [217][218]
What this result does not establish
The proposal calls its central idea a conjecture: that aspects of intelligence can be described precisely enough for a machine to simulate them. It sets a research agenda, rather than reporting that this had been achieved. [217][218]
Joseph Weizenbaum describes ELIZA, a program that responds to typed conversation by rearranging text. [219]
What changed, and the limits
What changed
ELIZA looks for keywords, matches patterns and uses a script to assemble a reply. The paper explains how a sentence about being unhappy can be turned into a follow-up question using these rules. [219]
What this result does not establish
Weizenbaum shows how the transformations can work without understanding the meaning of the text being rearranged. A reply that fits a conversation can give a misleading impression of what the program knows. [219]
A learning method adjusts a neural network's inner layers
Rumelhart, Hinton and Williams show how backpropagation can help neural networks learn useful internal patterns. [220]
What changed, and the limits
What changed
The method works backward from the difference between the desired answer and the network's answer. It calculates how to adjust the connections inside the network, allowing its intermediate layers to learn features useful for a task. [220]
What this result does not establish
The setup needs a specified task and desired outputs to compare against. Reducing that error is the learning objective; the method does not choose the purpose of the system. [220]
IBM's Deep Blue defeats reigning world chess champion Garry Kasparov in a six-game match. [221]
What changed, and the limits
What changed
IBM records a 3.5–2.5 result under standard tournament time controls. The system combines fast searches through possible chess positions with chess-specific evaluation, databases and advice from grandmasters. [221]
What this result does not establish
The result measures performance in chess under the match rules. The computer and its preparation were designed around that task; the match did not test conversation or everyday problem-solving. [221]
The ImageNet paper introduces a large collection of pictures for training and testing image-recognition systems. [222]
What changed, and the limits
What changed
The researchers gather web images, organize them using WordNet, a catalogue of word meanings, and ask people to check the labels. The 2009 paper reports 3.2 million images across 5,247 categories in the collection available at that point. [222]
What this result does not establish
The collection was still being built. Most of the paper's analysis focused on animal and vehicle categories, so those findings do not describe every kind of image or recognition task. [222]
A deep neural network improves image classification
Krizhevsky, Sutskever and Hinton report a winning ImageNet competition entry built with deep neural networks. [223]
What changed, and the limits
What changed
Their approach trains on labeled pictures using graphics processors, or GPUs, to speed up the calculations. It learns image features through several layers instead of relying only on features specified by hand. [223]
What this result does not establish
The competition asks for labels from a fixed set of image categories. The authors also note that even ImageNet cannot capture the full variety of object recognition in the world. [223]
DeepMind's AlphaGo wins four of five games against Go player Lee Sedol. [224]
What changed, and the limits
What changed
DeepMind describes a system combining neural networks with a search through possible moves. It first learns from expert games, then improves by playing against versions of itself and learning from the results. [224]
What this result does not establish
The result concerns the board game Go. Its training examples, possible moves and win-or-lose feedback are specific to that setting; the match did not test an unrestricted range of human tasks. [224]
The Transformer offers a different way to process text
Attention Is All You Need introduces the Transformer, a neural-network design that relates different parts of a text. [70]
What changed, and the limits
What changed
Its attention calculations let the model use information from other positions in a sequence. The authors report improved translation results and a design that allows more training calculations to happen in parallel. [70]
What this result does not establish
The original experiments cover English–German and English–French translation, plus a task that identifies sentence structure. They establish results for those tasks and conditions, rather than for every later Transformer application. [70]
Joy Buolamwini and Timnit Gebru find different error rates across skin-type and gender groups in three commercial face-analysis systems. [215]
What changed, and the limits
What changed
The Gender Shades study tests systems that assign gender labels to photographs. Using a dataset balanced by gender and skin type, the researchers find the most errors for darker-skinned women. The study examines performance separately for different groups. [215]
What this result does not establish
These results concern three systems and the dataset used in the 2018 study. They do not measure every face-analysis system, later versions or every aspect of fairness. [215]
BERT learns from text before adapting to particular tasks
BERT uses text on both sides of a word to learn numerical representations of text that can be adapted for language tasks. [225]
What changed, and the limits
What changed
The authors train a Transformer-based model on text, then adjust it for tasks such as answering questions. They report improved results across eleven language-processing tasks. [225]
What this result does not establish
The reported task results use additional training for each task. BERT's ability to learn useful text representations is distinct from a ready-to-use conversational assistant. [225]
RAG combines finding passages with writing an answer
A research paper combines a language generator with a system that retrieves relevant Wikipedia passages. [99]
What changed, and the limits
What changed
The authors call the approach retrieval-augmented generation, or RAG. Their model uses both what it learned during training and passages found in a separate index when producing an answer. [99]
What this result does not establish
The paper reports improvements on selected tests, not perfect factuality. Its authors warn that Wikipedia and other external sources can contain errors and bias. [99]
GPT-3 performs tasks from instructions and examples
OpenAI's GPT-3 paper studies how a large language model can attempt new tasks using examples placed in its input. [226]
What changed, and the limits
What changed
The researchers give the model instructions and a few demonstrations in text, without changing its learned settings for each task. They report results on translation, question answering and other language tests. [226]
What this result does not establish
The paper identifies tasks where this approach still struggles. It also checks overlap between web training data and test material, flagging some results while finding little effect on most of the tests examined. [226]
OpenAI's InstructGPT paper describes using people's example answers and preferences to tune a language model. [102]
What changed, and the limits
What changed
People first demonstrate desired responses, then rank alternative model outputs. The researchers use that feedback in further training and report that their evaluators preferred InstructGPT's answers to those from the underlying GPT-3 model. [102]
What this result does not establish
The preference result applies to the evaluators and prompts used in the study. The paper says InstructGPT still makes simple mistakes; being preferred by a reviewer does not establish that an answer is correct. [102]
Stanford's HELM project evaluates language models across several uses and measures, including fairness and efficiency. [101]
What changed, and the limits
What changed
The researchers use common test conditions and examine accuracy alongside reliability, bias, harmful language and other measures. They publish prompts and outputs so others can inspect the results. [101]
What this result does not establish
HELM's authors explicitly identify gaps in what they test, including underrepresented English dialects and some measures of trustworthiness. A broad evaluation still has boundaries. [101]
OpenAI introduces ChatGPT as a research preview that people can question through a conversation. [227]
What changed, and the limits
What changed
The announcement describes a model trained with example conversations and human feedback. Follow-up questions become part of the interaction, rather than each request being treated only as an isolated text completion. [227]
What this result does not establish
OpenAI's launch announcement warns about convincing but incorrect answers, sensitivity to wording and excessive verbosity. These are limitations reported for the version introduced in 2022. [227]
OpenAI's GPT-4 technical report describes a model that accepts pictures and text and produces text. [228]
What changed, and the limits
What changed
The report presents tests of this multimodal model on language tasks, images and exams. It describes next-token prediction followed by further training with human feedback. [228]
What this result does not establish
OpenAI says the model remains less capable than people in many real-world situations. The report also withholds details such as model size and training-data construction, limiting what readers can independently reconstruct. [228]
AlphaFold 3 predicts structures of interacting molecules
AlphaFold 3's authors describe a model for predicting how proteins and other biological molecules fit together. [229]
What changed, and the limits
What changed
The paper covers combinations that include proteins, DNA or RNA, and small molecules. The researchers report improved structure predictions over several specialized tools on the comparisons they test. [229]
What this result does not establish
The paper documents mistakes such as overlapping atoms and incorrect molecular geometry. It also describes limits in representing the different shapes a molecule can take. [229]
DeepSeek describes training aimed at reasoning tasks
DeepSeek's R1 paper reports using reward-based training to improve performance on reasoning tests. [230]
What changed, and the limits
What changed
The January 2025 paper describes R1-Zero and R1, with different training setups. R1 combines example data and reinforcement learning; the team reports results on mathematics, coding and other tasks, as well as training smaller models from R1-generated examples. [230]
What this result does not establish
These are the team's reported results. The same paper says R1 trails its earlier V3 model on some tool-use and multi-turn tasks, and can mix languages or respond poorly to certain prompts. [230]
Vernor Vinge's essay argues that greater-than-human intelligence could transform technological change and the limits of prediction. [67]
2002–2003 · Existential risk and advanced AI
Bostrom develops an existential-risk taxonomy and discusses advanced AI goals, including a paperclip example, in separate papers. [65][59]
2006–2009 · From Overcoming Bias to LessWrong
LessWrong's own history traces the community from the Overcoming Bias blog in 2006 to a separate site in 2009, seeded with the Sequences. [49]
2010 · The basilisk controversy
A LessWrong user named Roko posts the disputed thought experiment. The community retrospective describes rejection of the argument and a subsequent discussion ban. [52]
2011 · Effective altruism gets its name
The EA introduction dates the term's coinage to naming the Centre for Effective Altruism in Oxford. The project covers several causes, not just AI. [24]
2016–2019 · More concrete safety questions
Concrete Problems in AI Safety examines accident risks in 2016. The 2019 learned-optimization paper introduces mesa-optimization as a research question. [62][61]
2022 · e/acc lays out its outlook
The e/acc newsletter publishes its growth-focused self-description and reposts early principles opposing centralized deceleration. [25][26]
2023 · Pause and acceleration proposals
FLI's open letter requests a six-month frontier pause. Later that year, Buterin's techno-optimism essay introduces d/acc as a selective, defensive approach. [69][28]
2025–2026 · Proposals keep evolving
Buterin revisits d/acc in January 2025. PauseAI's April 2026 proposal sets out international safety and democratic-control conditions. [29][31]
Connections between ideas
lesswrong ↔ ea (interpretation)
One is a discussion community; the other is a project for finding effective ways to help. Their shared interest in reasoning does not make them the same institution or philosophy. [50][24]
ea ↔ longtermism (synthesis)
EA spans causes such as health, animals and catastrophic risks. Longtermism adds a particular emphasis on future generations. [24][30]
basilisk ↔ decision-theory (interpretation)
The basilisk tries to turn unusual decision-theory assumptions into an argument about future threats. The formal theories do not themselves establish that scenario. [53][54]
eacc ↔ dacc (synthesis)
Both use the language of building and progress. Buterin's d/acc makes defensive benefits and distributed power explicit selection criteria. [25][29]
orthogonality ↔ convergence (synthesis)
Different final goals can coexist with similar useful intermediate goals. These are complementary theses rather than contradictory claims. [60]
ai-ethics ↔ xrisk (interpretation)
Rights and present harms and long-run catastrophe are different dimensions of concern; this map cannot compress all of them into one risk score. [66][65]
training ↔ inference (synthesis)
Training adjusts learned settings. Inference uses the trained model to produce an output. They are different stages of working with a model. [90][87]
prompt ↔ fine-tuning (synthesis)
A prompt supplies instructions and examples as input. Fine-tuning changes learned parameters through additional training. [92]
rag ↔ hallucination (synthesis)
Retrieval supplies material that can support an answer. Check the generated claims against that material, since supplying sources does not by itself verify the result. [99][83]
tools ↔ agents (synthesis)
Tools provide operations an application can carry out. An agent can use a model to choose and repeat steps involving those tools, subject to the application's access controls. [97][81]
alignment ↔ safety (synthesis)
Alignment concerns intended behavior and goals. Safety also requires examining harms, failures and the environment in which a system operates. [102][62][77]
chatbot ↔ llm (synthesis)
The chatbot is the application you interact with. A language model can generate its replies, while the surrounding software manages context and tools. [109][97]
ai-governance ↔ eu-ai-act (synthesis)
Governance includes many ways of assigning responsibility and oversight. The EU AI Act is one legal framework within that wider topic. [106][114]
understanding ↔ grounding (interpretation)
The form-versus-meaning debate asks how language connects to what it refers to. [120]
anthropomorphism ↔ agents (interpretation)
An agent can have permission to use tools. Human-like wording does not establish human feelings or transfer responsibility away from its operators. [123]
evaluation ↔ benchmark (synthesis)
Evaluation is the wider process of checking a system. A benchmark is one reusable test within that process. [101]
reasoning-models ↔ chain-of-thought (synthesis)
A reasoning model may generate intermediate steps. Those steps can help with a task without reliably explaining how the answer was produced. [104][154][124]
Reading routes
Understand the debate
A short route through the major positions and ethical ideas. Editorial reading suggestion.
What we read: The publisher article, including its Key Facts and Trump Dismisses AI Concerns sections. Original social-media text was not independently retrieved.
No archived copy verified · Archive check 2026-09-15
The Wayback availability lookup returned HTTP 429 (rate limited). No capture was verified; this does not show that no archived copy exists.
Original posts mentioned in this source
These links help trace the reporting. Their original text was not independently retrieved. They are not additional verified sources.
Original not retrieved · Access checked 2026-09-15
The linked Truth Social page exposed a JavaScript prompt without the post text. The reporting remains the material read.
No archived copy verified · Archive check 2026-09-15
The Wayback availability lookup returned HTTP 429 (rate limited). No capture was verified; this does not show that no archived copy exists.
Read the publisher's accessible article, including Key Facts and Trump Dismisses AI Concerns; retrieval redirected to tollbit.forbes.com. It links September 14 Truth Social posts 117270591511950591 and 117269745153543631. Their original pages returned a JavaScript shell or browser security check; original post text was not independently retrieved. Vance's statements and reporters' claims about other actors are not assigned to Trump. The title differs between the indexed and retrieved versions.
What we read: AFP's report of the statements, not the original social-media posts.
No archived copy verified · Archive check 2026-09-15
The Wayback availability lookup returned HTTP 429 (rate limited). No capture was verified; this does not show that no archived copy exists.
Read AFP's own September 14 article, displayed at 23:36 with no time zone specified. Used for attributed statements about catastrophic AI risk and guardrails, not as direct verification of Truth Social posts or independent support for all contextual claims in the article.
Read the signed order, sections 1–5, directly on the White House site. Supports its stated policy and directions; does not independently establish completed implementation, security effectiveness, current legal status, or Trump's private beliefs. The voluntary model-evaluation process is distinct from an industry-wide training pause.
Read scaling discussion and the 00:36:46–00:59:56 Grok/alignment section of the host transcript. These are Musk's claims and plans, not verified engineering results or proof of alignment.
Read 01:23:13–01:29:39 on scaling, delayed open sourcing and Musk's account of safety disagreements. Historical self-report; claims about other people's motives are not adopted.
thiel-tyler-2024 · First-hand source Peter Thiel on Political Theology (Ep. 210) Conversations with Tyler / Mercatus Center · Published 2024-04-17 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Host transcript, recorded February 21, 2024. Read existential-risk discussion and audience questions on stopping AI, open source and substitution for workers. Political judgments are attributed to Thiel.
Read the host transcript, including the Scylla and Charybdis discussion. The page identifies recording on October 8, 2024, but no reliable publication date. Recording date is not substituted for publication. His theological framing is speculative.
Read the edited interview's AI discussion: growth, concentrated returns, labor substitution and skepticism of the label. The interview establishes his stated outlook, not its economic predictions.
thiel-spectator-2026 · Reporting or indirect copy Can Peter Thiel stop the Antichrist? The Spectator / Lara Brown and John Power · Published 2026-02-07 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the reporters' account of Thiel's Cambridge talk, in the February 7 issue. Supports continuity of his political-theological concern; no complete talk transcript was retrieved.
amodei-pacing · First-hand source We Must Pace the Frontier Dario Amodei · Published 2026-09 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
The essay itself specifies the month. The September 12 publication day is corroborated by the September 14 press roundup. A proposal and a commitment are not independent verification of implementation.
What we read: The article's reporting on Altman, Hassabis, Musk and LeCun, including links to their posts. The original post text was not independently retrieved.
No archived copy verified · Archive check 2026-09-15
The Wayback availability lookup returned HTTP 429 (rate limited). No capture was verified; this does not show that no archived copy exists.
Original posts mentioned in this source
These links help trace the reporting. Their original text was not independently retrieved. They are not additional verified sources.
The HTML article describes catastrophic-risk domains and release safeguards. No catastrophe probability is inferred.
andreessen-manifesto · First-hand source The Techno-Optimist Manifesto Andreessen Horowitz / Marc Andreessen · Published 2023-10-16 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the manifesto and footer. The footer explicitly attributes posts to individual personnel and excludes the views of a16z Capital Management and affiliates. Historical personal advocacy, not institutional evidence.
What we read: The affiliated announcement of the authors' argument, not the book itself.
No archived copy verified · Archive check 2026-09-15
The Wayback availability lookup returned HTTP 429 (rate limited). No capture was verified; this does not show that no archived copy exists.
A publisher/affiliated announcement of the authors' argument, not an empirical finding that extinction is certain.
miri-position · First-hand source Our view Machine Intelligence Research Institute · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Undated live institutional position checked on the review date. Used as context, not as a substitute for individual authorship.
pauseai-proposal · First-hand source PauseAI Proposal PauseAI · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
A proposed international pause. The page is undated; checkedOn is not a publication date.
superintelligence-statement · First-hand source Statement on Superintelligence Future of Life Institute · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
The statement text was accessible. The dynamic signatory list was not exposed in the retrieved text; Bengio's signature is corroborated separately.
What we read: The published interview transcript. Transcription errors are possible.
No archived copy verified · Archive check 2026-09-15
The Wayback availability lookup returned HTTP 429 (rate limited). No capture was verified; this does not show that no archived copy exists.
First-person interview transcript. The host notes that transcription errors are possible. Current opposition to the pacing essay is reported separately.
ea-definition · First-hand source What is effective altruism? Effective Altruism · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Definition of a philosophy/community, not a placement on the AI chart.
eacc-definition · First-hand source what the f* is e/acc e/acc newsletter · Published 2022-12-26 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Movement self-description. Participants need not share a single risk estimate.
eacc-tenets · First-hand source Notes on e/acc principles and tenets e/acc newsletter / BasedBeffJezos and bayeslord · Published 2022-10-31 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Date belongs to the repost. Read as a movement's own philosophical claims, not independent scientific validation. Includes adversarial usage of decel.
Read the AI doomers and effective altruism sections as examples of contested terminology. An advocate's description of opponents is not a neutral definition of their views. Transcript may contain errors.
dacc-original · First-hand source My techno-optimism Vitalik Buterin · Published 2023-11-27 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read via the author's eth.limo site because vitalik.ca did not load. Introduces defensive, differential and decentralization-focused acceleration. Used for ideas, not a new placement of the author.
dacc-update · First-hand source d/acc: one year later Vitalik Buterin · Published 2025-01-05 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the definition and distinctions: defensive development combined with distributed and democratic control. Hypothetical future examples are not observed outcomes.
longtermism-definition · First-hand source Longtermism William MacAskill · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
An advocate's account of the moral importance of future people. The page is undated; the book's publication date is not used as the page date.
What we read: The April 2026 proposal, including its pause, training-approval and democratic-control conditions.
No archived copy verified · Archive check 2026-09-15
The Wayback availability lookup returned HTTP 429 (rate limited). No capture was verified; this does not show that no archived copy exists.
Read the April 5th, 2026 version, including its global pause, training approvals and democratic-control conditions. Used for the movement's published proposal and terminology; its predictions are attributed advocacy.
cais-risk · First-hand source Statement on AI Extinction Risk Center for AI Safety · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the statement after following the safe.ai redirect. Establishes advocacy for treating extinction risk as a priority, not certainty of catastrophe or a common numerical probability. No new signatory claims are inferred.
Read introduction, risk-report requirements, external-review provisions and Appendix A. July 8 is the stated effective date. Separates company commitments from industry recommendations; not an implementation audit.
bengio-lawzero-2025 · First-hand source Introducing LawZero Yoshua Bengio · Published 2025-06-03 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the author's announcement, including his loss-of-control concerns. Describes his argument and research goals, not proof that the proposed approach is safe.
extreme-ai-risks-2024 · First-hand source Managing extreme AI risks amid rapid progress (version 3) Yoshua Bengio, Geoffrey Hinton, Stuart Russell and coauthors / arXiv · Published 2023-10-26 · Updated 2024-05-22 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read authors, societal-scale risks and governance/mitigation sections of version 3. Coauthorship supports a shared proposal, not identical personal beliefs or a present-day forecast.
Read executive summary and control-risk discussion. The document is dated December 6 and draws on earlier July testimony; the July date is not this document's date.
Read the interview publisher's summary and attributed control-risk quotation. The full audio/transcript was not reviewed; this record is reporting, not direct verification of every interview claim.
Read the journal abstract and publication metadata only; full-text retrieval failed. Used for the authors' recommendation to study defined tasks. Their wider historical thesis is not adopted as an atlas finding.
Read the author's slides, especially the opening and slides 25–33 on scrutiny, rights and teaching. April 16 is the talk date printed on the slides. Does not establish a frontier-training policy.
Read all three pages of the briefing, presented December 19, 2024. Its footnote explicitly says the views are individual and do not represent affiliated organizations.
whittaker-ndss-2024 · First-hand source AI, Encryption, and the Sins of the 90s Meredith Whittaker / Signal · Published 2024-02-27 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the keynote's AI, corporate-surveillance and concluding privacy arguments, pages 1–3 and 10–14. Date is the speech date. Used as her argument, not as proof that all AI systems have one business model.
eu-ai-office-overview · First-hand source European AI Office European Commission · Published undated · Updated 2026-09-08 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the Office's mandate, enforcement tasks and innovation role. The page gives a last-update date, not an original publication date. Describes an institutional mandate, not a catastrophe probability.
uk-aisi-agenda · First-hand source AISI Research Agenda UK AI Security Institute · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the research overview and Autonomous Systems section on catastrophic harm and loss of control. No reliable publication date was exposed. Research priorities do not by themselves establish a preferred development pace.
Read the announcement of the name change and its security-research focus. A government growth agenda is not automatically the institute's own frontier-pacing position.
deepseek-r1-release · First-hand source DeepSeek-R1 Release DeepSeek · Published 2025-01-20 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the release announcement and model-access descriptions. Product availability and benchmark claims do not establish an institutional stance on catastrophic AI risk.
Read the archived speech transcript, including its closing acknowledgment of safety concerns. Primary speech text hosted by a university archive; claims about competitors or regulations are attributed arguments.
altman-governance-2023 · First-hand source Governance of superintelligence Sam Altman, Greg Brockman and Ilya Sutskever / OpenAI · Published 2023-05-22 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the coauthored article on frontier growth limits, existential risk and lower-capability exemptions. Historical proposal, not evidence that a limit was implemented.
Read the publisher's episode description only. It describes advocacy for AI growth and concern about policy barriers; no full transcript was exposed. June 29 is the podcast publication date, not the June 25 event date.
lw-history · First-hand source A Brief History of LessWrong LessWrong / Ruby · Published 2019-06-01 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the site's own retrospective: Overcoming Bias in 2006, LessWrong in 2009, and the Sequences. An internal community account, not an independent history.
lw-welcome · First-hand source Welcome to LessWrong! LessWrong · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the site's introduction and intended reasoning practice. Self-description does not certify participants' accuracy or common beliefs.
lw-rationality · First-hand source What Do We Mean By Rationality? LessWrong / Eliezer Yudkowsky · Published 2009-03-16 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the epistemic/instrumental distinction and discussion of Bayesian reasoning. A proposed practice, not a verified trait of participants.
basilisk-history · First-hand source Roko's Basilisk LessWrong community wiki · Published undated · Updated 2022-11-30 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the community retrospective, objections and moderation account. Original deleted post not independently retrieved. Historical claims are attributed to this account.
Read the community response and quoted objections in the initial retrieval; a later fetch timed out. Used for disputed assumptions, not proof about hypothetical agents. Exact publication date was not independently established.
tdt-paper · First-hand source Timeless Decision Theory Eliezer Yudkowsky / MIRI · Published 2010 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read abstract and opening Newcomb discussion. Year follows the paper's citation. Proposed framework and formal examples, not an experimentally settled rule.
Read the authors' abstract and MIRI's October 22, 2017 introduction. Advantages are claimed within specified decision problems.
simulation-paper · First-hand source Are You Living in a Computer Simulation? Nick Bostrom / Philosophical Quarterly · Published 2003 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read abstract and introduction. Publication year differs from the first draft in 2001. Conditional philosophical argument, not a finding that our world is simulated.
Read abstract, Table 1 and autonomy discussion. Proposed distinctions between breadth, performance and autonomy; no current model classification adopted.
Read abstract and formulations of orthogonality and instrumental convergence. Theoretical theses with qualifications, not measured outcomes for a learning system.
Read abstract and revision history introducing mesa-optimization and the relation between learned and training objectives.
concrete-safety · First-hand source Concrete Problems in AI Safety Dario Amodei and coauthors / arXiv · Published 2016-06-21 · Updated 2016-07-25 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the abstract's five accident-risk problems including reward hacking and distributional shift. Does not give a general catastrophe probability.
goodhart-paper · First-hand source Categorizing Variants of Goodhart's Law David Manheim and Scott Garrabrant / arXiv · Published 2018-03-13 · Updated 2019-02-24 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read abstract distinguishing overoptimization mechanisms. A taxonomy, not a claim that every metric fails.
transhumanism-declaration · First-hand source The Transhumanist Declaration Humanity+ · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the declaration's commitments to enhancement, risk reduction and choice. Live page publication date unspecified.
Read overview of rights, dignity, fairness and oversight. Recommendation adoption date is not assigned to the live page.
vinge-singularity · First-hand source Technological Singularity Vernor Vinge / Carnegie Mellon University archive · Published 1993 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the author's essay via the academic archive after NASA retrieval failed. Speculative scenarios and historical forecasts, not a demonstrated trajectory.
utilitarianism-intro · First-hand source Introduction to Utilitarianism Utilitarianism.net · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read an advocate-authored philosophical introduction. Defines the view; does not establish its correctness or universal adoption among effective altruists.
Read the six-month request concerning systems more powerful than GPT-4. A historical advocacy milestone, not evidence of an implemented halt.
transformer-paper · First-hand source Attention Is All You Need Ashish Vaswani and coauthors / arXiv · Published 2017-06-12 · Updated 2023-08-02 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read abstract, model architecture, learned embeddings and next-token probabilities in the HTML paper; publication and revision dates checked against the arXiv abstract page. This is the original Transformer architecture, not a claim that every current LLM has its exact structure.
hf-generation · First-hand source Generation strategies Hugging Face Transformers documentation · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read greedy search, multinomial sampling and the custom generation loop separating model logits from token selection. Used to explain decoding choices; library defaults do not establish the behavior of every hosted chatbot. The live page does not establish a publication date.
pytorch-reproducibility · First-hand source Reproducibility PyTorch documentation · Published 2026-05-14 · Updated 2026-05-14 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the versioned page reached from the stable documentation: cross-release/platform limits, random seeds and deterministic algorithms. Dates follow the page's displayed Created On and Last Updated On fields, not the historical first publication of PyTorch's reproducibility guidance. These are framework constraints, not measurements of a particular chatbot service.
Read introduction, method overview and limitations including reconstruction errors, graph complexity, global circuits and mechanistic faithfulness. The authors' replacement-model analyses reveal selected mechanisms; they do not provide a complete explanation of all behavior. Later attention-tracing work is cited alongside this paper to avoid treating its missing-attention limitation as a permanent field-wide result.
Direct browser-tool retrieval failed; fetched the original publisher HTML successfully and read the introduction, case-study summaries, QK-attribution method, inhibitory-effect limitation and graph-construction tradeoffs. Extends earlier attribution graphs to attention; results are selected studies with open questions, not a complete model explanation.
Read LLM definition, code-generation workflow, generated-test caveats, inaccurate-code limitations and review/testing guidance. Vendor documentation establishes intended use and acknowledged limitations, not independent accuracy rates. Live page publication date unspecified.
weather-uncertainty · First-hand source Quantifying forecast uncertainty European Centre for Medium-Range Weather Forecasts · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the forecast-uncertainty explanation, initial-condition uncertainty and numerical-model approximations. Cited only for the weather side of an explicitly editorial analogy; it supplies no evidence that LLMs are meteorological or chaotic systems. Live page publication date unspecified.
ncsc-secure-ai · First-hand source Guidelines for secure AI system development UK National Cyber Security Centre and international partners · Published 2023-11-27 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read executive summary and lifecycle structure. Guidance addresses complete AI systems and recommends security throughout design, development, deployment and operation. Recommendations are not evidence of any organization's implementation.
Read threat modelling, task suitability, model selection, restricted actions and least privilege. Date follows the containing guideline publication. Used for design principles, not a certificate that any configuration is safe.
Read monitoring of behavior and inputs, privacy-aware logging, update evaluation and lessons learned. Date follows the containing guideline publication. Describes operational monitoring, not complete explanation of learned weights.
ncsc-agentic-risk · First-hand source Managing the cyber risk of agentic AI UK National Cyber Security Centre · Published 2026-08-20 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read autonomy, model safeguards, oversight, sandbox boundaries, network and credential restrictions, observability and emergency response. The publisher labels this interim practical advice based on its research; formal guidance may supersede it.
owasp-excessive-agency · First-hand source LLM06:2025 Excessive Agency OWASP Gen AI Security Project · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read agency definition, excessive functionality/permissions/autonomy, external authorization, approvals and monitoring limits. The 2025 label identifies the edition; the page does not establish its original publication date.
owasp-prompt-injection · First-hand source LLM01:2025 Prompt Injection OWASP Gen AI Security Project · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read direct and indirect injection definitions, contextual impact and mitigations including validation, privilege limits and adversarial testing. The 2025 label is an edition, not a verified publication date. Does not establish that a mitigation eliminates every attack.
Read introduction, section 2.2 on confabulation, and selected MEASURE actions 2.3, 2.5, 2.6, 2.7 and 2.9 concerning evaluation evidence, generalization, citations, generated-code review and safeguards. A voluntary risk-management profile; no claim that all 64 pages or every referenced study was reviewed.
Read abstract, rationale, research questions, limitations and conclusion. A research position paper: reasoning traces may add monitoring value while remaining incomplete and potentially fragile. Authors' views are not necessarily their institutions' positions; cited experiments were not all independently reviewed.
Read deployment approach, reasoning/tool-trace monitoring, asynchronous alerts, limitations and proposed control evaluations. A dated first-party account, not an independent audit or a claim about current coverage. The authors explicitly cannot establish a real-world missed-event rate from employee escalations alone.
otel-observability · First-hand source Observability primer OpenTelemetry documentation · Published undated · Updated 2026-04-23 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read observability, telemetry, logs and distributed-trace definitions. Updated date follows the displayed documentation modification, which references a spelling-related commit rather than a new research result. Used for software terminology, not complete access to model internals.
google-ml-glossary · First-hand source Machine Learning Glossary Google for Developers · Published undated · Updated 2026-04-10 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the artificial intelligence, deep model, context window, inference and chat entries. Used for terminology, not product performance claims. Updated date follows the earlier displayed page date; initial publication is unspecified. The chat entry was reread during the same-day beginner-content review. Also read the generalization and compute entries.
google-ml-intro · First-hand source What is Machine Learning? Google for Developers · Published undated · Updated 2026-01-27 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the introduction, model definition, supervised and unsupervised learning, and generative AI sections. Examples illustrate categories rather than measured accuracy. Updated date follows the page; initial publication is unspecified.
google-neural-layers · First-hand source Neural networks: Nodes and hidden layers Google for Developers · Published undated · Updated 2025-12-03 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the explanations of connected layers, numerical weights and biases, and calculations. Did not run the embedded exercises. Used for the mathematical structure, not a claim that an artificial network reproduces a human brain.
google-gradient-descent · First-hand source Linear regression: Gradient descent Google for Developers · Published undated · Updated 2026-02-03 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the iterative prediction, loss and parameter-update explanation. The page's guarantees for convex linear regression are not extended here to neural-network training. Updated date follows the page; initial publication is unspecified.
google-llm-intro · First-hand source LLMs: What's a large language model? Google for Developers · Published undated · Updated 2026-01-02 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read token prediction, encoder-only and decoder-only variants, and self-attention. Used for architecture and terminology; broad performance comparisons and claims about all LLMs on the teaching page are not adopted.
Read fine-tuning and prompt engineering, including the distinction between parameter updates and examples supplied as input. Used to distinguish these processes, without adopting general claims that fine-tuning is always necessary or improves every task.
Read numerical representations, distance as relative similarity, task dependence and the limits of human-readable dimensions. The food diagrams are teaching examples, not measurements reused in this atlas.
hf-tokenizers · First-hand source Tokenizers Hugging Face LLM Course · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read word, character and subword tokenization, encoding to numerical IDs and decoding. Used for the fact that token boundaries depend on the tokenizer; no fixed words-to-tokens conversion is assumed. Live page publication date unspecified.
hf-models · First-hand source Models Hugging Face LLM Course · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read architecture, weights, checkpoints, loading and saving. Used to distinguish a model's structure and learned values from the application around it. Example code was read, not executed; live page publication date unspecified.
hf-text-generation · First-hand source Text generation Hugging Face Transformers documentation · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read next-token generation, generation settings, temperature, sampling and prompt-format sections. Library options illustrate the process; defaults and suggested temperatures are not treated as universal chatbot behavior. Live page publication date unspecified.
hf-tool-use · First-hand source Tool use Hugging Face Transformers documentation · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read tool descriptions, model-generated call requests, application execution and returning results to the chat. Example functions were not run. Used for the separation between requesting and executing an action; live page publication date unspecified.
hf-multimodal · First-hand source Multimodal chat templates Hugging Face Transformers documentation · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read mixed text/image/audio/video inputs, preprocessing and model-specific video support. Used to explain input types; examples are not evidence of universal capabilities or accuracy. Live page publication date unspecified.
Read abstract and version history, plus version 4's results sections 4.3/4.4 and Broader Impact discussion during the history review. External passages can contain errors or bias. Used for the original approach; benchmark results do not establish accuracy for every system now called RAG.
Read the abstract and version history: position-dependent performance in multi-document question answering and key-value retrieval. Used to distinguish accepted context length from effective use of content, not to assign the same weakness to every current model.
helm-paper · First-hand source Holistic Evaluation of Language Models Percy Liang and coauthors / Stanford CRFM, arXiv · Published 2022-11-16 · Updated 2023-10-01 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the abstract and version history, including multiple use cases and metrics, standardized comparisons and acknowledged coverage gaps. Used for evaluation principles; historical model scores are not presented as current rankings.
Read abstract, section 3.1's demonstrations/comparisons/reward-model procedure and section 5.3's limitations. Human preference judgments and improved results on the authors' tasks do not establish universal alignment or safety. Publication date checked on the arXiv abstract page.
Read abstract and metadata describing model critiques, revisions and AI preference feedback guided by human-written principles. Used as an example of alignment methods; the paper's claims do not certify every output as harmless.
Read abstract, introduction and chain-of-thought/inference-time-compute discussion. Dates checked against the abstract page's version history. Used for training and extended reasoning examples; the authors' benchmark results are not a universal ranking or an explanation of all model internals.
osi-ai-definition · First-hand source The Open Source AI Definition – 1.0 Open Source Initiative · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Browser-tool fetch returned 403; fetched and read the original publisher HTML directly. Read the four freedoms and requirements for data information, code and parameters. This is OSI's definition, not a claim that every publisher uses the label identically. Live page publication date unspecified.
oecd-ai-principles · First-hand source AI principles Organisation for Economic Co-operation and Development · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read overview, human rights, transparency, safety, accountability and governance-policy recommendations. The principles were adopted in 2019 and updated in 2024; those dates are not assigned as the live page's publication date. Recommendations do not establish implementation or legal compliance.
Read publication metadata, the systemic/statistical/human bias taxonomy, contextual evaluation discussion and conclusion. Used to explain sources and assessment of bias; this does not establish a bias finding for any particular model or actor.
claude-introduction · First-hand source Introducing Claude Anthropic · Published 2023-03-14 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the launch introduction, chat/API distinction and model variants. Historical product identification only; customer testimonials and reliability claims are not treated as independent evidence or current specifications.
gemini-overview · First-hand source What is Gemini and how it works Google · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the introduction describing the Gemini app as an interface to Google's multimodal language models. Used for naming and app/model distinction, without endorsing performance claims. No exact publication date established.
grok-introduction · First-hand source Announcing Grok xAI (now hosted by SpaceXAI) · Published 2023-11-03 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the original introduction, Grok-1 model description and research section acknowledging false answers despite search access. Historical identification only; no launch specifications or benchmark rankings are presented as current.
llama-model-family · First-hand source The Llama 3 Herd of Models Llama team / Meta · Published 2024-07-23 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read Meta's abstract and displayed publication date, including separate pretrained and post-trained releases. Used to identify a model family, not endorse benchmark comparisons or describe the latest release.
Read the introduction, selected provenance/detection passages and Appendix D's audio/video examples. Date checked on NIST's landing page. No detector certified here.
merriam-webster-slop · First-hand source 2025 Word of the Year: Slop Merriam-Webster · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the dictionary publisher's definition and examples in its 2025 selection. The year identifies the selection; an exact article publication date was not established. Used for language usage, not a quality finding about particular content.
eu-ai-act-overview · First-hand source AI Act European Commission · Published undated · Updated 2026-08-03 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read risk categories, transparency, general-purpose model rules, enforcement and application timeline. Updated date is displayed on the page. This overview does not determine any particular application's legal duties; deadlines are deliberately not generalized here.
Read the race framing, policy pillars and stated proposals. A government announcement and example of political framing; its promised benefits and implementation are not independently established.
Read abstract, executive summary and sections 1.1–1.2, including disputed working assumptions and limits of behavior-based assessment. Version dates checked on arXiv. A proposed framework, not a consciousness test or a 2026 assessment of all models.
Read the task-based method, exposure categories, employment interpretation and adoption limitations. Estimates describe potential exposure, not observed job losses or a prediction about an individual worker.
chatgpt-memory · First-hand source Memory FAQ OpenAI Help Center · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read memory, context, controls and the legacy saved-memory explanation. Used as one provider's example; not generalized to all chatbots. Relative update wording was not converted to an invented exact date.
Read individual/business distinctions, training controls and feedback exceptions. Used to distinguish later training from conversation context, not to provide a complete privacy checklist. The page displays a relative update time; exact publication date is unspecified.
Read the paper and abstract. A position argument about meaning learned from form alone, not an experimental verdict on every multimodal or tool-connected system.
Read the full version 2 manuscript and arXiv submission history. First submitted October 14, 2022; the linked version 2 is dated October 27, 2022. Surveys competing accounts, benchmark shortcuts and open questions. Its model examples describe that period, not a current capability ranking.
Read the abstract and submission history. Reports board-state representations and interventions in a synthetic Othello task. This source alone does not establish human-like comprehension.
understanding-anthropomorphism · First-hand source AI Automatons: AI Systems Intended to Imitate Humans Alexandra Olteanu and coauthors / Microsoft Research · Published 2025-03 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the publication summary and accompanying research transcript. Discusses design choices, anthropomorphic cues and possible social effects. Does not prove a universal company motive.
Read methods, findings and limitations. Hint experiments used Claude 3.7 Sonnet and DeepSeek R1 on multiple-choice questions. Results do not cover every model, task or reasoning trace.
glossary-model-algorithm · First-hand source algorithm Paul E. Black / NIST Dictionary of Algorithms and Data Structures · Published undated · Updated 2020-11-09 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the definition and listed algorithm types. The entry was modified on 9 November 2020; the later HTML formatting date is not a content revision.
glossary-model-datasets · First-hand source Datasets: Dividing the original dataset Google for Developers · Published undated · Updated 2025-12-03 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read training, validation and test separation, duplicate examples and repeated test reuse. Course examples illustrate evaluation problems; no model was tested here.
glossary-model-pretraining · First-hand source How do Transformers work? Hugging Face LLM Course · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read self-supervised language modeling and transfer learning. Used for the pretraining distinction, without adopting historical dates, performance comparisons or claims of understanding from this teaching page.
glossary-model-posttraining · First-hand source The Llama 3 Herd of Models Llama Team / Meta, arXiv · Published 2024-07-31 · Updated 2024-08-15 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the two training stages, post-training methods and section 5.4.8 limitations in version 2. Dates follow arXiv submission history. One documented approach, not a universal training recipe.
glossary-model-supervised · First-hand source Supervised Learning Google for Developers · Published undated · Updated 2025-08-25 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read labeled examples, training, evaluation and inference. Used for the training procedure, without treating labels as infallible or adopting the page’s wording about understanding.
glossary-model-clustering · First-hand source What is clustering? Google for Developers · Published undated · Updated 2025-08-25 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read grouping of unlabeled examples and the choice of similarity measure. Clustering is one example of unsupervised learning. No privacy guarantee or natural group boundary is inferred.
Read the predictive-learning section and text masking examples. Used for how targets are formed. The authors’ forecasts about common sense and human-level intelligence are not adopted.
glossary-model-reinforcement · First-hand source Deep Reinforcement Learning David Silver / Google DeepMind · Published 2016-06-17 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the introduction defining reinforcement learning through trial, feedback and long-term rewards. Used for the training setup, without adopting human comparisons or generality claims.
glossary-model-deep-learning · First-hand source Deep Learning Ian Goodfellow, Yoshua Bengio and Aaron Courville · Published 2016 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read layered representations and computational depth in Chapter 1. The book’s citation page supplies the publication year. No agreed layer threshold or intelligence measure is inferred.
glossary-model-mixture-of-experts · First-hand source Mixtral of Experts Albert Q. Jiang and coauthors / arXiv · Published 2024-01-08 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the abstract and introduction describing token routing, expert blocks and active parameters. Publication date checked on arXiv. Mixtral illustrates sparse routing; its design is not assigned to every MoE. Also read section 5, Routing analysis, on whether routing follows subject areas.
glossary-model-distillation · First-hand source Distilling the Knowledge in a Neural Network Geoffrey Hinton, Oriol Vinyals and Jeff Dean / arXiv · Published 2015-03-09 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the introduction and distillation method, including probability targets and imperfect matching. Date checked on arXiv. The historical experiments do not establish performance for current distilled models.
glossary-model-quantization · First-hand source Quantization concepts Hugging Face Transformers documentation · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read lower-precision weights and activations, rounding, accuracy tradeoffs and hardware dependence. No universal speedup or accuracy loss is claimed. Live documentation publication date is unspecified.
glossary-model-overfitting · First-hand source Overfitting Google for Developers · Published undated · Updated 2025-12-03 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read training versus new-data performance, the underfitting comparison and generalization curves. Illustrative curves are not evidence about a particular model.
glossary-model-underfitting · First-hand source Machine Learning Glossary Google for Developers · Published undated · Updated 2026-04-10 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the underfitting entry and its listed causes. Updated date follows the glossary page; initial publication date is unspecified. This is a diagnostic concept, not a finding about any listed model.
glossary-model-loss · First-hand source Linear regression: Loss Google for Developers · Published undated · Updated 2026-01-05 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read error measures and their different treatment of large errors. Used for the idea of a training objective, without presenting squared error as the standard language-model objective.
glossary-eval-calibration · First-hand source On Calibration of Modern Neural Networks Chuan Guo and coauthors / PMLR · Published 2017-08 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the proceedings abstract and publication metadata. The experiments concern image and document classifiers, not the reliability of a chatbot saying it is certain.
glossary-eval-gpt3 · First-hand source Language Models are Few-Shot Learners Tom B. Brown and coauthors / arXiv · Published 2020-05-28 · Updated 2020-07-22 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the abstract, section 2 on zero-shot and few-shot settings, and section 4 on training-data overlap. Version 4 is dated 22 July 2020. The reported results concern GPT-3 and are not current model rankings; detected overlap did not uniformly inflate scores.
Read the abstract and version history. Used for the distinction between uncertainty in observations and uncertainty in the model. Its experiments concern computer vision, not a validated uncertainty measure for every LLM.
glossary-eval-rmf-characteristics · First-hand source AI Risks and Trustworthiness NIST AI Resource Center · Published 2023 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read sections 3.1 to 3.3 of the online AI RMF 1.0 excerpt: validity, reliability, robustness, safety and security. Guidance and definitions are not certification of a particular system.
glossary-eval-adversarial · First-hand source Explaining and Harnessing Adversarial Examples Ian J. Goodfellow, Jonathon Shlens and Christian Szegedy / arXiv · Published 2014-12-20 · Updated 2015-03-20 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the abstract and version history. Supports deliberately modified inputs that cause misclassification and adversarial training as a proposed response. Findings are scoped to the studied models and attacks.
Read the executive summary, attack-stage definitions, and sections 3.2.1 to 3.2.3 on generative-model poisoning and mitigations. Publication date comes from the NIST publication record. No claim is made that one defense stops all attacks.
Read the abstract and submission record, plus the authors' DeepMind research summary. The work studies automated test generation and presents it as one method among several, not an exhaustive safety test.
glossary-eval-induction · First-hand source In-context Learning and Induction Heads Catherine Olsson and coauthors / arXiv · Published 2022-09-24 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the abstract and submission record. The authors report causal evidence in small attention-only models and indirect or correlational evidence for their wider mechanism hypothesis. The glossary does not present that hypothesis as settled for all models.
Read the abstract and version history. The paper reports improvements on selected arithmetic, commonsense and symbolic tasks with intermediate-step examples. It does not establish that generated explanations faithfully reveal internal computation.
glossary-eval-guardrails · First-hand source Guardrail Types NVIDIA NeMo Guardrails documentation · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the descriptions and tables of input, retrieval, dialog, execution and output rails. Used as a concrete implementation example of the broader term. No publication date is displayed; the security FAQ link returned 404 and was not used.
Read the AI RMF 1.0 appendix on human roles, oversight, bias and differing outcomes of human-AI interaction. It describes both possible complementarity and amplified bias; it is guidance rather than a controlled experiment.
glossary-eval-human-errors · First-hand source The impact of AI errors in a human-in-the-loop process Ujué Agudo and coauthors / Cognitive Research: Principles and Implications · Published 2024-01-07 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the abstract, study procedures, results and general discussion. Two simulated judicial-decision experiments used purported AI advice; these are not a field trial of an LLM or a universal estimate of human oversight effectiveness.
Read the abstract and submission record. The study involved 28 pathology experts and reported improved overall performance alongside acceptance of some wrong advice. Abstract-only review; its error rate is not generalized to other users or tasks.
Read the abstract and version history. The paper studies retrieval-based overlap checks and a test-slot guessing method. Its reported scores concern particular models and benchmarks; the glossary does not treat a guessed answer alone as proof of training membership.
Read the abstract and version history. The experiments introduced biasing input features and examined generated explanations in GPT-3.5 and Claude 1.0. Their results are not a verdict on every model or every explanation.
glossary-wide-vision · First-hand source Cloud Vision API documentation Google Cloud · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the overview of image labeling, text extraction and object detection. Used to explain task types, not to claim that one product covers all computer vision or is equally accurate on every image.
Read the task definition and English/German transcription examples, including incorrect words. The examples illustrate the need to check a transcript; they do not establish error rates for current systems.
glossary-wide-tts · First-hand source Audio generation with a pipeline Hugging Face · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the speech-generation definition and examples. Used for the distinction between producing spoken audio, transcribing speech and generating music, without adopting product-quality claims.
glossary-wide-diffusion · First-hand source Denoising Diffusion Probabilistic Models Jonathan Ho, Ajay Jain and Pieter Abbeel / arXiv · Published 2020-06-19 · Updated 2020-12-16 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the abstract, introduction, forward/reverse process and sampling algorithm. The entry explains the denoising approach in this paper, without treating its benchmark results as current performance.
Read the abstract and introduction, including the separation between image compression and diffusion in a learned representation. Used as a concrete example of latent space, not a definition of every representation in AI.
glossary-wide-synthetic-data · First-hand source Welcome to the SDV! DataCebo / Synthetic Data Vault · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the library overview and its generation and evaluation workflow for tabular synthetic data. This is developer documentation; no commercial claim of quality or privacy is treated as independently verified.
Read the abstract and version history. The reported attacks show that passing the studied similarity-based tests does not establish anonymity. The entry does not claim that every synthetic dataset leaks personal information.
Read the abstract and version history. The paper distinguishes a reported cutoff from the effective cutoff for particular resources and topics. No specific current product cutoff is inferred.
glossary-wide-provenance · First-hand source PROV-Overview W3C · Published 2013-04-30 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the abstract and introduction to the PROV family, including entities, people, activities and derivation. Provenance supports assessment of reliability; the document does not make a recorded origin proof of truth.
glossary-wide-model-card · First-hand source Model Cards for Model Reporting Margaret Mitchell and colleagues / arXiv · Published 2018-10-05 · Updated 2019-01-14 · Material last read 2026-09-15
What we read: The abstract and version history, not the full paper.
No archive check recorded.
Read the abstract and version history, including intended uses, evaluation conditions and reporting across groups. This is a documentation proposal, not a certification that a model is safe or fair.
Read the system-card proposal, its distinction from model cards and its limitations section. Meta describes its own approach; publication of a card is not an independent audit of the system.
glossary-wide-gpu · First-hand source 1.1. Introduction NVIDIA / CUDA Programming Guide · Published undated · Updated 2026-09-09 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read sections 1.1.1 and 1.1.2 on GPU origins and parallel computation. The entry uses the architectural distinction without repeating vendor performance or energy comparisons.
Read the performance-metrics section distinguishing time to first token, subsequent token timing and throughput. Specific benchmark targets are not presented as universal user requirements.
Read the platform overview and deployment workflow for running models on devices. The entry describes the location of computation, without adopting blanket vendor claims about privacy, speed or capability.
glossary-balance-ea-crary · First-hand source Against ‘Effective Altruism’ Alice Crary / Radical Philosophy · Published 2021 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the institutional and philosophical critique and discussion of EA replies. These are Crary’s arguments, not findings about every participant or charity. Publication precision is the issue year.
glossary-balance-ea-donor-power · First-hand source Response to Effective Altruism Emma Saunders-Hastings / Boston Review · Published 2015-07-01 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the complete response on donor power and recipient choice. Historical charity recommendations are not repeated as current advice. The argument warns of a governance problem rather than rejecting every donation.
Read the definitions and replies on utilitarianism, uncertainty, systemic change and competing priorities. This is the movement website’s account of its principles, not an independent evaluation of its institutions.
glossary-balance-longtermism-setiya · First-hand source The New Moral Mathematics Kieran Setiya / Boston Review · Published 2022-08-15 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read Setiya’s review sections on population ethics and the priority given to future survival over present suffering. The entry attributes his ethical criticism and does not adopt the article’s empirical forecasts.
Read the abstract and submission history. The authors argue for clearer definitions, plural values and better risk methods. Their criticism of one influential framework is not a finding that all catastrophe research is invalid.
glossary-balance-utilitarianism-objections · First-hand source Objections to Utilitarianism and Responses Richard Yetter Chappell, Darius Meissner and William MacAskill / Utilitarianism.net · Published 2023 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the rights and demandingness objections and the authors’ response strategies. This textbook defends utilitarianism; its replies are arguments, not a resolution accepted by all moral philosophers.
Read the sections on delayed benefits, uncertain duration, weak enforcement, underground development and political compromise. These are PauseAI’s own possible failure scenarios, not measured outcomes. The page includes older threshold examples; current policy scope follows the April 2026 proposal. Publication date unspecified.
software-lens-nist-computer · First-hand source Computer National Institute of Standards and Technology · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the glossary definition identifying a computer as a device that processes digital data according to program instructions. The page attributes the definition to NIST SP 800-34 Rev. 1; that full contingency-planning document was not reviewed. Live glossary publication date unspecified.
software-lens-nist-software · First-hand source software National Institute of Standards and Technology · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the definitions of software as programs and associated data, and the distinction from physical hardware. The glossary aggregates several contextual definitions, including entries for related hardware and firmware terms; it is not a single universal definition. Referenced standards were not read in full. Live page publication date unspecified.
Read the updated definition and explanatory sections on techniques, human roles, autonomy, physical and virtual environments, inputs, models and outputs (pages 4 and 6 to 9). Publication date follows the OECD publication landing page. Used to distinguish models from systems and learned methods from knowledge-based approaches, not as a claim about human-like understanding or a legal classification of a particular product.
Read the host transcript, especially 00:43:04 to 00:46:47 and 00:56:10 to 01:04:10. The host describes editorially reworked transcripts. The proposed power cap lacks an implementation method; gradual release concerns deployment.
sutskever-superalignment-2023 · First-hand source Introducing Superalignment OpenAI / Jan Leike and Ilya Sutskever · Published 2023-07-05 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the coauthored introduction, research approach and footnotes. Used for historically expressed extinction concern and the distinction between a research aim and a solved control problem. The announced timetable is not treated as a result.
Read the issuer announcement linked from SSI, including Sutskever’s attributed research-scaling statement. Primary company announcement, not independent reporting or a capability test. SSI’s update is dated July 26; this release is dated July 27.
zuckerberg-future-2026 · First-hand source The Future is for Everyone Meta / Mark Zuckerberg · Published 2026-08-10 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the signed essay, especially the misuse, American leadership, existential-risk and control sections. Predictions and balance-of-power arguments are attributed to Zuckerberg; their effectiveness and announced governance are not independently verified.
zuckerberg-personal-2025 · First-hand source Personal Superintelligence Meta / Mark Zuckerberg · Published 2025-07-30 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the complete signed letter. It combines a building ambition with caution about novel safety risks and what to open source. Openness and release timing alone do not establish frontier-development pace.
Read the original interviewer’s introduction and episode highlights only. The full audio and transcript were not reviewed. Used only to record the recency search and retrieval limit, not as axis evidence or independent verification of product claims.
suleyman-humanist-2025 · First-hand source Towards Humanist Superintelligence Mustafa Suleyman · Published 2025-11-07 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the personal-site essay, including Containment is necessary and A safer superintelligence. This longer version is dated November 7; Microsoft’s shorter article is dated November 6. Claims about medical performance and future benefits are not adopted.
suleyman-code-2026 · First-hand source The Humanist AI Code of Conduct Mustafa Suleyman · Published 2026-09-15 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the complete personal essay from the original HTML with Python urllib after web retrieval failed. The page displays 15 September 2026. Used for stated control limits and future implementation; incident descriptions and consciousness claims are not independently established.
normal-technology-2025 · First-hand source AI as Normal Technology Knight First Amendment Institute / Arvind Narayanan and Sayash Kapoor · Published 2025-04-15 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the introduction and relevant passages in Parts I to IV on diffusion, capability and power, catastrophic misalignment, resilience and nonproliferation. The displayed publication date is April 15; the suggested citation instead says April 14. This record follows the displayed date and preserves the discrepancy here. Treat the essay as the authors' argument. Their September 2026 update qualifies its safety claims; cited incident reports and studies were not independently audited.
normal-technology-guide-2025 · First-hand source A guide to understanding AI as normal technology AI as Normal Technology / Arvind Narayanan and Sayash Kapoor · Published 2025-09-09 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the clarification of normal, restatement of the thesis, and response to Scott Alexander, including the distinction between economic and safety arguments. This is the authors' own explanation and reply, not independent verification of its forecasts. Linked conversations and all underlying empirical citations were not reviewed.
normal-technology-alexander-response-2025 · First-hand source AI As Profoundly Abnormal Technology AI Futures Project / Scott Alexander · Published 2025-07-24 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the opening, adoption by key actors, and sections on control, speculative risk and institutional assumptions. Used as a direct critique of the thesis, not as verification of the critic's forecasts or cited anecdotes. The older ai-futures.org link redirects to aifutures.org. The September 2025 reply and September 2026 revision are provided alongside this critique.
normal-technology-software-work-2026 · First-hand source Why AI hasn’t replaced software engineers, and won’t AI as Normal Technology / Arvind Narayanan and Sayash Kapoor · Published 2026-06-11 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the opening and discussion separating decisions, execution and delivery, including its explicit preference for accountability over slowing technical capabilities. Used for the dated public pace argument. Employment statistics, layoff reporting and underlying studies were not independently checked. The September 2026 essay supplies the more recent qualification about pausing experiments.
Read the framing of extraordinary intervention, its distinction between economic adoption and misuse, and the resilience section. The authors allow that precaution and temporary access restrictions can help while arguing for less restrictive defenses. Used as context for their own policy argument, not legal advice or verification of current law, capability gaps or historical examples.
normal-technology-research-2026 · First-hand source AI agents can't yet do open-ended AI research AI as Normal Technology / Sayash Kapoor and Arvind Narayanan · Published 2026-08-05 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the authors' research summary, its two-case design, stated limitations, and discussion of possible remaining bottlenecks. The linked paper, artifacts and agent logs were not audited. The newsletter reports the authors' own research but does not establish that current limitations are permanent or that all kinds of AI research are equally difficult.
Read the opening synthesis, Part 1 on control, governance and policy, and Part 3's explicit revisions and risk assessment, plus the policy discussion before Part 3. Used for the authors' advocacy and self-described changes. Their accounts of third-party incidents, product behavior, investment and law were not independently verified. The proposed record does not repeat those accounts as established facts or infer a whole-lab halt.
glossary-symbolic-stanford · First-hand source What is Traditional AI? Stanford HAI · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read Stanford's institutional definition of explicitly programmed rules and symbolic reasoning. No publication date shown; mental-state wording is not adopted.
glossary-expert-stanford · First-hand source What is an Expert System? Stanford HAI · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read knowledge-base, if-then-rule and limited-domain definitions. No publication date shown. Historical priority and broad interpretability claims are not adopted.
foundation-model-report · First-hand source On the Opportunities and Risks of Foundation Models Rishi Bommasani and coauthors / Stanford CRFM / arXiv · Published 2021-08-16 · Updated 2022-07-12 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read abstract, introduction and shared-model risk discussion. Dates follow arXiv versions. A terminology proposal and research synthesis.
scaling-kaplan · First-hand source Scaling Laws for Neural Language Models Jared Kaplan and coauthors / arXiv · Published 2020-01-23 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read abstract and submission record. Empirical relationships concern language-model prediction loss and training resources, not an AGI date.
Read abstract and submission record. Tests balance parameter count and training tokens under a fixed compute budget. Results are scoped to the studied setups.
Read definition and recursive-training setup. Checked the 2025 correction to a mathematical symbol. Findings are not evidence that every synthetic-data method fails.
Read abstract, introduction and language-model setup. Accumulating original and generated examples avoided collapse in tested settings. The original TinyStories text was itself synthetic.
data-work-typology · First-hand source A typology of artificial intelligence data work James Muldoon, Callum Cant, Boxi Wu and Mark Graham / Big Data & Society · Published 2024-03-18 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read abstract, fieldwork scope and data-work definitions. Includes computer-vision data workplaces and other fieldwork; not a representative census of all AI labor.
data-cascades-author-summary · First-hand source Data Cascades in Machine Learning Nithya Sambasivan / Google Research · Published 2021-06-04 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the researcher's account of the interview study, examples and data-work recommendations. Findings concern the studied projects, not every AI system.
ai-control-original · First-hand source AI Control: Improving Safety Despite Intentional Subversion Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan and Fabien Roger / Redwood Research / arXiv · Published 2023-12-12 · Updated 2024-07-23 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read abstract, introduction and sections 5.1.2 and 5.2. Programming-task experiments, with human review simulated by a model. Control is not declared solved.
Read abstract and version history. Reports prompt-injection attacks against monitors on two control benchmarks. Findings are scoped to tested protocols, not every possible safeguard.
ai-control-threats · First-hand source Prioritizing threats for AI control Ryan Greenblatt / Redwood Research · Published 2025-03-19 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the proposed threat categories, permission limits and blocking-review discussion. Prospective threat modeling and author priorities, not observed catastrophic events.
What we read: The presentation's title slide, used for the name and public handle only.
No archive check recorded.
Read the presentation's title slide, which displays Daniel Card and @Uk_Daniel_Card. Used only to corroborate the name and public handle. The retrieved slides do not establish a publication date, authenticate the supplied export or establish a current position on either map axis.
What we read: Selected text in a supplied export. Complete threads and media were not reviewed.
The linked original was not independently retrieved. Reading the supplied copy does not verify its authorship.
No archived copy verified · Archive check 2026-09-15
The Wayback availability lookup returned HTTP 429 (rate limited). No capture was verified; this does not show that no archived copy exists.
Read the text in a user-supplied export, matching Author Username to UK_Daniel_Card and retaining the original post ID and URL. The live X URL returned 403, so the original was not independently retrieved. The date follows the export's Created At field, which has no timezone. Repeated rows label this same authored post both Tweet and Quoted; other authors' quoted or reposted text was not attributed to Card. Media and complete threads were not reviewed. The post's allegations and technical conclusions were not independently verified. The read label applies only to the supplied export text, not the linked original. The author field is an attribution in the export, not independent verification of authorship.
What we read: Selected text in a supplied export. Complete threads and media were not reviewed.
The linked original was not independently retrieved. Reading the supplied copy does not verify its authorship.
No archived copy verified · Archive check 2026-09-15
The Wayback availability lookup returned HTTP 429 (rate limited). No capture was verified; this does not show that no archived copy exists.
Read the reply text in a user-supplied export after matching its author field to UK_Daniel_Card. The live original returned 403. The date follows the export's timestamp without assigning a timezone. Full conversation context and media were not retrieved. Used for his stated monitoring priority, not as proof that model-level monitoring is useless. The read label applies only to the supplied export text, not the linked original. The author field is an attribution in the export, not independent verification of authorship.
What we read: Selected text in a supplied export. Complete threads and media were not reviewed.
The linked original was not independently retrieved. Reading the supplied copy does not verify its authorship.
No archived copy verified · Archive check 2026-09-15
The Wayback availability lookup returned HTTP 429 (rate limited). No capture was verified; this does not show that no archived copy exists.
Read the text in a user-supplied export. Its Origin label was not treated as authorship evidence; attribution uses the UK_Daniel_Card author field. The live original returned 403. The date follows the export's timestamp, which has no timezone. Used for the argument about matching tools to tasks, not as a universal claim about model determinism or frontier development policy. The read label applies only to the supplied export text, not the linked original. The author field is an attribution in the export, not independent verification of authorship.
What we read: Selected text in a supplied export. Complete threads and media were not reviewed.
The linked original was not independently retrieved. Reading the supplied copy does not verify its authorship.
No archived copy verified · Archive check 2026-09-15
The Wayback availability lookup returned HTTP 429 (rate limited). No capture was verified; this does not show that no archived copy exists.
Read the text in a user-supplied export after matching the author field to UK_Daniel_Card. The live X original returned 403. The date follows the export's timestamp, which has no timezone. His warning against inferring the opposite of a criticism is retained as a qualification; the post does not provide an overall catastrophic-risk estimate. The read label applies only to the supplied export text, not the linked original. The author field is an attribution in the export, not independent verification of authorship.
card-pwndefend-risk-framing-2026 · First-hand source AI: Fear it, so I can sell you the cure! PwnDefend / Daniel Card · Published undated · Material last read 2026-09-15
What we read: The published commentary and its AI-assistance disclosure.
No archive check recorded.
Read the original page, especially its digital-nuclear-weapon analogy, risk framing and closing disclosure. The disclosure says the article was generated with Opus 4.8 from the author's prompts and judges it broadly on point. Treat it as published, AI-assisted commentary, not independently verified technical evidence. The permalink contains 2026/07/04, but no separate publication date was displayed in the retrieved text, so published remains null. A direct metadata retrieval attempt returned 403.
What we read: Selected text in a supplied export. Complete threads and media were not reviewed.
The linked original was not independently retrieved. Reading the supplied copy does not verify its authorship.
No archive check recorded.
The copied post quotes support for AI opportunities and involvement of cybersecurity practitioners, then explicitly signals agreement. The quoted wording is not presented as Card's original writing. This supports an endorsement of opportunities, not a preference on frontier capability growth. Read from the supplied export after matching Author Username to UK_Daniel_Card and deduplicating post IDs. The original X URL returned HTTP 403 on 15 September 2026. The read status covers the copied text, not the linked original, and does not independently authenticate authorship. Publication follows the export date without assigning a timezone. Full threads and media were not reviewed.
What we read: Selected text in a supplied export. Complete threads and media were not reviewed.
The linked original was not independently retrieved. Reading the supplied copy does not verify its authorship.
No archive check recorded.
The copied reply discusses differences between computing architectures, benefits for prototypes, energy costs and risks of deploying technology before addressing security. Used to describe the attributed argument, not to verify its technical claims or historical examples. Read from the supplied export after matching Author Username to UK_Daniel_Card and deduplicating post IDs. The original X URL returned HTTP 403 on 15 September 2026. The read status covers the copied text, not the linked original, and does not independently authenticate authorship. Publication follows the export date without assigning a timezone. Full threads and media were not reviewed.
Read the original abstract and conference citation metadata. The findings concern three commercial gender-classification systems and the study's dataset, not all facial analysis or current versions. The limitation states the scope of this historical evidence; no independent replication was performed.
development-turing · First-hand source Computing Machinery and Intelligence Alan M. Turing / Mind; text hosted by Simon Fraser University · Published 1950 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the original paper's university-hosted transcription, including the imitation game, objections, digital computers and learning machines. Used for Turing's proposed questions, not a claim that a contemporary model passes a universally agreed intelligence test.
Read the dated proposal, its stated conjecture and proposed topics. The source describes plans for summer 1956; the separate Dartmouth institutional history supports that the gathering took place.
development-dartmouth-history · First-hand source Our Story Dartmouth / AI at Dartmouth · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the institution's retrospective on the summer 1956 gathering and John McCarthy's organizing role. No publication date is shown. Its promotional claims and present-day research announcements are not used for the historical event.
Read the original paper's university-hosted text, particularly the abstract, keyword and reassembly rules, and the unhappy-sentence example showing transformations independent of meaning. Date retains the issue's month precision.
development-backpropagation · First-hand source Learning representations by back-propagating errors David E. Rumelhart, Geoffrey E. Hinton and Ronald J. Williams / Nature · Published 1986-10-09 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Visually read the author-hosted scan's opening page, page 533: publication date, abstract and explanation of desired outputs and hidden units. Used for the paper's method and task setup; not a claim that this paper was the first invention of backpropagation.
development-deep-blue · First-hand source Deep Blue IBM History · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read IBM's account of the May 1997 rematch, score, search hardware and chess-specific preparation. This is the builder's retrospective; claims about wider industrial benefits are not adopted. The page supplies a match month, not an exact day.
Read the project-hosted original paper's abstract and sections 1–3 on size, WordNet categories, image collection, human checking and the subset used for analysis. Historical counts describe the paper's version, not the current dataset.
Read the proceedings PDF's abstract, introduction and competition-results section. The PDF describes the 2012 winning entry and the limits of image datasets; its numbers differ from the older abstract text on the proceedings landing page, so the PDF is authoritative here.
development-alphago · First-hand source AlphaGo Google DeepMind · Published undated · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the builder's retrospective sections Our approach and The matches, covering expert games, self-play, search and the March 2016 Lee Sedol result. Broad promotional claims about creativity or solving other domains are not adopted.
Read the abstract and submission history: bidirectional text representations, task-specific fine-tuning and reported results on eleven language tasks. The timeline date is the original preprint, not the later revision.
development-gpt3 · First-hand source Language Models are Few-Shot Learners Tom B. Brown and coauthors / OpenAI, arXiv · Published 2020-05-28 · Updated 2020-07-22 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the abstract, sections 4–5 and appendix C on training/test overlap, and the submission history. Covers text-supplied demonstrations, weak tasks and overlap checks that find small effects on most tests but flag some results. Results are the model developers' report.
development-chatgpt · First-hand source Introducing ChatGPT OpenAI · Published 2022-11-30 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the original announcement, Methods and Limitations. The page now explicitly labels itself the 2022 introduction. Claims concern that launch version, including dialogue training and reported errors, not every later ChatGPT model.
development-gpt4 · First-hand source GPT-4 Technical Report OpenAI / arXiv · Published 2023-03-15 · Updated 2024-03-04 · Material last read 2026-09-15
Read scope is described in the source note.
No archive check recorded.
Read the abstract, introduction and Scope and Limitations section; checked publication and revision dates against the abstract page. The report supports image/text inputs, training overview and withheld technical details. Performance statements are attributed to OpenAI.
Read publication metadata, abstract, main description and Model limitations, including molecular geometry, overlapping atoms and limited conformations. Comparisons are the authors' study results; no clinical or drug-development outcome is inferred.
Read the pinned first version's abstract and section 5, including training variants, smaller models and limitations in tool calls, multi-turn tasks, language mixing and prompting. This historical record intentionally uses v1, not the revised January 2026 text. Results are the developers' report.