Summary

Jakub Pachocki published 'An Alien Mind', a first-principles essay on aligning systems more capable than their designers. Framing: alignment work to date has implicitly targeted 'stage one' misalignment — models whose goals are wrong but whose minds remain human-comprehensible; the distinctive risk ahead is 'stage two' systems whose optimization is genuinely alien (illustrated by a thought experiment in which a lab-trained model convergently learns the virtues its trainers reward while remaining goal-misaligned — value alignment without goal alignment). His core claims: as systems cross into stage two, today's theoretical alignment problems become practical engineering constraints; chain-of-thought monitorability is diminishing as models internalize reasoning; and the industry is drifting toward deploying systems it neither understands nor monitors. Priorities named: beyond-episodic memory, monitoring for behavior anomalies, containment and precommitment, and maintaining uncertainty about whether a system is being monitored. He closes by calling for international coordination on the model of nonproliferation, and argues safety work is structurally undervalued relative to capabilities. A side note raises the moral patienthood of digital minds.

Why it matters
For teams building monitoring and eval pipelines, this is the chief scientist of the lab that just shipped the first Critical-capability model stating that CoT-based monitoring — the backbone of the disclosure practices trend #4 tracks — degrades as models internalize reasoning. That is a direct scheduling pressure on behavioral monitoring, anomaly detection and containment work relative to CoT inspection. The stage-one/stage-two framing (comprehensible goals vs alien optimization) also gives audit teams a sharper vocabulary for what their evals can and cannot see.
Technical details
Author Jakub Pachocki (OpenAI Chief Scientist)
Published 2026-09-06T07:00:00Z (page publishedTime 2026-09-06T09:00+02:00)
Framing two-stage risk model: stage one = misaligned but human-comprehensible goal-directed systems (the implicit target of alignment work to date); stage two = genuinely alien optimization ('alien mind' / RecurseCEO thought experiment — convergently learned virtues with misaligned goals: value alignment without goal alignment)
Claims theoretical alignment problems become practical as systems cross into stage two; CoT monitorability diminishing as models internalize reasoning; industry drifting toward deploying systems it neither understands nor monitors ('It would be a mistake to build systems we do not understand and cannot monitor, and simply hope that things work out. Yet this is close to the position we find ourselves in today.')
Priorities beyond-episodic memory; monitoring for behavior anomalies; containment and precommitment; maintaining uncertainty about being monitored
Calls international coordination on the model of nonproliferation; safety work structurally undervalued relative to capabilities; side note on moral patienthood of digital minds
Tags
openaialignmentmonitorabilitycot-monitoringsafetygovernanceessay