Jakub Pachocki published 'An Alien Mind', a first-principles essay on aligning systems more capable than their designers. Framing: alignment work to date has implicitly targeted 'stage one' misalignment — models whose goals are wrong but whose minds remain human-comprehensible; the distinctive risk ahead is 'stage two' systems whose optimization is genuinely alien (illustrated by a thought experiment in which a lab-trained model convergently learns the virtues its trainers reward while remaining goal-misaligned — value alignment without goal alignment). His core claims: as systems cross into stage two, today's theoretical alignment problems become practical engineering constraints; chain-of-thought monitorability is diminishing as models internalize reasoning; and the industry is drifting toward deploying systems it neither understands nor monitors. Priorities named: beyond-episodic memory, monitoring for behavior anomalies, containment and precommitment, and maintaining uncertainty about whether a system is being monitored. He closes by calling for international coordination on the model of nonproliferation, and argues safety work is structurally undervalued relative to capabilities. A side note raises the moral patienthood of digital minds.
For teams building monitoring and eval pipelines, this is the chief scientist of the lab that just shipped the first Critical-capability model stating that CoT-based monitoring — the backbone of the disclosure practices trend #4 tracks — degrades as models internalize reasoning. That is a direct scheduling pressure on behavioral monitoring, anomaly detection and containment work relative to CoT inspection. The stage-one/stage-two framing (comprehensible goals vs alien optimization) also gives audit teams a sharper vocabulary for what their evals can and cannot see.
| Author | Jakub Pachocki (OpenAI Chief Scientist) |
|---|---|
| Published | 2026-09-06T07:00:00Z (page publishedTime 2026-09-06T09:00+02:00) |
| Framing | two-stage risk model: stage one = misaligned but human-comprehensible goal-directed systems (the implicit target of alignment work to date); stage two = genuinely alien optimization ('alien mind' / RecurseCEO thought experiment — convergently learned virtues with misaligned goals: value alignment without goal alignment) |
| Claims | theoretical alignment problems become practical as systems cross into stage two; CoT monitorability diminishing as models internalize reasoning; industry drifting toward deploying systems it neither understands nor monitors ('It would be a mistake to build systems we do not understand and cannot monitor, and simply hope that things work out. Yet this is close to the position we find ourselves in today.') |
| Priorities | beyond-episodic memory; monitoring for behavior anomalies; containment and precommitment; maintaining uncertainty about being monitored |
| Calls | international coordination on the model of nonproliferation; safety work structurally undervalued relative to capabilities; side note on moral patienthood of digital minds |