“Over the past few months my team has been sailing against the wind. OpenAI's core product priorities have been racing ahead of safety work.”
“Building smarter-than-human machines is an inherently dangerous endeavor. OpenAI is ill-prepared for this.”
“Safety culture and processes have taken a back seat to shiny products.”
“I joined Anthropic because I believe it is currently the best place to do alignment research that matters.”
“The superalignment problem is not solved. We do not have the tools to align a system smarter than us. We need to build them.”
“If you believe superintelligence is coming and could be catastrophic, then working on alignment is the most important thing you can do.”
Jan Leike’s public departure from OpenAI in May 2024 was one of the most significant safety-focused exits in AI history. As co-lead of OpenAI’s Superalignment team (alongside Ilya Sutskever), Leike had been tasked with solving what many consider AI’s hardest problem: ensuring that systems smarter than humans remain aligned with human values. His decision to leave and his public explanation sent shockwaves through the AI community.
Leike’s resignation statement was notable for its directness. Rather than offering diplomatic platitudes about pursuing new opportunities, he explicitly stated that OpenAI’s leadership was not prioritizing safety adequately. His claim that the core product team was racing ahead while safety work was under-resourced directly contradicted OpenAI’s public narrative about treating safety as a top priority.
The contrast between OpenAI’s public commitments (dedicating 20% of compute to alignment) and Leike’s account of internal reality (his team sailing against the wind) crystallized skepticism about whether AI labs could be trusted to self-regulate on safety. If the person OpenAI had hired specifically to lead alignment felt the company was not taking it seriously, external observers had reason for concern.
Leike’s subsequent move to Anthropic was interpreted as a vote of confidence in Anthropic’s safety-first approach and a vote of no-confidence in OpenAI’s. His statements about Anthropic being the best place for alignment research that matters implicitly compared the two organizations’ cultures and priorities.
His technical position — that alignment of superintelligent systems is an unsolved problem requiring dedicated research investment — represents the consensus view among alignment researchers. The question his departure raised was whether commercial AI labs would actually make that investment when it competed with product development for resources.