Co-founder of OpenAI and creator of PPO (Proximal Policy Optimization), the reinforcement learning algorithm that made RLHF possible. Joined Anthropic in August 2024 after nine years at OpenAI.
Schulman's departure from OpenAI to Anthropic was one of the most significant talent moves in AI history. The man who invented PPO — the algorithm that literally makes RLHF work — leaving the company he co-founded signals something profound about OpenAI's direction. His stated reason (wanting to focus on alignment research) is credible given his track record, and Anthropic gaining the architect of modern RL-based alignment is a massive win for their research bench. This is the kind of move that compounds over years.
No recent LinkedIn activity tracked yet.