Quotes · John Schulman
AI QUOTES

John Schulman

Research Scientist, Anthropic

NOTABLE QUOTES 5 quotes

“I've decided to pursue alignment research at Anthropic, where I believe I can do the most impactful work on this critical problem.”

August 6, 2024 · Departure announcement on X · Source

“PPO was designed to be simple, general, and stable. Those are the properties that matter most for a training algorithm.”

July 20, 2017 · PPO paper publication · Source

“Reinforcement learning from human feedback is not alignment. It's a first step, and we need many more steps.”

November 30, 2023 · NeurIPS presentation

“The gap between making models that are helpful and making models that are genuinely aligned with human values is enormous.”

March 15, 2024 · AI safety workshop

“I believe the alignment problem is solvable, but only if we dedicate serious resources to it now.”

August 6, 2024 · Departure announcement on X

John Schulman is one of the most technically influential figures in modern AI, primarily through his creation of PPO (Proximal Policy Optimization), the algorithm that underpins RLHF and makes current language models follow instructions effectively. His departure from OpenAI — the company he co-founded — to join Anthropic in August 2024 represented one of the most significant talent movements in AI history.

Schulman’s public statements are relatively rare compared to other AI leaders, reflecting a researcher’s temperament rather than a CEO’s. When he does speak publicly, his words carry technical precision. His observation that RLHF is not alignment but merely a first step reflects a deep understanding of the gap between making models that appear aligned (helpful, harmless, honest in typical interactions) and making models that are robustly aligned under adversarial conditions or novel situations.

His departure from OpenAI, combined with Jan Leike’s earlier exit, created a narrative of alignment-focused researchers leaving the company they helped found because they felt safety was being deprioritized. As the creator of the algorithm that makes ChatGPT work, Schulman’s implicit criticism of OpenAI’s direction carried significant weight.

The move to Anthropic aligned with Schulman’s stated belief that alignment research requires dedicated focus and resources. His technical background makes him uniquely positioned to bridge the gap between the capabilities research that produces more powerful models and the alignment research that aims to make those models safe.

Schulman represents the quiet technical voice in AI discourse — someone whose code has shaped the entire field but who speaks publicly only when he feels it matters enough to break his usual silence.