Earned a BS in Physics from Caltech (2010), then pursued a PhD in EECS at UC Berkeley under Pieter Abbeel. Co-founded OpenAI in late 2015, shortly before completing his PhD. At OpenAI, he developed TRPO (2015) and PPO (Proximal Policy Optimization, 2017), which became the most widely used reinforcement learning algorithm -- and the core algorithm behind RLHF that makes ChatGPT work (23,000+ citations). Co-led OpenAI's post-training team (2022-2024), directly overseeing ChatGPT's development. Joined Anthropic in August 2024 to focus on AI safety. Later moved to Thinking Machines Lab as Chief Scientist (February 2025), where he said the lab plans to ship its own frontier models in 2026, backed by a March 2026 NVIDIA partnership for 1 GW of Vera Rubin compute.