AI 学者图谱AI Scholar Graph John Schulman

John SchulmanJohn Schulman

Anthropic

PPO 算法创造者 · OpenAI 联合创始人Creator of PPO · Co-founder of OpenAI

rlsafety
加州理工学院物理学学士(2010),后在 UC Berkeley 攻读 EECS 博士,导师 Pieter Abbeel。2015年底联合创办 OpenAI(博士尚未完成)。在 OpenAI 开发了 TRPO(2015)和 PPO(近端策略优化,2017),PPO 成为使用最广泛的强化学习算法,也是 RLHF 背后让 ChatGPT 成为可能的核心算法(引用 23000+)。2022-2024年联合领导 OpenAI 后训练团队,直接负责 ChatGPT 的开发。2024年8月加入 Anthropic 专注 AI 安全。2025年2月转任 Thinking Machines Lab 首席科学家,并透露该实验室计划 2026 年发布自研前沿模型,2026 年 3 月与 NVIDIA 达成合作、部署 1 GW Vera Rubin 算力支撑。
Earned a BS in Physics from Caltech (2010), then pursued a PhD in EECS at UC Berkeley under Pieter Abbeel. Co-founded OpenAI in late 2015, shortly before completing his PhD. At OpenAI, he developed TRPO (2015) and PPO (Proximal Policy Optimization, 2017), which became the most widely used reinforcement learning algorithm -- and the core algorithm behind RLHF that makes ChatGPT work (23,000+ citations). Co-led OpenAI's post-training team (2022-2024), directly overseeing ChatGPT's development. Joined Anthropic in August 2024 to focus on AI safety. Later moved to Thinking Machines Lab as Chief Scientist (February 2025), where he said the lab plans to ship its own frontier models in 2026, backed by a March 2026 NVIDIA partnership for 1 GW of Vera Rubin compute.

同一人物 · 姊妹站:听 TA 的播客访谈(AI Podcast) · 读 TA 的论文与长文(AI Paper)

在关系图谱中查看 John Schulman →View John Schulman in the graph →

时间线Timeline

关系网络Connections

翁丽莲 Lilian WengOpenAI 同事Ilya Sutskever联合创办 OpenAISam AltmanOpenAI 管理层Dario Amodei均为前 OpenAI,现 AnthropicPieter AbbeelUC Berkeley 博士导师Elon MuskOpenAI 2015 创始成员Andrej KarpathyOpenAI 2015 创始团队Greg BrockmanOpenAI 2015 联合创始人Mira MuratiSchulman 2025 从 Anthropic 加入 Thinking Machines内森·兰伯特 Nathan LambertRLHF / 后训练研究脉络保罗·克里斯蒂亚诺 Paul ChristianoRLHF 渊源 ← 返回完整图谱← Back to the graph