美国 AI 研究者,在艾伦人工智能研究院(Ai2)主导后训练工作,参与构建了开源的 Tulu 与 OLMo 模型及训练配方家族。他在 UC Berkeley 获博士学位,此前在 Hugging Face 从事 RLHF。通过广受关注的 Interconnects 通讯,他成为讲解「基于人类反馈的强化学习」、偏好微调以及开源与闭源模型开发机制最清晰的公开声音之一。
American AI researcher who leads post-training work at the Allen Institute for AI (Ai2), where he helped build the open Tulu and OLMo model and recipe families. He earned his PhD at UC Berkeley and previously worked on RLHF at Hugging Face. Through his widely read Interconnects newsletter he has become one of the clearest public explainers of reinforcement learning from human feedback, preference tuning, and the mechanics of open versus closed model development.