北京大学人工智能研究院助理教授 · 智源大模型安全中心负责人Assistant Professor PKU · Head of LLM Safety, BAAI
safetyrlagents
北京大学人工智能研究院助理教授,兼任北京智源人工智能研究院大模型安全研究中心负责人。研究多智能体强化学习与 AI 对齐,主导 PKU-Beaver/Safe-RLHF——最早带显式安全约束的开源 RLHF 框架之一,以及中文大模型价值对齐评测基准。中国 AI 安全研究群体的核心人物。
Assistant professor at Peking University's Institute for AI and head of the large-model safety research centre at BAAI (Beijing Academy of AI). Works on multi-agent reinforcement learning and AI alignment, including PKU-Beaver/Safe-RLHF — one of the first open-source RLHF frameworks with explicit safety constraints — and value-alignment benchmarks for Chinese LLMs. A central figure in China's emerging AI-safety research community.