Chinese University of Hong Kong (Shenzhen) PhD. Core researcher at DeepSeek. First author of DeepSeek-R1 (2025), the reasoning model that showed reinforcement learning could teach LLMs to reason without supervised fine-tuning, reaching performance competitive with OpenAI o1 at much lower cost. Also key contributor to DeepSeek-Coder and DeepSeek-V2/V3.