GPT-3 论文(2020)《Language Models are Few-Shot Learners》的第一作者。该论文展示了一个 1750 亿参数的模型可以仅凭上下文中的少量示例就完成任务,无需微调。这篇论文是一个分水岭时刻,证实了 Scaling 假说,直接引发了大语言模型浪潮,催生了 ChatGPT 和整个生成式 AI 浪潮。
Lead author of the GPT-3 paper (2020), 'Language Models are Few-Shot Learners', which demonstrated that a 175-billion parameter model could perform tasks with just a few examples in context, without fine-tuning. This paper was a watershed moment that proved the scaling hypothesis and directly sparked the large language model revolution, leading to ChatGPT and the entire generative AI wave.