Creator of HELM benchmark · Founded Together AICreator of HELM benchmark · Founded Together AI
nlpsafetysystems
Stanford professor who became one of the most important voices in LLM evaluation and accountability. Created HELM (Holistic Evaluation of Language Models, 2022), the most comprehensive and standardized benchmark for evaluating large language models across multiple dimensions including accuracy, robustness, fairness, and toxicity. Led the Stanford Alpaca project (2023), demonstrating that a small model fine-tuned on GPT-generated data could match larger models. Co-founded Together AI to democratize access to open-source AI infrastructure.
Stanford professor who became one of the most important voices in LLM evaluation and accountability. Created HELM (Holistic Evaluation of Language Models, 2022), the most comprehensive and standardized benchmark for evaluating large language models across multiple dimensions including accuracy, robustness, fairness, and toxicity. Led the Stanford Alpaca project (2023), demonstrating that a small model fine-tuned on GPT-generated data could match larger models. Co-founded Together AI to democratize access to open-source AI infrastructure.