Draft:Evan Frick

Evan Frick

Evan Frick (born July 17, 2002) is an American machine learning researcher and computer scientist known for his contributions to the evaluation and post-training of large language models (LLMs). He is a researcher at Arena.ai and a contributor to the LMSYS Org (Large Model Systems Organization), specifically working on the Chatbot Arena benchmarking platform.[1]

Education

Frick attended the University of California, Berkeley, where he earned a B.A. in Computer Science (2024) and an M.S. in Electrical Engineering and Computer Science (2025). During his graduate studies, he conducted research on reward modeling and reinforcement learning from human feedback (RLHF) under the supervision of Jiantao Jiao and in collaboration with Ion Stoica.[2]

Career and research

Frick’s research focuses on the "post-training" phase of AI development, specifically how to align model outputs with human preferences. As a member of the LMSYS Org team, he co-developed Arena-Hard and the Prompt-to-Leaderboard (P2L) system, which aim to provide more nuanced evaluations than traditional static benchmarks.[3]

He has held positions at:

Arena.ai: Member of Technical Staff focusing on preference modeling. Nexusflow: Machine Learning Engineer, where he contributed to the development of the "Athene" series of open-weights models. Berkeley AI Research (BAIR): Research contributor to the Starling-7B project, a model series utilizing Reinforcement Learning from AI Feedback (RLAIF). His work has been published in machine learning conferences including the International Conference on Machine Learning (ICML) and the International Conference on Learning Representations (ICLR).[4]

Selected publications

Frick, E., et al. (2025). "Prompt-to-Leaderboard: Prompt-Adaptive LLM Evaluations." Proceedings of the 42nd ICML. Zhu, B., Frick, E., et al. (2024). "Starling-7B: Improving Helpfulness and Harmlessness with RLAIF." Conference on Language Modeling (COLM). Lambert, N., ... Frick, E., et al. (2024). "How to Evaluate Reward Models for RLHF." ICLR 2024.

References

  1. ^ Norton, Rosa (2025-05-06). "As companies pour billions into AI, a ranking system by UC Berkeley students has all eyes on it". Berkeley News. Retrieved 2024-11-17.
  2. ^ "Reward Modeling for Human Preferences". EECS at UC Berkeley. Retrieved 2024-11-17.
  3. ^ "How Early Access to NVIDIA GB200 Systems Helped LMArena Build a Model to Evaluate LLMs". NVIDIA Technical Blog. 2025-06-18.
  4. ^ Frick, Evan; Chen, Connor; et al. (2025). "Prompt-to-Leaderboard: Prompt-Adaptive LLM Evaluations". Proceedings of the 42nd International Conference on Machine Learning. 267. PMLR: 17672–17689.

Official website Evan Frick publications indexed by Google Scholar

Content Disclaimer

Informasi ini disarikan dari Wikipedia dan disajikan kembali untuk tujuan edukasi. Konten tersedia di bawah lisensi CC BY-SA 3.0. Kami tidak bertanggung jawab atas ketidakakuratan data yang bersumber dari kontribusi publik tersebut.

  1. The information displayed on this website is sourced in part or in whole from Wikipedia and has been adapted for the purpose of restating it. We strive to provide accurate and relevant information, however:
  2. There is no guarantee of absolute accuracy. Wikipedia is an open, collaborative project that can be edited by anyone, so information is subject to change.
  3. It is not intended to constitute professional advice. The content displayed is for informational and educational purposes only. For important decisions (e.g., medical, legal, or financial), please consult a professional.
  4. Content copyright. Wikipedia is licensed under the Creative Commons Attribution-ShareAlike License (CC BY-SA). This means that content may be reused with appropriate attribution and shared under a similar license.
  5. Responsible use. Any risk arising from the use of information from this website is entirely the responsibility of the user.