Academic center · Science policy · Berkeley, California · Est. 2016
The Center for Human-Compatible Artificial Intelligence (CHAI) is a multi-institution research group based at UC Berkeley that aims to “reorient the general thrust of AI research towards provably beneficial systems.” CHAI frames the core problem as one of control: advanced AI systems can pursue objectives in ways that may be unpredictable to humans, so CHAI’s work emphasizes technical approaches intended to rule out harmful outcomes by design (including provable guarantees under specified assumptions).
Because what counts as “beneficial” depends on human preferences and social realities, CHAI explicitly includes social-scientific components alongside AI research. Institutionally, CHAI is positioned as an academic hub for both research and outreach: its progress reporting describes research outputs across areas including inverse/value learning and control, assistance and cooperation games, robustness and vulnerabilities, multi-agent cooperation, and social impacts—alongside advising for policy and regulation discussions.
Khanh Nguyen, Benjamin Plaut, Tu Trinh, and Mohamad Danesh introduce a fundamental coordination problem called Learning to Yield and Request Control (YRC), where the objective is to learn a strategy that determines when to act autonomously and when to seek expert assistance. They
Computational Frameworks for Human CareBrian Christian, CHAI Affiliate, has published an article titled “<a href="https://www.amacad.org/sites/default/files/publication/downloads/daedalus_wi25_12_christian.pdf">Computational Frameworks for Human Care</a>” in the most recent issue of Daedalus, the journal of the Americ
A Practical Definition of Political Neutrality for AIThere is an urgent need for a clear, consistent, and practical definition of political neutrality for AI systems.
RvS: What is Essential for Offline RL via Supervised Learning?Scott Emmons, PhD student, was an author on “RvS: What is Essential for Offline RL via Supervised Learning?”
Getting By Goal Misgeneralization With a Little Help From a Mentor“Tu Trinh, Ben Plaut, Khanh Nguyen, and Mohamad Danesh wrote the paper, “Getting By Goal Misgeneralization With a Little Help From a Mentor.” This paper explores whether goal misgeneralization can be mitigated by allowing an agent to ask for help when it is uncertain. The answer
Linear Probe Penalties Reduce LLM SycophancyVisiting ETH MsC student Henry Papadatos and supervising CHAI PhD student Rachel Freedman publish an article “Linear Probe Penalties Reduce LLM Sycophancy” at the NeurIPS SoLaR workshop. The paper demonstrates a generalizable methodology for reducing unwanted LLM behaviors that a
Rachel Freedman selected as inaugural Cooperative AI FellowRachel Freedman, PhD Student, has been selected as one of the fellows for <a href="https://www.cooperativeai.com/phd-fellowship/2025">Cooperative AI’s PhD Fellow Program</a>.
Representative Social Choice: From Learning Theory to AI AlignmentTianyi Qiu, CHAI Intern, wrote this paper which was accepted by NeurIPS 2024 Pluralistic Alignment Workshop. <a href="https://arxiv.org/pdf/2410.23953">Here</a> is the link to the paper.
Getting By Goal Misgeneralization With a Little Help From a MentorKhanh Nguyen, Mohamad Danesh, Ben Plaut, and Alina Trinh wrote <a href="https://arxiv.org/pdf/2410.21052">this paper</a> which was presented at Towards Safe & Trustworthy Agents Workshop at NeurIPS 2024.
Language-Guided World Models: A Model-Based Approach to AI ControlKhanh Nguyen, CHAI Postdoctoral Fellow, published a paper at the Fourth International Combined Workshop on Spatial Language Understanding and Grounded Communication for Robotics (ACL 2024).
“The Alignment Problem” Wins Xingdu Book AwardBrian Christian’s book “The Alignment Problem” was announced as the sole winner in the New Knowledge Category for Imported Editions at the Xingdu Book Award ceremony in China. The Chinese translation was published this past year by Hunan Science & Technology Press.
Social Choice Should Guide AI Alignment in Dealing with Diverse Human FeedbackRachel Feedman, CHAI Phd Student, and Wes Holliday, CHAI Affiliate, published a paper at the International Conference on Machine Learning
CHAI introduces a coordination problem called Learning to Yield and Request Control (YRC) and describes an open-source benchmark and evaluation approach for studying when agents should act autonomously vs. seek expert assistance.
Computational Frameworks for Human CareCHAI affiliate Brian Christian publishes an article titled “Computational Frameworks for Human Care” in Daedalus, framed around how alignment has progressed toward care-like relationships and the implications for human care understanding.
A Practical Definition of Political Neutrality for AICHAI announces a current research project to build political neutrality evaluations and presents a practical implementable definition of political neutrality for AI.
RvS: What is Essential for Offline RL via Supervised Learning?CHAI posts that Scott Emmons, PhD student, is an author on “RvS: What is Essential for Offline RL via Supervised Learning?”
Getting By Goal Misgeneralization With a Little Help From a MentorCHAI describes research on whether goal misgeneralization can be mitigated by allowing an agent to ask for help under uncertainty, along with noted weaknesses and future directions.
Linear Probe Penalties Reduce LLM SycophancyCHAI posts that visiting ETH MsC student Henry Papadatos and supervising CHAI PhD student Rachel Freedman published a paper at the NeurIPS SoLaR workshop describing a methodology intended to reduce unwanted LLM behaviors including sycophancy.