Research institute · Research orgs · London · Est. 2013
The Center on Long-Term Risk (CLR) is a UK-registered nonprofit research organization working on worst-case risks from advanced artificial intelligence, with an emphasis on how “harmful propensities” in AI systems could interact to produce catastrophic outcomes. Its stated mission is to address worst-case risks from the development and deployment of advanced AI systems, and it currently focuses on (1) harmful propensities in AI systems and (2) cooperation failures between such systems.
CLR’s work combines interdisciplinary research, publishing, grantmaking (through the CLR Fund), and community-building programs aimed at training and supporting researchers and others working on s-risk reduction.
Inoculation prompting applies one prompt to every training example, which leaves a backdoor through which similar prompts still elicit the undesired trait, and…
Taboo “equilibrium”: Less confused frames for research on AI bargaining — Center on Long-Term RiskCommon frames on bargaining problems — especially the assumption that agents play an equilibrium — are confused, and get in the way of research on safe Pareto…
Summer Update 2026 — Center on Long-Term RiskWe're writing to share our research progress and program updates from the first half of 2026.
Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values — Center on Long-Term RiskPeople use language models for practical questions whose answers are difficult to verify.
Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors — Center on Long-Term RiskInoculation prompting is a selective-generalization technique used against Emergent Misalignment.
A high-level model of AI bargaining — Center on Long-Term RiskAdvanced AIs might be capable of various credible commitments unavailable to humans, which they could use when bargaining with each other.
Evals for “SPI-incompatible” behavior and reasoning — Center on Long-Term RiskIn Part I of CLR's safe Pareto improvements (SPI) agenda, we gave our high-level strategy for evaluating models for SPI-incompatible behavior and reasoning.
Strategic Obfuscation of Deceptive Reasoning in Language Models — Center on Long-Term RiskLarge language models can exhibit different behaviors during training versus deployment, a phenomenon known as alignment faking.
Safe Pareto Improvements Research Agenda — Center on Long-Term RiskSafe Pareto improvements (SPIs) are modifications to agents’ bargaining strategies that make all parties better off, regardless of their original strategies.
Shaping the exploration of the motivation-space matters for AI safety — Center on Long-Term RiskWe argue that shaping RL exploration, and especially the exploration of the motivation-space, is understudied in AI safety and could be influential in…
Model Persona Research Agenda — Center on Long-Term RiskCLR’s overall mission is to reduce the risk of astronomical suffering from powerful AI, or s-risks.
Concrete Research Ideas on AI Personas — Center on Long-Term RiskWe have previously explained some high-level reasons for working on understanding how personas emerge in LLMs.
CLR reported research progress in the first half of 2026 across two streams: Model Personas (empirical) and Safe Pareto Improvements (conceptual). It also announced that the Summer Research Fellowship began with eight fellows and that CLR launched the Research Affiliates program in June.
Annual Review & Fundraiser 2025CLR summarized 2025 activities and plans for 2026, including a fundraising target. It described major organizational transition in 2025: Jesse Clifton stepped down as Executive Director in January (with Tristan Cook and Mia Taylor stepping into Managing Director and Research Director roles), Mia Taylor later departed in August, and Tristan Cook continued as Managing Director with Niels Warncke leading empirical research. It also described CLR’s clarified empirical and conceptual research agendas for 2026.
Summer Update 2025CLR reported organizational and research leadership transitions during 2025, including Mia Taylor’s decision to leave at the end of August, Tristan Cook taking over leadership after the transition, and progress on emerging “personas” and “strategic readiness” agendas.
Center on Long-Term Risk: 2025 PlansCLR laid out its 2025 strategy, including building expertise and collaborative relationships with the AI safety community through an externally legible empirical agenda, continuing macrostrategy work via a “strategic readiness” agenda, and targeting research hiring plus community-building activities.