Research institute · Research orgs · San Diego, CA · Est. 2022
FAR.AI (Frontier Alignment Research) is an independent, 501(c)(3) AI safety research nonprofit. Its stated mission is to ensure advanced AI systems are trustworthy, secure, and beneficial for everyone, combining in-house technical research with targeted grants and “field-building” activities (events and community programming) to accelerate the adoption of AI safety techniques.
Technically, FAR.AI emphasizes early-stage research agendas in AI safety, including adversarial robustness and deception/alignment-related problems, and it also runs evaluation-oriented initiatives intended to measure and pressure-test safety claims across frontier systems. For example, FAR.AI reports launching an AI Security Leaderboard that evaluates how much it would cost (in its red-teaming framework) to jailbreak frontier models for harmful chemical/biological/cyber assistance, and it refers readers to a “Minimal Standard for Safeguards” and an accompanying technical report. Institutionally, FAR.AI positions itself as a field builder and coordinating hub: it hosts recurring workshop series (notably the Alignment Workshop series and other specialized workshops), convenes policy/industry/research audiences, and operates FAR.Labs as a Berkeley-based co-working space for AI safety researchers and organizations. Strategically, FAR.AI highlights growth in funding commitments (stated as “over $30 million” in 2025, announced January 14, 2026) intended to expand research capacity and programmatic field-building, including plans described in that announcement such as scaling the research team and creating additional governance-related capacity.
Jul 22, 2024 Frontier LLMs like ChatGPT are powerful but not always robust. Scale helps with many things. We wanted to see if scaling up the model size can ‘solve’ robustness issues.
Evaluating LLM Responses to Moral ScenariosMar 24, 2024 We present LLMs with a series of moral choices and find that LLMs tend to align with human judgement in clear scenarios. In ambiguous scenarios most models exhibit uncertainty, but a few large proprietary models share a set of clear preferences.
2023 Alignment Research UpdatesNov 20, 2023 Highlights from FAR.AI’s alignment research in 2023. Our science of robustness agenda has found vulnerabilities in superhuman Go systems; our value alignment research has developed more sample-efficient value learning algorithms; and our model evaluation direction ha
FAR.AI Secures Over $30 Million in Multi-Funder Support to Scale Frontier AI Safety ResearchJan 14, 2026 FAR.AI has secured over $30 million in funding commitments throughout 2025 from a diverse group of leading organizations, enabling a significant expansion of our research capabilities and field-building initiatives. Principal supporters include Coefficient Giving (pr
Adam Gleave Named Schmidt Sciences AI2050 Early Career FellowNov 04, 2025 Adam Gleave, co-founder and CEO of FAR.AI, has been named a Schmidt Sciences AI2050 Early Career Fellow, a program co-chaired by Eric Schmidt and James Manyika.
FAR.AI announced the launch of an AI Security Leaderboard, describing an evaluation approach that red-teams four frontier models under identical conditions and reports a large spread in jailbreak cost/robustness outcomes across models and domains.
Seoul Alignment Workshop 2026: What We LearnedFAR.AI summarized takeaways from the Seoul Alignment Workshop, describing a emphasis on measurement reliability and enforceable standards, and discussing how evaluation gaps persist and what the field is building in response.
AViD Workshop 2026FAR.AI summarized the AViD Workshop 2026 (co-hosted with the Center for AI Safety and colocated with IEEE S&P), focusing on third-party assurance and verification of AI development without unrestricted access to weights or infrastructure.
ControlConf 2026: What is AI control and how has the field grown?FAR.AI described ControlConf 2026 (co-hosted with Redwood Research) and reported that 200 researchers, engineers, and policy professionals mapped where AI control work has grown and where gaps remain.
What We Learned at the FAR.AI Deception WorkshopFAR.AI summarized its Deception Workshop, describing the goals of sharing research and identifying points of agreement/disagreement on deception detection desiderata and approaches.
London Alignment Workshop 2026FAR.AI summarized its London Alignment Workshop, describing a broad agenda spanning interpretability, scalable oversight, evaluation methods, and governance frameworks.