Research institute · Research orgs · Covina · Est. 2021
Alignment Research Center (ARC) is a U.S. non-profit research organization focused on “intent alignment”/alignment research for future machine learning systems. ARC’s stated technical agenda centers on building scalable methods to estimate neural network behavior from the model’s internal structure (mechanistic analysis of weights) rather than from black-box sampling over many inputs.
ARC’s homepage describes its current research focus as developing a theoretical foundation for mechanistic explanations of neural network behavior, and emphasizes computational efficiency relative to sampling, with applications including predicting out-of-distribution performance and detecting anomalies. ARC also operates a white-box model-evaluation direction formerly branded as “ARC Evals,” noting that evaluators are now at METR.
Over the past 15 months or so, ARC's technical agenda has developed quite a bit. The advent of the Matching Sampling Principle (MSP), and ideas like it, has begotten a host of concrete technical problems; progress on those problems has given us more philosophical clarity on the b
Announcing the ARC White-Box Estimation ChallengeARC has teamed up with AIcrowd to launch the ARC White-Box Estimation Challenge , a contest to improve upon our estimation algorithms for random MLPs . The warm-up round begins this week, and later rounds will have a total prize pool of at least $100,000. We are very grateful
Mechanistic estimation for expectations of random productsWe have developed some relatively general methods for mechanistic estimation competitive with sampling by studying problems that are expressible as expectations of random products . This includes several different estimation problems, such as random halfspace intersections, rando
Mechanistic estimation for wide random MLPsThis post covers joint work with Wilson Wu, George Robinson, Mike Winer, Victor Lecomte and Paul Christiano. Thanks to Geoffrey Irving and Jess Riedel for comments on the post. In ARC's latest paper, we study the following problem: given a randomly initialized multilayer perceptr
AlgZoo: uninterpreted models with fewer than 1,500 parametersThis post covers work done by several researchers at, visitors to and collaborators of ARC, including Zihao Chen, George Robinson, David Matolcsi, Jacob Stavrianos, Jiawei Li and Michael Sklar. Thanks to Aryan Bhatt, Gabriel Wu, Jiawei Li, Lee Sharkey, Victor Lecomte and Zihao Ch
Competing with samplingIn 2025, ARC has been making conceptual and theoretical progress at the fastest pace that I've seen since I first interned in 2022. Most of this progress has come about because of a re-orientation around a more specific goal: outperforming random sampling when it comes to underst
Obstacles in ARC's research agendaFormer ARC researcher David Matolcsi has put together a sequence of posts that explores ARC's big-picture vision for our research and examines several obstacles that we face. We think these posts will be useful to readers who are interested in digging into the details of our big-
A computational no-coincidence principleIn a recent paper in Annals of Mathematics and Philosophy, Fields medalist Timothy Gowers asks why mathematicians sometimes believe that unproved statements are likely to be true. For example, it is unknown whether \(\pi\) is a normal number (which, roughly speaking, means that e
A bird's eye view of ARC's researchOver the last few months, ARC has released a number of pieces of research. While some of these can be independently motivated, there is also a more unified research vision behind them. The purpose of this post is to try to convey some of that vision and how our individual
Low Probability Estimation in Language ModelsARC recently released our first empirical paper: Estimating the Probabilities of Rare Language Model Outputs . In this work, we construct a simple setting for low probability estimation — single-token argmax sampling in transformers — and use it to compare the performance of vari
Research update: Towards a Law of Iterated Expectations for Heuristic EstimatorsLast week, ARC released a paper called Towards a Law of Iterated Expectations for Heuristic Estimators , which follows up on previous work on formalizing the presumption of independence . Most of the work described here was done in 2023. A brief table of contents for this post: W
Estimating Tail Risk in Neural NetworksMachine learning systems are typically trained to maximize average-case performance. However, this method of training can fail to meaningfully control the probability of tail events that might cause significant harm. For instance, while an artificial intelligence (AI) assistant m
Michael Winer posts an updated picture of ARC’s research pipeline for aligning powerful AI, including a role for mechanistic estimators, structure monitoring during training, and estimation-based safety relevance rather than waiting for rare catastrophic behaviors in samples.
Announcing the ARC White-Box Estimation ChallengeARC announces a white-box estimation challenge with AIcrowd to improve ARC’s estimation algorithms for random MLPs, including details on the warm-up round and stated prize pool baseline.
Mechanistic estimation for expectations of random productsARC publishes an interim technical update introducing general methods for mechanistic estimation competitive with sampling, framed via expectations of random products (including examples such as random halfspace intersections, random #3-SAT, and random permanents).
Mechanistic estimation for wide random MLPsARC describes joint work studying how to estimate expected outputs of randomly initialized wide ReLU MLPs under Gaussian input without running the network, presenting theoretical and empirical comparisons versus Monte Carlo sampling.
AlgZoo: uninterpreted models with fewer than 1,500 parametersARC publishes AlgZoo, sharing small models (RNNs/transformers) trained on algorithmic tasks intended to serve as challenging test cases for interpretability visions; ARC describes the smallest models it believes it has substantially understood and larger ones it has not fully understood.
ARC authors describe the motivation and semi-formalization (Matching Sampling Principle) for an agenda focused on out-performing random sampling when estimating neural network outputs, tying it to preventing AI misalignment and describing related progress and recruiting links.