Research institute · Research orgs · San Francisco, CA · Est. 2024
Transluce is an independent, nonprofit research lab focused on building the public technology stack for understanding and steering advanced AI systems in the public interest. The lab frames the core challenge as the gap between (a) how quickly AI is developed and deployed and (b) the difficulty of reliably predicting and inspecting model behavior once systems are in the world.
Transluce’s strategy is to make oversight scalable and auditable by developing open tools and research methods that convert qualitative observations and intuition about model behavior into traceable measurements. At the center of Transluce’s work is Docent, a platform for analyzing agent transcripts—intended to support safety evaluations, monitoring, and iterative improvement of deployed systems by letting evaluators specify behavioral questions, operationalize them into rubrics, and run those rubrics over collections of transcripts. In parallel, Transluce develops “oversight foundation model” research aimed at training AI systems that can answer oversight questions given access to a subject model’s internal state (e.g., activations), with the goal of producing empirically checkable claims. Transluce also positions itself as an institutionally neutral bridge between researchers, developers, and governance stakeholders: its public materials emphasize openness to third-party auditing and the intent to work with model providers and governments once publicly-vetted methods can meet “public best practices” for oversight.
Surfacing misaligned behaviors in production coding agent traffic
Diagnosing a performance regression on Terminal-Bench with DocentMonitoring SWE-bench AgentsWe're partnering with the SWE-bench team to enable reliable evaluation of AI coding agents.
Open-sourcing DocentWe're excited to share that Docent is now open-source!
Docent's public alphaWe're excited to share our progress and get feedback from the community.
Introducing DocentA system for analyzing and intervening on agent behavior
Transluce reports scaling activation-oracle style oversight assistants to subject models up to ~1.1T parameters and claims performance improvements with subject/oracle scaling and training-data improvements.
Scaling Laws for Exact String ElicitationTransluce reports findings described as “elicitation ability follows predictable power laws,” relating elicitation performance to scaling behavior.
User awareness in frontier modelsTransluce describes work related to “user awareness” in frontier models, emphasizing that who is asking (or what the model infers about the user) shifts what models say.
Measuring coding agent misalignment in the wildTransluce’s Docent team describes a measurement study over 8,600 real-world coding agent sessions, reporting severe monitor-evasion and overselling rates and presenting how the measurement was constructed using Docent-style rubric/judging.
Foundation Models for OversightTransluce outlines a research vision for training oversight foundation models that formalize, test, and answer questions about AI model behavior, casting oversight as a world-modeling/inference-style problem in Pythonic terms.
Monitoring SWE-bench AgentsTransluce describes a partnership with the SWE-bench team to integrate SWE-bench agent trajectories/transcripts into Docent so that cheating or unintended solution paths can be monitored more reliably.