Academic center · Science policy · Stanford University (Stanford, California) · Est. 2021
The Stanford Center for Research on Foundation Models (CRFM) is an interdisciplinary initiative at Stanford’s Institute for Human-Centered Artificial Intelligence (HAI) focused on advancing research on foundation models—spanning technical evaluation and transparency, domain-specific model development, and research-to-policy analysis of societal impacts. CRFM brings together faculty, students, researchers, and engineers across more than a dozen Stanford departments to study how foundation models are built and deployed, and to develop open, reproducible tools and benchmarks intended to make model behavior more measurable and comparable over time.
CRFM’s technical work is strongly reflected in its evaluation infrastructure (notably HELM and its modality- and domain-specific variants), while its policy-facing work includes transparency assessment through the Foundation Model Transparency Index (FMTI). As an academic center within Stanford HAI, CRFM positions itself as a convening and research hub that can systematically characterize model capabilities and risks in ways that are difficult for closed, proprietary ecosystems to provide.
No indexed openings right now. Check the careers page ↗.
We present HELM Arabic Enterprise, a leaderboard for transparent, reproducible evaluation of large language models on Arabic-language benchmarks designed around enterprise use cases. The leaderboard was developed in collaboration with Arabic.AI and builds on the HELM evaluation m
HELM ArabicAs part of our efforts to better understand the multilingual capabilities of large language models (LLMs), we present HELM Arabic, a leaderboard for transparent and reproducible evaluation of LLMs on Arabic language benchmarks. This leaderboard was produced in collaboration with
HELM Long ContextReliable and Efficient Amortized Model-Based EvaluationTLDR: We enhance the reliability and efficiency of language model evaluation by introducing IRT-based adaptive testing, which has been integrated into the HELM framework.
Surprisingly Fast AI-Generated Kernels We Didn’t Mean to Publish (Yet)TL;DR
BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity SystemsHELM Capabilities: Evaluating LMs Capability by CapabilityGeneral-Purpose AI Needs Coordinated Flaw ReportingHELM Safety: Towards Standardized Safety Evaluations of Language Models*Work done while at Stanford CRFM
Advancing Customizable Benchmarking in HELM via Unitxt IntegrationThe Holistic Evaluation of Language Models (HELM) framework is an open source framework for reproducible and transparent benchmarking of language models that is widely adopted by academia and industry. To meet HELM users’ needs for more powerful benchmarking features, we are prou
CRFM describes an IRT/Rasch-model-based approach with adaptive testing integrated into HELM to reduce evaluation cost while preserving reliability for language-model evaluation.
Holistic Evaluation of Vision-Language Models (VHELM v1.0) releasedCRFM introduces VHELM v1.0, extending HELM to evaluate prominent vision-language models with released prompts and raw predictions, starting from three scenarios/datasets and six VLMs.
Foundation Model Transparency Index (October 18, 2023) announcedCRFM (with Stanford HAI coverage) introduces FMTI as an index rating the transparency of foundation model companies using 100 transparency indicators and shared scoring methodology.
Foundation Model Transparency Index (December 2025) — 2025 edition released with updated indicators and company scoringCRFM reports that the 2025 FMTI is the third edition, with updated indicators reflecting changes in the AI ecosystem, and publishes overall findings including a decline in mean scores versus 2024.
HELM Capabilities (HELM benchmark/leaderboard)CRFM introduces HELM Capabilities as a new leaderboard/benchmark intended to measure general capabilities with prompt-level transparency and reproducibility via HELM.
HELM leaderboards (framework and platform overview)CRFM’s HELM hub page consolidates multiple leaderboards, including capabilities, safety, modality-specific evaluations, and lightweight evaluation variants.