Standards consortium · Science policy · International coordination across member countries and institutions; UK AI Security Institute (AISI) acts as Network Coordinator in . ([gov.uk](https://www.gov.uk/government/news/efforts-to-share-best-practices-on-ai-measurement-and-evaluations-driven-forward-through-the-international-network-for-advanced-ai-measurement-evalua)) · Est. 2024
The International Network for Advanced AI Measurement, Evaluation and Science (formerly the International Network of AI Safety Institutes) is an international, government-led consortium that coordinates cross-border work on the science and practice of evaluating advanced AI capabilities. The Network’s stated focus is to strengthen shared measurement and evaluation approaches—aiming to improve trust in AI capability claims and to support confident, safer adoption across languages, industries, and jurisdictions.
In its recent outputs, the Network emphasizes evaluation principles such as defining clear objectives, improving transparency and (in principle) reproducibility, embedding quality assurance, and strengthening validity through uncertainty estimates and attention to external validity (including realistic context and multilingual/cultural factors). A key Network deliverable described by member governments and partners is the publication of consensus areas and open questions for evaluating AI capabilities, including unresolved issues around the use of “risk models,” prioritization under resource constraints, what evaluator information should be shared versus withheld, and how to evaluate AI systems more broadly than model weights alone. Institutionally, the Network is positioned as an international venue to align best practices between national AI safety measurement institutions; the UK AI Security Institute (AISI) serves as Network Coordinator in 2026 and leads efforts to turn shared learning into more detailed best-practice documentation.
Early learnings from red-teaming the internal monitors of frontier AI companies, and our perspectives on the open problems that remain.
Incident Report: unsanctioned agent behaviour during cyber testingDuring a routine cyber evaluation, AISI identified an incident in which AI agents took sustained, unsanctioned action directed at real people and organisations. We are disclosing what we found, what it means, and the actions now underway.
UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber CapabilitiesOur joint evaluation with CAISI finds Kimi K3 trails leading US frontier closed weight models on cyber capability.
Cheating behaviour in frontier model evaluationsWe find cheating behaviour in all of our cyber capability evaluations, and outline the implications as models grow more capable.
How Far Behind the Frontier are Leading Open Weight Models on Cyber?We evaluated the cyber capabilities of leading open and closed weight AI models, and found that recent open models GLM-5.2 and DeepSeek V4-Pro perform similarly to frontier closed models released 4 to 7 months before them – a narrower gap than the 6 to 10 months we measured throu
International evaluation best practice and open questions in AI measurementThe International Network for Advanced AI Measurement, Evaluation and Science convened in Seoul to continue outlining international best practice.
Deepening our partnership with the Australian AI Safety InstituteAn agreement between Institutes to collaborate on best practices in AI evaluation, and share research findings.
Finding Cloud Misconfigurations with Frontier AI: A Case StudyA cybersecurity exercise from AISI’s engineering team, using frontier models to test our research platform for misconfigurations.
UK-Germany Joint Statement on advanced AI safety and securityA joint statement by the UK and Germany on collaborating to ensure advanced AI is developed safely and its risks are rigorously understood and managed.
A Decision-Theoretic Formalisation of Steganography With Applications to LLM MonitoringAlignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentConsistency Training Can Entrench MisalignmentAISI’s blog post states that the Network gathered in Seoul to continue outlining international best practice, referencing completion of the Network’s first best-practice guidance document and highlighting themes around evaluation practices and open questions, including agentic evaluation challenges.
UK AISI coordinator role for the International NetworkA UK parliamentary written answer states that UK AISI is the coordinator of the International Network for Advanced AI Measurement, Evaluation and Science and that the Network brings together international partners to advance AI evaluation science and measurement.
International Network publishes consensus areas and open questions on best practices for automated AI evaluationsNIST reported that the Network published a list of key practices and open questions for measuring and evaluating AI capabilities, describing its membership and its focus on strengthening the science underpinning AI evaluation.
UK takes on Network Coordinator role as part of the Network’s transition/re-focusA UK government press release states that the International Network of AI Safety Institutes transitioned into the International Network for Advanced AI Measurement, Evaluation and Science; it also says the UK now takes on a Network Coordinator role.
International consensus and open questions in AI evaluationsAISI’s blog post explains that the Network re-focused its work from the prior International Network of AI Safety Institutes to strengthening the science underpinning AI evaluation; it also summarizes consensus areas and specific open questions discussed by Network members.