AI safety is the technical and policy field concerned with ensuring artificial intelligence systems behave reliably, stay under meaningful human control, and do not cause serious harm, particularly as their capabilities approach or exceed human-level performance in specific domains. It covers core technical problems including alignment (making a system pursue the goals its developers actually intended), robustness (resilience to errors, adversarial inputs, and unfamiliar situations), interpretability (understanding why a model produced a given output), and oversight (the ability to monitor, correct, or halt a system). For governance teams, AI safety matters because it supplies the evaluation methods, red-teaming practices, and risk thresholds that underpin frontier-model release decisions, vendor due diligence, and regulatory obligations for high-risk AI. The field has become increasingly institutionalised through government bodies and voluntary industry commitments, but several of the highest-profile national institutes were renamed and re-scoped in 2025, shifting emphasis from broad "safety" toward security and standards, a change governance practitioners should track closely.
Run the free AI Health CheckAI Safety, the field concerned with ensuring AI systems behave reliably and as intended, particularly as their capabilities approach or exceed human-level performance in defined domains.
AI safety as a technical field focuses on alignment (does the AI pursue the goals we actually want), robustness (does it behave reliably under unusual inputs or adversarial pressure), interpretability (can we understand what it is doing), and oversight (can we maintain control). Frontier AI safety is increasingly institutionalised through national bodies, including AI safety institutes in Japan and Singapore, Australia's now-operational AI Safety Institute (backed by around AUD 29.9 million in funding, with Dr Kate Conroy appointed inaugural General Manager in May 2026), the UK's AI Security Institute (renamed from the AI Safety Institute on 14 February 2025, retaining the AISI abbreviation but shifting focus toward crime and national-security threats), and the US Center for AI Standards and Innovation (renamed from the US AI Safety Institute on 3 June 2025, refocused on voluntary standards and national-security-relevant capability evaluation). Frontier labs also publish their own voluntary scaling frameworks, such as Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework, and Google DeepMind's Frontier Safety Framework, with industry bodies like the Frontier Model Forum promoting shared practice across them rather than issuing any single lab's framework.
Source: UK AI Security Institute (AISI); US Center for AI Standards and Innovation (CAISI); Australian AI Safety Institute; Anthropic Responsible Scaling Policy; Frontier Model Forum
Alignment
Ensuring a system's actual objectives and behaviour match what its developers and users intended, including in novel situations the system wasn't explicitly trained for.
Robustness
Performance that holds up under adversarial inputs, distribution shift, and edge cases, rather than failing silently or unpredictably in production.
Interpretability
Methods for understanding a model's internal reasoning well enough to explain, predict, or audit its outputs, increasingly relevant to explainability obligations under emerging AI regulation.
Oversight and control
Human-in-the-loop review, monitoring, and the practical ability to intervene, pause, or shut down a system before harm compounds.
Evaluations and red-teaming
Structured testing, capability evaluations, dangerous-capability testing, adversarial probing, conducted before a model is released or deployed.
Misuse prevention
Safeguards against a system being deliberately used to cause harm, for example in cyberattacks, biological or chemical weapons uplift, or large-scale fraud.
Governments moved quickly after the 2023 Bletchley Park AI Safety Summit to stand up technical institutes for evaluating frontier AI models. By 2025, several of the most prominent ones had been renamed and re-scoped, which matters for anyone citing them in governance documentation.
United Kingdom, AI Security Institute (AISI)
Launched in November 2023 as the AI Safety Institute following the Bletchley Park summit. Renamed the AI Security Institute on 14 February 2025 under the Department for Science, Innovation and Technology, shifting its stated focus toward crime, national security threats, and cyberattacks while keeping the AISI abbreviation. Site: aisi.gov.uk.
United States, Center for AI Standards and Innovation (CAISI)
The US AI Safety Institute, housed at NIST, was renamed the Center for AI Standards and Innovation on 4 June 2025. CAISI's remit centres on voluntary standards development, evaluating national-security-relevant AI capabilities (cyber, bio, chemical), and assessing foreign and adversary AI systems, following the current administration's January 2025 revocation of the prior administration's Executive Order 14110.
Australia, AI Safety Institute
Australia's AI Safety Institute has launched and is now operational, backed by roughly AUD 29.9 million in funding. Dr Kate Conroy was appointed as its inaugural General Manager in May 2026, with a mandate covering monitoring frontier AI capabilities, assessing risks, and sharing findings with domestic regulators and international partners. Unlike the UK and US bodies, Australia's institute has retained the "safety" name.
International Network of AI Safety Institutes
Launched November 2024 with Australia, Canada, the European Commission, France, Japan, Kenya, South Korea, Singapore, the UK, and the US as founding members, coordinating joint model testing and shared risk-assessment approaches across jurisdictions, even as individual members' domestic institutes have been renamed.
Frontier AI developers also publish their own internal safety frameworks, which set capability thresholds that trigger additional safeguards before a more capable model is trained or released. These go by different names at each lab, Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework, and Google DeepMind's Frontier Safety Framework are the best known, but share a common structure of risk tiers tied to escalating safeguards.
The Frontier Model Forum, an industry body founded in 2023 by Anthropic, Google DeepMind, Microsoft, and OpenAI, promotes shared safety research and best practices across labs, though the underlying frameworks remain company-specific commitments rather than a jointly issued standard. Because these commitments are voluntary and self-governed, they are frequently referenced in AI governance and procurement due diligence as one signal of a vendor's safety maturity, but they carry no independent enforcement mechanism of their own.
These three terms are often used interchangeably, but they answer different questions. AI safety asks a technical question, does the system behave as intended, and can we detect and correct it when it doesn't? AI ethics asks a normative question, what values, fairness standards, and societal impacts should the system embody? AI governance is the broader organisational layer, the policies, oversight structures, and compliance processes that translate safety evaluations and ethical principles into enforceable practice inside a company or under a regulatory regime. In practice, a mature AI governance program treats safety evaluations and red-teaming findings as key inputs to its risk register, rather than as a substitute for governance itself.
What is AI safety in simple terms?
AI safety is the field of research and engineering focused on making sure AI systems do what they're intended to do, don't cause unintended harm, and can be monitored or stopped if something goes wrong, especially as models become more capable.
What is the difference between AI safety and AI ethics?
AI safety is primarily a technical discipline asking whether a system behaves reliably and as intended. AI ethics is a normative discipline asking what values, fairness principles, and societal outcomes a system should reflect. The two overlap but are not the same field.
What happened to the UK AI Safety Institute?
It was renamed the AI Security Institute on 14 February 2025, still operating as UK AISI under the Department for Science, Innovation and Technology, with a stated shift toward security threats such as cyberattacks, fraud, and child exploitation rather than broader AI safety research.
Is the US AI Safety Institute still called that?
No. The US Department of Commerce renamed it the Center for AI Standards and Innovation (CAISI) on 4 June 2025. CAISI sits within NIST and focuses on voluntary standards, national-security-relevant capability evaluations, and assessment of foreign AI systems.
Does Australia have an AI Safety Institute?
Yes, Australia's AI Safety Institute launched in early 2026, backed by roughly AUD 29.9 million in funding, with Dr Kate Conroy appointed as its inaugural General Manager in May 2026. Unlike the UK and US counterparts, it retained the "AI Safety Institute" name.
What is a Responsible Scaling Policy?
A Responsible Scaling Policy is Anthropic's internal framework that ties increasing model capability to escalating safety and security safeguards before training or deploying more powerful models. Other labs publish comparable frameworks under different names, such as OpenAI's Preparedness Framework and Google DeepMind's Frontier Safety Framework.
Last reviewed July 2026
This page is general information about What Is AI Safety?, not legal, regulatory, or professional advice, and does not capture every nuance or exception. Requirements change and can be fact-specific. Always verify against primary sources and your own qualified legal counsel before relying on it.