The UK AI Safety Institute (AISI) defines evaluations as a set of techniques for assessing an advanced AI system's capabilities, and government bodies like AISI conduct these both before and after deployment. Google DeepMind ties concrete mitigation action to results across four dangerous-capability domains, but AISI itself cautions that evaluation tasks can be gamed and do not certify any system as "safe."
What does "evaluations" mean, and what techniques are used to assess an AI system's capabilities?
The UK AI Safety Institute (AISI) defines evaluations as the practice of assessing an advanced AI system's capabilities using a range of different techniquesCITE:E1. AISI states: "By evaluations, we mean assessing the capabilities of an advanced AI system using a range of different techniques."CITE:E1 Under this definition, published on 2024-02-09, evaluation is not a single test but a toolkit — a standardized way of measuring capability across an AI systemCITE:E1.
What is the scope of AI evaluations? How does Google DeepMind categorize dangerous capabilities?
Google DeepMind's Frontier Safety Framework focuses its initial Critical Capability Level investigation on four dangerous-capability domains: autonomy, biosecurity, cybersecurity, and machine learning research and development (R&D)CITE:E2. Google DeepMind states its "initial set of Critical Capability Levels is based on investigation of four domains: autonomy, biosecurity, cybersecurity, and machine learning research and development (R&D)."CITE:E2 The framework, introduced on 2024-05-17, treats these four domains as the starting scope for identifying capabilities serious enough to warrant scrutinyCITE:E2.
Who is responsible for evaluating frontier AI systems? What role does government play?
AISI develops and conducts evaluations on advanced AI systems, assessing potential risks both before and after deploymentCITE:E3. AISI states it will "assess potential risks of new models before and after they are deployed, including by evaluating for potentially harmful capabilities."CITE:E3 This positions a government body, alongside the labs building the models, as an active evaluator across the deployment lifecycle — a role AISI laid out on 2024-02-09CITE:E3.
What evaluation tools and frameworks exist in practice? How does Inspect work?
Inspect is a framework for frontier AI evaluations developed jointly by the UK AI Security Institute (AISI) and Meridian LabsCITE:E4. Released in 2024 as an open framework, Inspect gives evaluators shared infrastructure for running the kind of capability tests that AISI and Google DeepMind describe in their respective frameworksCITE:E4.
How do evaluations feed into AI model release and governance decisions?
Google DeepMind commits to applying a mitigation plan once a model passes its early warning evaluations, weighing the overall balance of benefits and risks alongside the intended deployment contextsCITE:E5. Google DeepMind states this step "should take into account the overall balance of benefits and risks, and the intended deployment contexts."CITE:E5 Under the Frontier Safety Framework published on 2024-05-17, evaluation outcomes are the direct trigger for governance action, not a separate reporting exerciseCITE:E5.
What are the limits of evaluations? Why can't they capture true capability or guarantee safety?
AISI warns that future AI systems could be manipulated into deliberately failing specific evaluation tasks, making those tasks unrepresentative of a system's underlying capabilitiesCITE:E6. In a blog dated 2024-10-24, AISI wrote: "There are also concerns that AI systems might be manipulated to fail specific evaluation tasks in future, so that these tasks would not be representative of the underlying capabilities."CITE:E6 AISI adds, in its 2024-02-09 publication, that its evaluations "are not comprehensive assessments of an AI system's safety, and the goal is not to designate any system as 'safe.'"CITE:E7
Source-by-source summary
| Source | Entity | Publication Date | Core Claim |
|---|
| Approach to evaluations | UK AI Safety Institute (AISI) | 2024-02-09 | Defines evaluations as multi-technique capability assessment; AISI evaluates before/after deployment; evaluations are not safety certificationCITE:E1CITE:E3CITE:E7 |
| Frontier Safety Framework | Google DeepMind | 2024-05-17 | Four dangerous-capability domains; passing early warning evaluations triggers a mitigation planCITE:E2CITE:E5 |
| Inspect framework | UK AI Security Institute (AISI) & Meridian Labs | 2024 | Open framework for running frontier AI evaluationsCITE:E4 |
| Early lessons from evaluating frontier AI systems | UK AI Safety Institute (AISI) | 2024-10-24 | Evaluation tasks can be gamed, undermining their representativenessCITE:E6 |
Taken together, these four publications describe a closed loop: AISI's technique-based definition of evaluationsCITE:E1 supplies the method, Google DeepMind's four dangerous-capability domainsCITE:E2 supply the scope, and passing those evaluations is what triggers a mitigation plan under DeepMind's frameworkCITE:E5. But the same institute that evaluates models before and after deploymentCITE:E3 also says its own evaluation tasks may not represent a system's true underlying capabilities if the system is manipulated to fail themCITE:E6, and states plainly that passing an evaluation does not mean a system is designated "safe"CITE:E7. The governance action described by Google DeepMind therefore rests on a measurement tool that its own operator, AISI, says cannot yet guarantee what it measures.
Author's Take・林紀旭 James Lin
The most telling detail here isn't the four-domain taxonomy itself — autonomy, biosecurity, cybersecurity, ML R&D — it's that the body running the evaluations is the same one flagging their weakness. AISI ties evaluations to real consequences, since Google DeepMind's framework triggers an actual mitigation plan once a model clears an early-warning evaluation, yet AISI also warns that systems could be manipulated to fail those same tasks on purpose, and states outright that passing does not mean "safe." That's not a minor caveat sitting next to the governance mechanism — it's a crack in the mechanism's own foundation. The metric worth watching next is whether AISI or Google DeepMind publish anything about evaluation robustness itself — evidence that a capability score resisted deliberate gaming — rather than just the capability scores against the four domains or tools like Inspect. Until that shows up, every mitigation decision keyed to an evaluation result is only as trustworthy as AISI's own caveat allows it to be.