AIFEATURE

What AI Model Evaluations Actually Measure — And What They Can't

林紀旭 James LinEditor-in-Chief
Published · Updated
The UK AI Safety Institute (AISI) defines evaluations as a set of techniques for assessing an advanced AI system's capabilities, and government bodies like AISI conduct these both before and after deployment. Google DeepMind ties concrete mitigation action to results across four dangerous-capability domains, but AISI itself cautions that evaluation tasks can be gamed and do not certify any system as "safe."

What does "evaluations" mean, and what techniques are used to assess an AI system's capabilities?

The UK AI Safety Institute (AISI) defines evaluations as the practice of assessing an advanced AI system's capabilities using a range of different techniquesCITE:E1. AISI states: "By evaluations, we mean assessing the capabilities of an advanced AI system using a range of different techniques."CITE:E1 Under this definition, published on 2024-02-09, evaluation is not a single test but a toolkit — a standardized way of measuring capability across an AI systemCITE:E1.

What is the scope of AI evaluations? How does Google DeepMind categorize dangerous capabilities?

Google DeepMind's Frontier Safety Framework focuses its initial Critical Capability Level investigation on four dangerous-capability domains: autonomy, biosecurity, cybersecurity, and machine learning research and development (R&D)CITE:E2. Google DeepMind states its "initial set of Critical Capability Levels is based on investigation of four domains: autonomy, biosecurity, cybersecurity, and machine learning research and development (R&D)."CITE:E2 The framework, introduced on 2024-05-17, treats these four domains as the starting scope for identifying capabilities serious enough to warrant scrutinyCITE:E2.

Who is responsible for evaluating frontier AI systems? What role does government play?

AISI develops and conducts evaluations on advanced AI systems, assessing potential risks both before and after deploymentCITE:E3. AISI states it will "assess potential risks of new models before and after they are deployed, including by evaluating for potentially harmful capabilities."CITE:E3 This positions a government body, alongside the labs building the models, as an active evaluator across the deployment lifecycle — a role AISI laid out on 2024-02-09CITE:E3.

What evaluation tools and frameworks exist in practice? How does Inspect work?

Inspect is a framework for frontier AI evaluations developed jointly by the UK AI Security Institute (AISI) and Meridian LabsCITE:E4. Released in 2024 as an open framework, Inspect gives evaluators shared infrastructure for running the kind of capability tests that AISI and Google DeepMind describe in their respective frameworksCITE:E4.

How do evaluations feed into AI model release and governance decisions?

Google DeepMind commits to applying a mitigation plan once a model passes its early warning evaluations, weighing the overall balance of benefits and risks alongside the intended deployment contextsCITE:E5. Google DeepMind states this step "should take into account the overall balance of benefits and risks, and the intended deployment contexts."CITE:E5 Under the Frontier Safety Framework published on 2024-05-17, evaluation outcomes are the direct trigger for governance action, not a separate reporting exerciseCITE:E5.

What are the limits of evaluations? Why can't they capture true capability or guarantee safety?

AISI warns that future AI systems could be manipulated into deliberately failing specific evaluation tasks, making those tasks unrepresentative of a system's underlying capabilitiesCITE:E6. In a blog dated 2024-10-24, AISI wrote: "There are also concerns that AI systems might be manipulated to fail specific evaluation tasks in future, so that these tasks would not be representative of the underlying capabilities."CITE:E6 AISI adds, in its 2024-02-09 publication, that its evaluations "are not comprehensive assessments of an AI system's safety, and the goal is not to designate any system as 'safe.'"CITE:E7

Source-by-source summary

SourceEntityPublication DateCore Claim
Approach to evaluationsUK AI Safety Institute (AISI)2024-02-09Defines evaluations as multi-technique capability assessment; AISI evaluates before/after deployment; evaluations are not safety certificationCITE:E1CITE:E3CITE:E7
Frontier Safety FrameworkGoogle DeepMind2024-05-17Four dangerous-capability domains; passing early warning evaluations triggers a mitigation planCITE:E2CITE:E5
Inspect frameworkUK AI Security Institute (AISI) & Meridian Labs2024Open framework for running frontier AI evaluationsCITE:E4
Early lessons from evaluating frontier AI systemsUK AI Safety Institute (AISI)2024-10-24Evaluation tasks can be gamed, undermining their representativenessCITE:E6

Taken together, these four publications describe a closed loop: AISI's technique-based definition of evaluationsCITE:E1 supplies the method, Google DeepMind's four dangerous-capability domainsCITE:E2 supply the scope, and passing those evaluations is what triggers a mitigation plan under DeepMind's frameworkCITE:E5. But the same institute that evaluates models before and after deploymentCITE:E3 also says its own evaluation tasks may not represent a system's true underlying capabilities if the system is manipulated to fail themCITE:E6, and states plainly that passing an evaluation does not mean a system is designated "safe"CITE:E7. The governance action described by Google DeepMind therefore rests on a measurement tool that its own operator, AISI, says cannot yet guarantee what it measures.

📊 Evidence

FAQ

What does "evaluations" mean, and what techniques are used to assess an AI system's capabilities?

The UK AI Safety Institute (AISI) defines evaluations as the practice of assessing an advanced AI system's capabilities using a range of different techniquesCIT…

What is the scope of AI evaluations? How does Google DeepMind categorize dangerous capabilities?

Google DeepMind's Frontier Safety Framework focuses its initial Critical Capability Level investigation on four dangerous-capability domains: autonomy, biosecur…

Who is responsible for evaluating frontier AI systems? What role does government play?

AISI develops and conducts evaluations on advanced AI systems, assessing potential risks both before and after deploymentCITE:E3.

What evaluation tools and frameworks exist in practice? How does Inspect work?

Inspect is a framework for frontier AI evaluations developed jointly by the UK AI Security Institute (AISI) and Meridian LabsCITE:E4.

📎 Sources

  1. gov.uk
  2. deepmind.google
  3. inspect.aisi.org.uk
  4. aisi.gov.uk

Related data

Author's Take林紀旭 James Lin

The most telling detail here isn't the four-domain taxonomy itself — autonomy, biosecurity, cybersecurity, ML R&D — it's that the body running the evaluations is the same one flagging their weakness. AISI ties evaluations to real consequences, since Google DeepMind's framework triggers an actual mitigation plan once a model clears an early-warning evaluation, yet AISI also warns that systems could be manipulated to fail those same tasks on purpose, and states outright that passing does not mean "safe." That's not a minor caveat sitting next to the governance mechanism — it's a crack in the mechanism's own foundation. The metric worth watching next is whether AISI or Google DeepMind publish anything about evaluation robustness itself — evidence that a capability score resisted deliberate gaming — rather than just the capability scores against the four domains or tools like Inspect. Until that shows up, every mitigation decision keyed to an evaluation result is only as trustworthy as AISI's own caveat allows it to be.

林紀旭 James LinEditor-in-Chief

Related

BRIEF

GreenTrans Unveils GT5X, GT3X Quadruped Robots, Targets 100% Taiwan-Made Content by 2027

GreenTrans (綠捷), the robotics subsidiary of China Motor (中華車), unveiled quadruped robots GT5X and GT3X at SEMICON Taiwan 2026, targeting 100% Taiwan-made content by 2027. The robots combine an in-house-designed control unit and battery management system, NVIDIA's Jetson Orin and Isaac Lab platforms, and a new LFP battery developed with Formosa Smart Energy (台塑新智能), while GreenTrans's inspection robots are already deployed in semiconductor fabs.

EffectStory 編輯部 ·
BRIEF

Nvidia Confirms $12.93 Billion Acquisition of Hugging Face

Nvidia confirmed on September 3, 2026 that it agreed to buy Hugging Face for $12.93 billion, exactly $12,930,300,000, gaining the open-source AI hosting platform used by over 18 million developers. CEO Jensen Huang pledged the platform will stay open, with no Nvidia compute required to build on or deploy through it.

EffectStory 編輯部 ·
BRIEF

NVIDIA to Subscribe US$3.5 Billion of MediaTek's Record US$3.9 Billion Convertible Bond

NVIDIA will subscribe US$3.5 billion of MediaTek's US$3.9 billion offshore convertible bond offering, the largest such issuance in Taiwan's capital market history, deepening cooperation in AI infrastructure, edge AI computing, and automotive platforms while marking NVIDIA's first major investment in a Taiwanese company.

EffectStory 編輯部 ·