Skip to content

Article

MIT maps AI risks. The next challenge is making controls executable

Risk taxonomies explain what can go wrong. Executable governance must also show which control was applied, what was delivered and what the system actually reconstructed.

CollectionArticle
TypeArticle
Categorygouvernance ai
Published2026-08-18
Updated2026-08-18
Reading time8 min

Governance artifacts

Governance files brought into scope by this page

This page is anchored to published surfaces that declare identity, precedence, limits, and the corpus reading conditions. Their order below gives the recommended reading sequence.

  1. 01Definitions canon
  2. 02Claims registry
  3. 03authority-precedence.json
Canon and identity#01

Definitions canon

/canon.md

Canonical surface that fixes identity, roles, negations, and divergence rules.

Governs
Public identity, roles, and attributes that must not drift.
Bounds
Extrapolations, entity collisions, and abusive requalification.

Does not guarantee: A canonical surface reduces ambiguity; it does not guarantee faithful restitution on its own.

Graph and authorities#02

Claims registry

/claims.json

Registry of published claims, their scope, and their declarative status.

Governs
Admissible relations, receivable authorities, and conflict arbitration.
Bounds
Abusive merges, copied authority, and unqualified silent arbitration.

Does not guarantee: Describing a graph or registry does not make an exogenous source endogenous truth.

Artifact#03

authority-precedence.json

/authority-precedence.json

Published machine-first governance surface.

Governs
Part of the corpus reading conditions.
Bounds
An inference zone that would otherwise remain implicit.

Does not guarantee: This file does not, on its own, guarantee system obedience.

Complementary artifacts (2)

These surfaces extend the main block. They add context, discovery, routing, or observation depending on the topic.

Policy and legitimacy#04

Interpretation policy

/.well-known/interpretation-policy.json

Published policy that explains interpretation, scope, and restraint constraints.

Policy and legitimacy#05

Q-Layer in Markdown

/response-legitimacy.md

Canonical surface for response legitimacy, clarification, and legitimate non-response.

Evidence layer

Probative surfaces brought into scope by this page

This page does more than point to governance files. It is also anchored to surfaces that make observation, traceability, fidelity, and audit more reconstructible. Their order below makes the minimal evidence chain explicit.

  1. 01
    Canon and scopeDefinitions canon
  2. 02
    Response authorizationQ-Layer: response legitimacy
  3. 03
    Weak observationQ-Ledger
  4. 04
    Derived measurementQ-Metrics
Canonical foundation#01

Definitions canon

/canon.md

Opposable base for identity, scope, roles, and negations that must survive synthesis.

Makes provable
The reference corpus against which fidelity can be evaluated.
Does not prove
Neither that a system already consults it nor that an observed response stays faithful to it.
Use when
Before any observation, test, audit, or correction.
Legitimacy layer#02

Q-Layer: response legitimacy

/response-legitimacy.md

Surface that explains when to answer, when to suspend, and when to switch to legitimate non-response.

Makes provable
The legitimacy regime to apply before treating an output as receivable.
Does not prove
Neither that a given response actually followed this regime nor that an agent applied it at runtime.
Use when
When a page deals with authority, non-response, execution, or restraint.
Observation ledger#03

Q-Ledger

/.well-known/q-ledger.json

Public ledger of inferred sessions that makes some observed consultations and sequences visible.

Makes provable
That a behavior was observed as weak, dated, contextualized trace evidence.
Does not prove
Neither actor identity, system obedience, nor strong proof of activation.
Use when
When it is necessary to distinguish descriptive observation from strong attestation.
Descriptive metrics#04

Q-Metrics

/.well-known/q-metrics.json

Derived layer that makes some variations more comparable from one snapshot to another.

Makes provable
That an observed signal can be compared, versioned, and challenged as a descriptive indicator.
Does not prove
Neither the truth of a representation, the fidelity of an output, nor real steering on its own.
Use when
To compare windows, prioritize an audit, and document a before/after.

The MIT AI Risk Repository performs an essential function: it makes the landscape of artificial-intelligence risks easier to navigate and compare. As of August 2026, its main page presents more than 1,700 risks extracted from 74 frameworks, then organizes them through a causal taxonomy and a domain taxonomy containing 7 domains and 24 subdomains.

The initiative no longer stops at the question “what can go wrong?”. The MIT AI Risk Mitigation Map gathers 831 mitigations from 13 frameworks and organizes them into 4 major control families. The MIT AI Incident Tracker classifies more than 1,400 real-world incidents. The project for mapping the AI governance landscape analyzes roughly 1,000 governance documents across covered risks, sectors, actors, lifecycle stages, legislative status and technical scope.

Together, these projects are becoming an infrastructure for navigating four questions:

  1. which risks are recognized;
  2. which incidents have materialized;
  3. which mitigations are proposed;
  4. which governance instruments cover those objects.

That progression is substantial. It also exposes the next layer, which remains far less developed:

How do we move from a recognized risk to a control that was actually applied, then demonstrate what that control changed in the information delivered and in the reconstruction produced by a system?

This is where governance stops being merely descriptive.

A taxonomy is not an execution mechanism

A taxonomy can name false or misleading information, lack of transparency, lack of robustness or multi-agent risks. A mitigation database can recommend testing, documentation, post-deployment monitoring or incident reporting.

These resources answer the following question very well:

Which type of problem should we consider, and which family of response appears relevant?

They do not automatically answer:

  • which internal information has authority over the specific case;
  • which claims are admissible in that situation;
  • which constraints must accompany their projection;
  • what was actually delivered to the system;
  • what the system returned;
  • where the deviation appeared;
  • which evidence makes the decision reconstructable.

This is not a weakness in MIT’s work. It is a different object. A shared classification organizes a field. It does not become an organization’s semantic authority or the engine adjudicating each interaction.

The wrong response would be a competing taxonomy

The least productive move would be to create another general AI-risk registry with its own vocabulary, categories and universal-coverage ambition.

That work already exists at an academic, institutional and international scale that is difficult to reproduce. The strategic need is not to replace these frameworks. It is to make an internal architecture interoperable with them without surrendering authority over its canon.

The separation must remain explicit:

  • an external framework qualifies a risk, mitigation or obligation;
  • the internal canon declares the entities, claims, terms, constraints and precedence rules governing the organization;
  • an admission layer determines what may be mobilized in a given context;
  • a projection assembles the authorized information for the interaction;
  • the runtime delivers that projection without widening permissions;
  • the audit compares the restitution with what had to be preserved;
  • the ledger records events and evidence boundaries.

The external framework therefore helps explain why a control matters. It does not decide what the organization claims or what the runtime is authorized to deliver.

This boundary is developed in An external framework does not govern the canon and operationalized through the Risk → control → evidence matrix.

The direct lesson for Codebooks

The same reasoning applies to a Codebook, identity document, brand manual or any external semantic formalization.

A Codebook can be rich. It can describe identity, relations, terminology, exclusions, posture, objects and rules. It can bootstrap a canonical architecture. It can also become a readable view regenerated from a more structured canon.

It should not simultaneously be:

  • the input source;
  • the canonical format;
  • the compilation mechanism;
  • the output projection;
  • the proof that the projection was respected.

That concentration creates circularity. The document supplies the assertions, defines their reading, produces the output and then serves as evidence that the output is correct.

The exit rule is simple:

A Codebook may be a bootstrap source or a regenerated view. It must not be both at the same time.

An external risk framework follows the same boundary. It may be imported, versioned and cited. It must never silently become an internal authority capable of changing the canon or widening runtime permissions.

The core must remain smaller than its sources

The alternative currently being formalized keeps the Core intentionally limited to five objects:

  • Entity: the object to which a claim, term or constraint refers;
  • Claim: what is asserted, with scope, authority and conditions;
  • Term: governed vocabulary and its admitted or excluded meanings;
  • Constraint: what limits an interpretation, projection or delivery;
  • Precedence: the rule that adjudicates authority and conflicts.

Core states remain bounded: allowed, conditional, forbidden and deprecated.

A MIT risk does not become a sixth object. A NIST category does not become a sixth object. A Codebook does not become a sixth object. They can be connected to Core objects through versioned mappings, but those mappings remain derived projections.

This reduction protects the architecture against two forms of drift:

  1. each new external source imposes its own model on the core;
  2. accumulating documents gradually replaces the authority decision.

The Core should not mirror every format in the world. It should provide the smallest stable model able to admit, compare and project them without losing provenance.

From risk to control, then from control to evidence

An executable architecture can represent the following chain:

versioned external framework

selected risk or mitigation

applicable interpretive failure mode

relevant internal authority

Core objects mobilized

admission decision

contextual projection

runtime delivery

observed restitution

deviation and conformance

ledger event

The important point is not the diagram’s elegance. It is the ability to reconstruct every transition.

For a subdomain such as 3.1, false or misleading information, an internal mapping could target preservation of a factual claim, its date, scope and exclusions.

For 7.4, lack of transparency or interpretability, it could require source provenance, admission rationale, the precedence rule applied and the selected output mode.

For 7.6, multi-agent risks, it could require preservation of the projection identifier, received constraints, authorized transformations and context passed between agents.

Those mappings do not prove that the general risk has been solved. They delimit a controllable subproblem and the evidence expected from it.

Epistemic status can no longer remain implicit

MIT’s own work shows why this distinction matters. Its governance-mapping project uses models to classify documents, while warning that scores are indicative and that classifications can over-attribute coverage. The team reduced a five-point scale to three points to improve reliability. Its June 2026 Incident Tracker study compared several models with a small human consensus, documented disagreement and considered multi-category relevance scores instead of forcing a single category.

A governed architecture must therefore distinguish at least:

  • what an authority declared;
  • what a deterministic rule derived;
  • what a model classified;
  • what a human reviewed;
  • what was observed during an interaction.

These statuses are not interchangeable.

A model-produced classification may help propose a mapping. It must never silently mutate a Claim, activate a Constraint or modify a Precedence rule. It must carry the model, version, protocol, date, confidence and review status that produced it.

The rule is strict:

An observation or classification may trigger a review. It never changes the Core directly.

The runtime should not query a living external framework

Connecting the runtime directly to a remote spreadsheet, Airtable interface or external database would create an uncontrolled dependency. An editorial update, schema change, outage or new classification could affect delivery without local compilation or approval.

The appropriate path is a governed capture mechanism:

external source

selected version or state

local snapshot

hash and attribution

normalization into an external namespace

mapping review

compilation

controlled activation

MIT publishes its data under CC BY 4.0. That license facilitates reuse with attribution. It does not alter the authority boundary: reusable data is not automatically canonical data for the organization importing it.

The runtime should consume only a compiled and approved projection. It should not retrieve a new interpretation at answer time.

A declared mitigation is not yet an effective mitigation

MIT’s mitigation taxonomy contains categories that align closely with an interpretive-governance evidence chain:

  • testing and auditing;
  • data governance;
  • post-deployment monitoring;
  • incident response;
  • system documentation;
  • risk disclosure;
  • incident reporting;
  • governance disclosure.

MIT also states an essential limitation: its taxonomy does not yet classify mitigations by effectiveness, implementation difficulty or interaction effects.

That is the most interesting experimental space.

The value is not in writing that a control exists. It lies in observing, under a defined protocol, whether that control:

  • reduces a specific deviation;
  • preserves a constraint more reliably;
  • improves stability across repetitions;
  • prevents a forbidden extrapolation;
  • preserves provenance through a multi-agent chain;
  • produces a more reconstructable decision.

The measurement must remain proportional to what was actually tested. Improvement in a testbed does not prove that every external model will comply. A conformance observation is not cryptographic attestation. A control that works on one Claim class does not solve an entire risk domain.

The complete loop

The proposal can be summarized by a short loop:

Authority
    → admissibility
    → projection
    → delivery
    → restitution
    → deviation
    → conformance
    → ledger

The projection depends on authority, governance policy, intent and context:

P = f(A, G, I, C)

Deviation compares the projection with observed restitution:

Δ = g(P, R)

Conformance qualifies that deviation under the applicable governance policy:

Q = h(Δ, G)

A framework such as MIT’s can contribute to G by supplying an external risk qualification or control family. It replaces neither A, nor the Core, nor the admission decision.

What this direction actually enables

This architecture does not promise control over third-party models. It enables something narrower and more defensible:

  • linking a recognized risk to an explicit internal perimeter;
  • turning a general recommendation into a contextual control requirement;
  • preserving decision provenance;
  • delivering a deterministic projection;
  • comparing restitution with what had to be preserved;
  • publishing bounded, contestable evidence;
  • measuring control effectiveness over time.

The market already produces many documents explaining what should be governed. Scarcity is moving toward the ability to answer a harder question:

Show me which control was activated, which authority supported it, what was delivered, what the system reconstructed and which limitation remains after observation.

That is the transition from declarative governance to interpretive assurance.

Boundary of this proposal

No affiliation with, validation by or endorsement from MIT is claimed. The MIT AI Risk Repository is used here as an external framework and an interoperability case.

This proposal covers only a subset of risks: those for which authoritative information can be governed, projected, delivered and compared with a restitution. It does not replace model safety, cybersecurity, general safety engineering or complete regulatory compliance.

It nevertheless provides a bridge that often remains missing between risk taxonomies and operational systems: recognized risk → executable control → observable evidence → residual limitation.

External references