Governance artifacts
Governance files brought into scope by this page
This page is anchored to published surfaces that declare identity, precedence, limits, and the corpus reading conditions. Their order below gives the recommended reading sequence.
Definitions canon
/canon.md
Canonical surface that fixes identity, roles, negations, and divergence rules.
- Governs
- Public identity, roles, and attributes that must not drift.
- Bounds
- Extrapolations, entity collisions, and abusive requalification.
Does not guarantee: A canonical surface reduces ambiguity; it does not guarantee faithful restitution on its own.
Public AI manifest
/ai-manifest.json
Structured inventory of the surfaces, registries, and modules that extend the canonical entrypoint.
- Governs
- Access order across surfaces and initial precedence.
- Bounds
- Free readings that bypass the canon or the published order.
Does not guarantee: This surface publishes a reading order; it does not force execution or obedience.
Canonical AI entrypoint
/.well-known/ai-governance.json
Neutral entrypoint that declares the governance map, precedence chain, and the surfaces to read first.
- Governs
- Access order across surfaces and initial precedence.
- Bounds
- Free readings that bypass the canon or the published order.
Does not guarantee: This surface publishes a reading order; it does not force execution or obedience.
Complementary artifacts (2)
These surfaces extend the main block. They add context, discovery, routing, or observation depending on the topic.
Q-Ledger JSON
/.well-known/q-ledger.json
Machine-first journal of observations, baselines, and versioned gaps.
Q-Metrics JSON
/.well-known/q-metrics.json
Descriptive metrics surface for observing gaps, snapshots, and comparisons.
Evidence layer
Probative surfaces brought into scope by this page
This page does more than point to governance files. It is also anchored to surfaces that make observation, traceability, fidelity, and audit more reconstructible. Their order below makes the minimal evidence chain explicit.
- 01Canon and scopeDefinitions canon
- 02Response authorizationQ-Layer: response legitimacy
- 03Weak observationQ-Ledger
- 04Derived measurementQ-Metrics
Definitions canon
/canon.md
Opposable base for identity, scope, roles, and negations that must survive synthesis.
- Makes provable
- The reference corpus against which fidelity can be evaluated.
- Does not prove
- Neither that a system already consults it nor that an observed response stays faithful to it.
- Use when
- Before any observation, test, audit, or correction.
Q-Layer: response legitimacy
/response-legitimacy.md
Surface that explains when to answer, when to suspend, and when to switch to legitimate non-response.
- Makes provable
- The legitimacy regime to apply before treating an output as receivable.
- Does not prove
- Neither that a given response actually followed this regime nor that an agent applied it at runtime.
- Use when
- When a page deals with authority, non-response, execution, or restraint.
Q-Ledger
/.well-known/q-ledger.json
Public ledger of inferred sessions that makes some observed consultations and sequences visible.
- Makes provable
- That a behavior was observed as weak, dated, contextualized trace evidence.
- Does not prove
- Neither actor identity, system obedience, nor strong proof of activation.
- Use when
- When it is necessary to distinguish descriptive observation from strong attestation.
Q-Metrics
/.well-known/q-metrics.json
Derived layer that makes some variations more comparable from one snapshot to another.
- Makes provable
- That an observed signal can be compared, versioned, and challenged as a descriptive indicator.
- Does not prove
- Neither the truth of a representation, the fidelity of an output, nor real steering on its own.
- Use when
- To compare windows, prioritize an audit, and document a before/after.
Complementary probative surfaces (1)
These artifacts extend the main chain. They help qualify an audit, an evidence level, a citation, or a version trajectory.
Q-Attest protocol
/.well-known/q-attest-protocol.md
Optional specification that cleanly separates inferred sessions from validated attestations.
Causal mesh
CCL chain declared for this surface
This block separates the triggering situation, latent need, canonical surfaces, anti-fusion clarifications, evidence and declared bridges that govern the causal reading.
The causal chain declares situated relevance. It does not create a promise, result guarantee, implicit offer, or citation obligation.
Triggering situation
An intervention is followed by a change in mentions, citations, source selection, or reconstruction fidelity.
Problem or risk
Without an observation unit, estimand, counterfactual, and exogenous-event log, the change cannot be properly attributed.
Latent need
An evaluation design that separates open-visibility effects from controlled-fidelity effects and qualifies the actual level of evidence.
Intended consequence
Make performance claims contestable, reproducible, and proportional to the identification design.
Declared service bridge
The protocol can frame an audit, experiment, or program reassessment without guaranteeing a visibility gain or positive effect.
Non-derivation boundaries
- Each study must define one primary estimand before post-intervention observation.
- Controlled conditions and the open Web produce different types of evidence.
- Model and answer-surface changes must be logged as exogenous events.
- A non-identifiable verdict is valid and must remain publishable.
Governing doctrine
Without a counterfactual, an increase in mentions or citations cannot be attributed to a GEO intervention.
The GEO market now produces a simple and costly confusion: descriptive indicators are treated as steering instruments.
Consequence frameworks
Protocol for measuring LLM perception drift from a baseline, a canon, multi-model outputs, and a documented canon-output gap.
Framework for building an observability layer around interpretive stability, using metrics, logs, and evidence without confusing observation with attestation.
Interpretation integrity audit is the disciplined process by which a declared canon is compared with real model outputs under bounded conditions.
Evidence surfaces
Canonical definition of proof of fidelity: the minimum evidence required to show that an AI output remains faithful to the canon rather than merely plausible.
AI answer auditability requires tracing inference, implicit authority, legitimate refusals, and unknowns across an interpreted system.
Next reading routes
Without a counterfactual, an increase in mentions or citations cannot be attributed to a GEO intervention.
The GEO market now produces a simple and costly confusion: descriptive indicators are treated as steering instruments.
Protocol for measuring LLM perception drift from a baseline, a canon, multi-model outputs, and a documented canon-output gap.
Machine-readable artifacts
Evidence artifacts
Forbidden derivations
before_after_as_causalityvisibility_effect_as_fidelity_effectcontrolled_effect_as_open_web_effectaggregate_score_as_estimandpost_hoc_prompt_selectionomitted_null_result
Protocol status
This document proposes an evaluation method. It is neither a performance promise, a universal certification, nor an adopted industry standard.
Its purpose is narrower: to prevent an observed change after a GEO intervention from being automatically presented as an effect caused by that intervention.
The protocol operationalizes Causal attribution in GEO requires a counterfactual and complements GEO metrics do not govern representation.
1. Begin by separating two causal questions
A study must never merge the following two objects.
1.1 Effect on open-Web visibility
The first question is:
To what extent did the intervention change the probability that an entity, domain, page, or source appears in generative answers observed on the open Web?
Possible outcomes include:
- entity presence rate;
- domain citation rate;
- selection of a canonical source;
- share of voice within an explicitly defined panel;
- presence within an intent class;
- order or prominence of appearance.
This object is heavily exposed to exogenous change: interface adoption, answer-surface rollouts, model changes, citation policies, query composition, competition, seasonality, and personalization.
1.2 Effect on reconstruction fidelity in a controlled environment
The second question is:
To what extent do a documentary layer, governance files, or a runtime improve reconstruction fidelity when introduced under controlled conditions?
Possible outcomes include:
- accuracy of critical attributes;
- preservation of exclusions;
- preservation of relationships between entities;
- compliance with source classes;
- reduction in unauthorized inferences;
- stability across wording, languages, and repetitions;
- critical-error rate.
An improvement measured here is evidence about fidelity within the tested setup. It does not automatically prove an increase in visibility in ChatGPT, Gemini, Perplexity, Google AI Overviews, or any other open surface.
2. Define the estimand before measurement
The estimand is the precise effect the study seeks to estimate. It must be formulated before post-intervention observation.
A defensible formulation specifies at least:
- the unit affected by the intervention;
- the intervention and its version;
- the primary outcome;
- the no-intervention comparator;
- the time window;
- systems and models;
- languages and regions;
- the query or prompt panel;
- the population to which the conclusion may generalize.
Open-Web example:
The eight-week average effect of publishing a restructured canonical corpus on the domain citation rate within a preregistered panel of French-language commercial queries, compared with a group of similar unchanged pages.
Controlled-fidelity example:
The average effect of adding versioned governance files and runtime on reconstruction-fidelity scores, while holding the source corpus, prompts, models, parameters, and repetitions constant.
A study that merely announces “improve AI visibility” does not define an estimand. It defines a commercial ambition.
3. Fix the observation unit
The observation must be represented as a cube rather than an opaque average:
E × I × P × S × A × L × T
where:
- E is the entity, page, or exposed unit;
- I is the intervention and its version;
- P is the prompt, query, or intent family;
- S is the system, product, and model;
- A is the access channel: interface, API, browsing, retrieval, or agent;
- L is the language, region, and local settings;
- T is time, including date, hour, and observation window.
An “AI share of voice” metric that does not expose these dimensions cannot support replication, diagnosis, or serious attribution.
4. Preregister the intervention and outcomes
Before the post-intervention window begins, record:
- the primary hypothesis;
- the primary estimand;
- secondary outcomes;
- the full prompt panel or the rules that generate it;
- treated and control units;
- the collection window;
- planned exclusions;
- scoring method;
- sensitivity analyses;
- verdict criteria.
Preregistration prevents after-the-fact selection of prompts, models, or dates that produce the most favorable curve.
Any later change must be logged as an amendment and separated from confirmatory analysis.
5. Freeze and log the intervention
A GEO intervention is not a label. It must be described as a versioned set of observable actions.
The log must include:
- URLs created, changed, merged, redirected, or removed;
- content, architecture, entity, and structured-data changes;
- governance files or machine surfaces actually published;
- dates of availability, indexability, and propagation;
- changes to internal linking and canonicals;
- known external interventions;
- concurrent campaigns, public relations, or awareness gains;
- versions of the tested runtime or corpus.
Without this log, the “treatment” remains undefined and its effect cannot be isolated.
6. Establish a usable baseline
A baseline is not a number captured the day before launch. It must document pre-intervention level, trend, and variability.
It should ideally cover:
- multiple observation periods;
- the same units, prompts, systems, languages, and channels as the post period;
- enough repetitions to estimate dispersion;
- known model incidents and changes;
- behavior of control units;
- relevant seasonal trends.
A short baseline may serve as a descriptive snapshot. It does not magically become a counterfactual.
7. Construct the counterfactual
The counterfactual represents what would plausibly have happened without the intervention. No design is perfect, but some are much more credible than a simple before-and-after comparison.
7.1 Random assignment with a control group
When possible, randomly allocate similar pages, entities, markets, corpora, or units to treatment and control.
This is the strongest design for balancing observed and unobserved factors, provided the sample is large enough, execution is stable, and contamination between groups is limited.
7.2 Staggered rollout and difference-in-differences
When the intervention is rolled out at different times across comparable units, measure the treated group’s change relative to the control group’s change:
DiD effect = (Y treated, post − Y treated, pre) − (Y control, post − Y control, pre)
This design notably requires a credible parallel-trends assumption before treatment and careful handling of heterogeneous effects across treatment dates.
7.3 Matched or synthetic control
When few units are treated, construct a comparator from untreated units that resemble the pre-intervention behavior.
A synthetic control can combine several control units to reproduce the treated unit’s prior trajectory. Pre-intervention fit and placebo tests then become essential.
7.4 Controlled interrupted time series
When the intervention occurs at a clear date and many time observations are available, estimate changes in level and slope, ideally relative to an unexposed comparison series.
This design must account for autocorrelation, seasonality, concurrent interventions, and events coinciding with the interruption.
7.5 Paired test under controlled conditions
To assess interpretive governance, compare the same tasks in two conditions:
- corpus or runtime without the governance layer;
- corpus or runtime with the versioned governance layer.
Hold models, prompts, parameters, tools, source documents, languages, repetitions, and evaluation rules constant.
The paired fidelity effect can be expressed as:
ΔF = F governed − F ungoverned
This result measures an effect in the testbed. It must not be rewritten as a “market visibility gain.”
8. Preserve a representative prompt panel
The panel must be defined through intent families, not an opportunistic list of favorable wording.
Depending on the object, it should cover:
- direct identification;
- category definition;
- unbranded solution search;
- comparison;
- attribute verification;
- scope and exclusions;
- history and freshness;
- reputation and risk;
- ambiguous or adversarial variants;
- relevant languages and regions.
Document the panel’s origin: search data, interviews, usage logs, commercial queries, user research, or expert construction.
The protocol must preserve the denominator. Publishing “30 citations” without the number of eligible answers, tested queries, and access conditions supports no reliable interpretation.
9. Capture outputs and conditions
For every observation, retain at least:
- study identifier;
- unit identifier;
- treated or control status;
- intervention version;
- exact prompt and intent family;
- system, product, model, and observable version;
- date, time, language, region, and access channel;
- browsing or retrieval status;
- session, memory, and personalization state when known;
- complete answer;
- citations, URLs, order, and associated excerpts;
- refusals, warnings, and no-answer states;
- score by dimension and rationale;
- evaluator identity or scoring-system version;
- fingerprint or journal-integrity mechanism when risk justifies it.
When a parameter is not observable, record “not observable” rather than infer it.
10. Measure without compressing distinct realities
10.1 Visibility outcomes
Descriptive outcomes may include:
- mention rate per eligible answer;
- domain citation rate;
- canonical-source citation rate;
- share of voice with a published denominator;
- coverage by intent family;
- prominence or appearance order;
- recommendation probability for a defined need;
- source diversity and concentration.
10.2 Fidelity outcomes
Reconstruction outcomes may include:
- correct-identity rate;
- category accuracy;
- scope preservation;
- accuracy of critical attributes;
- preservation of exclusions;
- relationship accuracy;
- temporal freshness;
- correct-attribution rate;
- unauthorized-inference rate;
- critical-error rate;
- canon-output gap;
- cross-prompt, cross-model, cross-language, and temporal stability.
10.3 Governability outcomes
A study may also document:
- ability to trace an answer to its sources;
- speed of drift detection;
- ability to version a correction;
- persistence or resolution after correction;
- number of attributable and reproducible incidents.
No global score should hide a critical error, missing source, or structural distinction between visibility and fidelity.
11. Log exogenous events
The event log must track any factor capable of affecting results independently of the intervention:
- launch or expansion of a generative surface;
- model, citation-policy, or retrieval change;
- major change to the index or accessible corpus;
- interface, account, or personalization change;
- advertising campaign, public relations, or media event;
- major competitor change;
- seasonality and industry events;
- outage, rate limit, unavailability, or API change;
- simultaneous site change not included in the intervention.
These events must not be improvised as explanations after an unfavorable result. They must be collected under a rule defined in advance.
12. Test assumptions and try to falsify the effect
A credible analysis does not merely seek confirmation. It also looks for conditions under which the effect disappears or becomes incompatible with the proposed explanation.
Recommended checks include:
- inspection of pre-intervention trends;
- placebo dates;
- placebo units;
- negative-control queries that should not be affected;
- negative-control entities;
- negative-control outcomes that should not move;
- sensitivity to removing a prompt family;
- sensitivity across models and channels;
- spillover analysis into the control group;
- checks for panel-composition changes;
- re-estimation around model updates;
- publication of null and contrary results.
A conclusion that survives only one particular selection of prompts or dates must be described as fragile.
13. Evaluate fidelity with a hybrid method
Fidelity must not depend on a single automated judge.
The protocol recommends three layers:
- Deterministic assertions for attributes, identifiers, exclusions, dates, relationships, and verifiable values.
- Expert rubric for completeness, scope, ambiguity, qualification, and context-sensitive inference.
- Model-assisted evaluation to accelerate coding, detect gaps, and propose rationales, without giving it sole final authority over sensitive claims.
Evaluators should be blind to treatment condition when possible. Disagreements must be measured, adjudicated, and retained.
14. Qualify the level of evidence
The report must use one of the following verdicts.
14.1 Non-identifiable
The data or design cannot distinguish the intervention’s effect from concurrent changes.
14.2 Observed variation
An indicator changed. No additional temporal or causal relationship is asserted.
14.3 Temporal association
The change follows the intervention and remains compatible with it, but the counterfactual is insufficient.
14.4 Plausible contribution
Several elements converge toward a likely role for the intervention, but important alternative explanations remain.
14.5 Causally supported effect within the studied scope
The identification design, controls, and sensitivity analyses support an effect bounded to the studied units, systems, languages, queries, outcomes, and windows.
14.6 No detectable effect
The protocol detected no effect compatible with the study’s threshold and power. This verdict does not prove an exactly zero effect in every context.
14.7 Adverse effect
The intervention is associated or causally linked, depending on the design, to deterioration in the primary outcome or an increase in critical errors.
15. Mandatory minimum report
Every published result should present:
- the causal question and estimand;
- the exact intervention and version;
- the identification design;
- treated and control units;
- the prompt panel and its origin;
- systems, models, languages, regions, and channels;
- baseline and post period;
- primary, secondary, and critical results;
- uncertainty intervals or relevant dispersion;
- known exogenous events;
- placebo tests and sensitivity analyses;
- exclusions and missing data;
- null or adverse results;
- limits to generalization;
- the standardized verdict.
Full logs may be restricted for confidentiality, but their structure, integrity, and sampling rules must remain auditable.
16. Claims prohibited without matching evidence
The protocol prohibits the following reformulations when they exceed the evidence level:
- “We generated X% more citations” from a simple before-and-after comparison;
- “AI visibility increased because of our method” without a credible comparator;
- “Our score proves that the brand is better understood” without fidelity measurement;
- “Improvement in our testbed proves an open-Web increase”;
- “No detection proves that the intervention has no effect”;
- “The results represent the market” when the panel, systems, or denominator are opaque.
Acceptable wording must match the verdict: observation, association, plausible contribution, or supported effect within an explicit scope.
17. Structural limitations
Even when properly applied, the protocol faces several limits:
- interfaces and models may change during the study;
- some versions, parameters, and retrieval systems remain unobservable;
- nondeterminism requires repetitions and raises cost;
- interventions may contaminate control groups;
- observed queries never perfectly represent all use;
- a controlled experiment may have high internal validity but weak open-Web generalizability;
- an open-Web study may be realistic but less causally identifiable;
- no detectable effect may reflect a small effect, insufficient power, or a short window.
Rigor does not make these limits disappear. It makes them visible before the conclusion is sold.
Methodological references
- Miguel A. Hernán and James M. Robins, Causal Inference: What If: https://miguelhernan.org/whatifbook
- Brantly Callaway and Pedro H. C. Sant’Anna, “Difference-in-Differences with multiple time periods”: https://doi.org/10.1016/j.jeconom.2020.12.001
- Alberto Abadie, “Using Synthetic Controls: Feasibility, Data Requirements, and Methodological Aspects”: https://doi.org/10.1257/jel.20191450
- James L. Bernal, Steven Cummins, and Antonio Gasparrini, “Interrupted time series regression for the evaluation of public health interventions”: https://doi.org/10.1093/ije/dyw098
Final reading rule
The protocol does not turn a metric into proof. It defines the conditions under which a claim may cautiously move from observation toward attribution.