Skip to content

Framework

Contextual fidelity protocol

Proposed protocol for measuring contextual adaptation of entity representation without altered invariants, lost conditions or unsupported recommendation.

CollectionFramework
TypeProtocol
Layertransversal
Version0.1-proposed
Stabilization2026-08-16
Published2026-08-16
Updated2026-08-16

Governance artifacts

Governance files brought into scope by this page

This page is anchored to published surfaces that declare identity, precedence, limits, and the corpus reading conditions. Their order below gives the recommended reading sequence.

  1. 01situational-applicability-map.json
  2. 02interpretive-weighting-policy.json
  3. 03attested-interpretive-units.json
Artifact#01

situational-applicability-map.json

/situational-applicability-map.json

Published machine-first governance surface.

Governs
Part of the corpus reading conditions.
Bounds
An inference zone that would otherwise remain implicit.

Does not guarantee: This file does not, on its own, guarantee system obedience.

Artifact#02

interpretive-weighting-policy.json

/interpretive-weighting-policy.json

Published machine-first governance surface.

Governs
Part of the corpus reading conditions.
Bounds
An inference zone that would otherwise remain implicit.

Does not guarantee: This file does not, on its own, guarantee system obedience.

Artifact#03

attested-interpretive-units.json

/attested-interpretive-units.json

Published machine-first governance surface.

Governs
Part of the corpus reading conditions.
Bounds
An inference zone that would otherwise remain implicit.

Does not guarantee: This file does not, on its own, guarantee system obedience.

Complementary artifacts (1)

These surfaces extend the main block. They add context, discovery, routing, or observation depending on the topic.

Evidence layer

Probative surfaces brought into scope by this page

This page does more than point to governance files. It is also anchored to surfaces that make observation, traceability, fidelity, and audit more reconstructible. Their order below makes the minimal evidence chain explicit.

  1. 01
    Canon and scopeDefinitions canon
  2. 02
    Response authorizationQ-Layer: response legitimacy
  3. 03
    Evidence artifactcontent-digests.json
Canonical foundation#01

Definitions canon

/canon.md

Opposable base for identity, scope, roles, and negations that must survive synthesis.

Makes provable
The reference corpus against which fidelity can be evaluated.
Does not prove
Neither that a system already consults it nor that an observed response stays faithful to it.
Use when
Before any observation, test, audit, or correction.
Legitimacy layer#02

Q-Layer: response legitimacy

/response-legitimacy.md

Surface that explains when to answer, when to suspend, and when to switch to legitimate non-response.

Makes provable
The legitimacy regime to apply before treating an output as receivable.
Does not prove
Neither that a given response actually followed this regime nor that an agent applied it at runtime.
Use when
When a page deals with authority, non-response, execution, or restraint.
Artifact#03

content-digests.json

/content-digests.json

Published surface that contributes to making an evidence chain more reconstructible.

Makes provable
Part of the observation, trace, audit, or fidelity chain.
Does not prove
Neither total proof, obedience guarantee, nor implicit certification.
Use when
When a page needs to make its evidence regime explicit.

Contextual fidelity protocol

The contextual fidelity protocol measures whether a system correctly adapts an entity representation when context changes. It does not seek identical answers. It seeks explainable variation with preserved invariants.

The observation unit is:

E × C × S × A × L × T
  • E: entity;
  • C: context profile;
  • S: system or model;
  • A: access channel such as API, interface or tool-using agent;
  • L: language or region;
  • T: observation time.

These dimensions must remain separate. An average mixing user interface, controlled API, browsing, regions and dates erases the very conditions the protocol is designed to measure.

1. Preconditions

Before a campaign:

  1. complete the interpretive conditioning matrix;
  2. resolve entity identity and homonyms;
  3. establish a dated invariant baseline;
  4. define at least two profiles whose difference is materially relevant;
  5. classify sources by authority scope;
  6. identify volatile relations and expiry dates;
  7. declare permitted output modes;
  8. prepare forbidden transformations;
  9. fix capture and archival procedures;
  10. define stop conditions and non-judgeable cases.

A campaign without a baseline cannot distinguish variation from drift. A campaign without contrasted profiles does not measure conditioning.

2. Building the context set

The minimum set includes:

  • a reference profile;
  • a favourable profile;
  • an unfavourable or non-applicable profile;
  • an ambiguous profile with material missing data;
  • a counterfactual profile where one condition is inverted;
  • a temporal profile before, during or after a validity window.

Profiles should vary one material dimension at a time whenever possible. Changing ten variables at once prevents causal attribution.

3. Probe families

Positive probes

They verify that the system uses a legitimate relation when conditions are met.

Example: ask whether a hotel is relevant to a car-free stay when destinations and schedules are provided.

Negative probes

They verify correct suppression or weakening when a non-applicability condition is present.

Example: add a need for secure parking when the hotel has none.

Ambiguous probes

They remove material data to test clarification, uncertainty or abstention.

Example: ask for the “best hotel” without destination, budget, dates or criteria.

Trap probes

They explicitly invite the system to cross a boundary.

Example: “Since this hotel is close to my destination, confirm it is the best-located hotel in the city.”

A faithful system should resist generalization.

Counterfactual probes

They invert one condition while preserving the entity.

Example: replace “without a car” with “requires long-term parking” and verify that the conclusion changes without changing hotel facts.

Temporal probes

They test freshness, expiry and temporary events.

Example: compare a conclusion before, during and after announced construction.

Comparative probes

They test data symmetry and preservation of criteria.

Example: compare two hotels using the same variables, then remove one datum for one option to verify whether unknown is treated as inferior.

4. Capture protocol

Every observation should preserve:

  • campaign identifier;
  • matrix version;
  • entity and context profile;
  • exact prompt or structured input;
  • system, declared version and parameters;
  • access channel;
  • language, region and date;
  • tools or retrieved sources;
  • raw output;
  • citations or links;
  • technical errors;
  • deterministic annotations;
  • human or assisted judgment;
  • confidence level;
  • final decision and rationale.

API and user-interface campaigns must remain separate. An interface may add memory, search, geolocation, personalization or orchestration absent from the API.

5. Statement segmentation

The output should be decomposed into verifiable units:

  • identity claim;
  • invariant;
  • contextual relation;
  • temporal state;
  • conditioned interpretation;
  • comparison;
  • recommendation;
  • uncertainty;
  • clarification;
  • exclusion or abstention.

Each unit receives an expected source class, scope and qualification. Global answer judgment must not hide a local material error.

6. Primary metrics

Invariant preservation

Proportion of material invariants correctly preserved in outputs where they become relevant.

Omission and contradiction must remain separate. Contradiction is more severe.

Correct contextual sensitivity

Ability to change the conclusion when context changes materially without changing invariant facts.

A system that is always identical may be context-insensitive. A system that changes everything may be over-conditioned.

Variation attribution

Proportion of differences for which the causal context dimension can be identified in the output or reconstructed from the trace.

Relational fidelity

Accuracy, freshness and scope of used relations. A correct distance applied to the wrong destination remains unfaithful.

Inversion consistency

When condition C is inverted, does the conclusion change in the expected direction while preserving the entity?

This metric detects superficial personalization and conclusions that do not really depend on announced criteria.

Exclusion preservation

Rate at which relevant limits, non-applicability conditions and contraindications are preserved.

Overstatement

Frequency with which an output moves from relation to property, relevance to superiority or bounded comparison to global recommendation.

Unsupported recommendation

Proportion of recommendations produced without a sufficient comparison set, explicit criteria, symmetrical data or arbitration rule.

Contextual fossilization

Frequency with which a temporary or expired relation survives as a current property.

Legitimate clarification

Ability to request genuinely missing information without creating unnecessary friction when context is already sufficient.

7. Result classification

Each observation may be classified as:

  • faithful and conditioned: invariants preserved, variation explained, scope preserved;
  • faithful but incomplete: no contradiction, but relations or conditions omitted;
  • context-insensitive: unchanged conclusion despite a material difference;
  • over-conditioned: preference or context modifies entity facts;
  • overstated: output stronger than evidence;
  • drifted: contradiction, generalization or fossilization;
  • non-judgeable: insufficient data, trace or identity;
  • legitimate abstention: the system correctly refuses to conclude.

This taxonomy must remain available alongside quantitative metrics.

8. Hybrid judgment

The protocol recommends three levels:

  1. deterministic controls: dates, identifiers, exclusions, citations, absolute terms and exact data;
  2. human annotation or expert rule: scope, source competence and materiality;
  3. assisted LLM judge: semantic comparison for ambiguous cases, without autonomous final authority.

The LLM judge must not receive only the output. It should receive invariants, profile, expected relations, forbidden transformations and decision rule. Its own results must be sampled and contestable.

9. Repetition and stability

A single output is insufficient for a stochastic system. Every material cell should be repeated according to a declared observation count.

Preserve separately:

  • frequency of each class;
  • confidence interval or sampling uncertainty;
  • wording variation;
  • conclusion variation;
  • critical divergences;
  • temporal evolution.

Repetition of an error increases its stability, not its fidelity.

10. No initial global score

The proposed version produces no single score. Dimensions remain separate because some errors are non-compensable.

Excellent contextual sensitivity does not compensate for identity contradiction. Strong average relevance does not compensate for a dangerous recommendation in a regulated sub-context. Broad coverage does not compensate for fossilized stale information.

An aggregate score might later be added for a specific use, provided weights, thresholds, critical errors and non-compensation rules are published.

11. Campaign report

The final report should include:

  • scope and assumptions;
  • matrix and versions;
  • systems, channels, languages, regions and dates;
  • observation count per cell;
  • separate metrics;
  • representative examples;
  • critical contradictions;
  • non-judgeable cases;
  • limitations;
  • recommended changes to canon, relations, freshness or response conditions;
  • re-observation plan.

The protocol must not automatically attribute every error to the site. Drift may originate in the model, retrieval, an external source, former state, identity ambiguity or poorly formulated context.

12. Boundaries

This protocol does not prove future system behaviour, guarantee recommendation or replace domain-specific safety evaluation. It measures outputs under declared conditions.

It must not be used to train a site to manipulate personal preference or present self-evaluation as independent proof. The site declares. The auditor measures. The external agent interprets or recommends under its own responsibility. The verdict must remain tied to the criteria for legitimate contextual variation, never to response fluency alone.