Before measuring drift, establish an AI perception baseline
A brand may appear better or worse represented from one week to the next. Without a reference state, that impression cannot distinguish model change, prompt change, new sources or normal output variability.
A baseline is not an ideal portrait. It is a dated set of observations preserved before intervention.
Baseline components
| Component | Content |
|---|---|
| Entity | Names, identifiers, domains, products, subsidiaries and former names |
| Canon | Current claims, boundaries, exclusions and version |
| External sources | Reputation, criticism, decisions and structuring descriptions |
| Prompts | Intent families, wording and variables |
| Conditions | Model, product, date, language, region, browsing and session |
| Outputs | Full answers, citations and captures |
| Coding | Identity, category, scope, sentiment, sources and recommendability |
Common failure: baseline after correction
If measurement starts only after the site was corrected, there is no defensible “before.” Current output can be observed, but change cannot be attributed to the intervention. Reconstructing history from memory or selected screenshots introduces bias.
Freeze the initial state, including inconvenient errors. Its value comes from preserving what must later be compared.
Example
A company launches a new position. Before publication, models mostly describe it as a general agency. The baseline records 40 prompts across definition, comparison, recommendation and customer problems in French and English.
After the redesign, the same protocol shows that the new category appears more often in French while generic English prompts still select the former competitor set. The result is a localized improvement and residual cross-language drift, not a global success claim.
Baseline versus absolute truth
The baseline canon must be bounded. It can establish identity, offer and intended positioning. It cannot store “good reputation” as official truth. External claims remain separate.
It may also contain unproven ambitions. A brand can wish to be recognized as a leader without comparative evidence. The audit should mark the claim as intended positioning, not an invariant that systems must adopt.
Size and frequency
There is no universal prompt count. The sample depends on intents, languages, models and risk. A local business may begin with 20 to 40 carefully selected prompts; a multilingual regulated organization may require hundreds.
Re-observation cadence increases around rebrands, incidents and migrations. Protocol changes are versioned, never silently merged.
Minimum deliverable
A baseline contains a protocol ID, prompt inventory, canon state, output log, coding grid, source map, limits, hypotheses and next observation date.
It does not predict future answers. It supplies the zero point without which “drift,” “improvement” and “successful correction” remain unproven statements.