A Measurement Model for Brand Visibility in ChatGPT Answers: Sampling Noise, Day Effects and Level Drift, and Their Consequences for Monitoring
Brands track a visibility rate in generative engines: the share of answers to a prompt that mention them. Such rates are reported as point values, although every answer is a stochastic draw and the engine changes from one day to the next. We study the consumer web interface of ChatGPT with eight months of production monitoring (January to September 2026) from GetMint, an AI-visibility platform: more than a million answers to tens of thousands of prompts, most of them observed once a day. We model the daily true rate of a prompt as a slowly drifting level plus a transient day effect, observed through binomial draws, with rare common shocks shared by all prompts on the surface. Within a day we take the binomial as the working model of sampling: a random split of a day's runs cannot test it, its dispersion being fixed by construction, while a chronological split shows that the rate already moves within the day, by about 10 points between its two halves. The binomial gives a closed-form run abacus for the day's rate. Across days the picture is different. On days with replicated runs, where the sampling variance is estimated within the day under that model, the daily true rate of a mid-rate prompt moves by 14 points between consecutive days and by 25 to 33 points at gaps of one to six weeks. This movement decomposes into a transient day effect of 14 to 16 points, which fades within two to three days, and a level drift of about 4 to 4.5 points per day, with compatible values on two observation windows. We show analytically and empirically that the usual one-run-per-day estimator with a window plug-in measures movement minus twice the between-day variance of the rate, sums to zero over the pairs of a series and crosses zero at their mean gap whatever the engine does, and that likelihood estimation with one run per day does not recover a non-persistent day effect: replicated days are necessary. The consequences for monitoring follow. No number of runs taken on a single day measures a prompt's level to better than about ±28 points; precision comes from pooling days and prompts. A 20-prompt topic pooled over 14 days is known to about ±7 to ±8 points in calm periods when the drift of its prompts is idiosyncratic, and a same-day burst of 600 answers to about ±7 as well, not the ±4 that sampling alone promises. A behavioural sentinel dates common shocks (seven on ChatGPT web from May to July 2026, six of them unannounced, and an unannounced model version change on 8 August). A state-space filter using the sentinel's flags tracks the level with a 95 % band whose coverage we check on simulated paths, under the model and under a persistent day effect (94 to 95 % in calm periods, about one half while a level is ramping); a 14-day moving average approximates the filter. Three estimators are given to attribute a movement to a brand's own action, against the model's drift and against control prompts. Aggregate tables and the analysis code are available from the author; raw answers are proprietary.
Authors
- Amine Aziz Alaoui
Institutions
- Gemini Computers (United States) (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23156370
- Primary Topic
- Digital Marketing and Social Media
- Type
- preprint