Beyond the Screenshot: Establishing a Reproducible Protocol for Measuring Brand Visibility in Generative AI

A single screenshot of a ChatGPT response mentioning a brand is statistically insignificant, functioning merely as a sample size of one. When the same query is repeated moments later, the model frequently returns an entirely different list of entities, rendering anecdotal evidence insufficient for any serious market analysis. Without knowing the exact parameters, volume, and frequency of these queries, it is impossible to distinguish between meaningful signal and algorithmic noise. To address this, the data intelligence firm Brasil GEO has developed a rigorous, auditable protocol for measuring brand visibility in large language models (LLMs), moving the industry toward a standard of scientific reproducibility.
The Problem with Anecdotal AI Monitoring
In the current landscape of digital marketing, "Generative Engine Optimization" (GEO) has become a critical pursuit for brands looking to maintain relevance. However, because LLMs are probabilistic—meaning they generate responses based on weights and patterns rather than static database lookups—a query about "the best software for logistics" may yield a different set of vendors in every session.
If a brand manager relies on a single screenshot to prove their dominance, they are essentially viewing a snapshot of a moving target. This lack of consistency has led to widespread frustration among marketing teams, who often struggle to justify investments in AI visibility because they lack a baseline for what constitutes a "successful" result. The Brasil GEO framework posits that visibility is not a static fact but a distribution variable. By measuring the "Mention Rate"—the number of successful brand mentions divided by the total number of controlled executions—organizations can finally apply statistical rigor to their AI performance metrics.
Defining the Core Variables for Reproducibility
For a visibility report to be considered auditable, four specific variables must remain fixed. If any of these are altered without documentation, the integrity of the longitudinal data is compromised.
- The Query Corpus: A fixed bank of 30 to 40 queries, modeled on actual customer intent, must remain unchanged throughout the measurement window. Brasil GEO, for instance, operated with a bank of 37 queries until August 2026, transitioning to 51 fixed queries in late September. Modifying the phrasing of these questions midway through an observation window invalidates any comparative analysis.
- Model Parameters: Precision requires defining the engine, the "temperature" (a setting that controls the randomness of the model’s output), the specific time of collection, and the frequency of execution. Current best practices dictate a temperature setting of zero to minimize creative variance and ensure that the output is as deterministic as possible across different models, such as ChatGPT, Claude, Gemini, and Perplexity.
- The Temporal Window: Every visibility metric must be bound by a defined start and end date. Crucially, any intervention—such as optimizing a website’s landing page to better answer a specific query—must be logged with a timestamp. This allows researchers to create a clear "before-and-after" cut in the data.
- The Denominator: Transparency requires publishing the ratio of collected queries versus anticipated queries, broken down by model. Providing an aggregate average often hides significant discrepancies in how different models process the same information.
Statistical Requirements and Sample Sizes
The necessity for a high "N" (number of executions) is derived from principles of statistical distribution. Research indicates that below a certain threshold of executions, the natural variance in an AI’s output exceeds the effect the researcher intends to measure.
According to the internal protocols adopted by Brasil GEO, a minimum of five executions per query is required for continuous monitoring. However, for a rigorous "before-and-after" analysis of a site intervention, a minimum of 30 executions is recommended. This aligns with broader industry findings; for example, a landmark study conducted in January 2026 by Rand Fishkin’s SparkToro and Gumshoe.ai analyzed 2,961 prompts across 600 volunteers. The study found that even dominant category leaders appeared in only 55% to 77% of responses. This highlights a critical reality: even the most prominent brands fail to appear in a significant fraction of AI responses, proving that "perfect" visibility is a fallacy in the age of generative search.
The Seven-Step Pipeline for Data Integrity
To ensure that results are not tainted by noise or technical errors, practitioners should follow a standardized seven-step pipeline:

- Establish the Bank: Curate 30–40 customer-centric queries with "frozen" text.
- Define Entities: Map the target brand alongside 15 to 25 competitors, accounting for all known spellings and common misspellings. For short acronyms, the protocol must enforce domain-specific context.
- Set Protocol: Define engines, temperature, and specific execution schedules.
- Establish Baseline: Run at least 30 executions per query before implementing any changes to the digital strategy.
- Document Intervention: Record the precise date and time a page begins to successfully address a query in the bank.
- Report Denominators: Explicitly declare the ratio of collected versus planned queries for every dataset.
- Comparative Analysis: Evaluate the mean of the new window against the previous baseline; if the variation falls within the previously observed noise range, the change is considered statistically insignificant.
The Challenge of False Negatives and Entity Consistency
One of the greatest risks in this monitoring process is the prevalence of "false negatives"—where a brand is present in the training data but fails to be identified by the monitoring system. Brasil GEO’s internal audits revealed that 11 out of 21 tracked competitors were scanned over 60 times without a single detection. This confirmed a total absence of brand visibility, but only after a thorough review of spelling variants and entity aliases.
Conversely, short acronyms often trigger "false positives." To mitigate this, a detection system must look for context within a 120-character window surrounding the acronym. Furthermore, there is an upstream dependency on how the model perceives the brand. If information about an entity is fragmented or contradictory across various data sources, the mention rate becomes a measure of "algorithmic confusion" rather than brand authority. Industry standards now aim for an Entity Consistency Score of 0.9 or higher to ensure the AI has a coherent understanding of the entity in question.
Scaling the Protocol: The Brasil GEO Index
The efficacy of this structured approach is demonstrated in the "Brasil GEO Index," an initiative that applies this protocol to an entire market. In a recent data capture dated September 9, 2026, the index analyzed 88,911 responses across 127 entities. The report observed a 36.1% overall citation rate, with a 95% confidence interval ranging from 35.8% to 36.4%.
To maintain the integrity of the data, the index includes 16 fictitious entities as a control group for false positives. Of the 90 days in the observation window, 54 days yielded usable, recorded data. Rather than using projections or interpolations to fill the remaining 36 days, the team left those entries blank, ensuring that the final output remained rooted in observed, rather than estimated, performance.
Broader Implications for Digital Strategy
The shift toward auditable, protocol-driven visibility measurement marks a maturation point for the digital marketing industry. As generative AI becomes the primary interface for consumer search, the "black box" nature of these systems will no longer be an acceptable excuse for poor data transparency.
By treating AI visibility as a statistical distribution rather than a search-engine ranking, firms like Brasil GEO are providing a roadmap for CMOs and digital strategists to quantify their brand’s presence in a way that is repeatable and defensible. As models continue to evolve, the reliance on the "four fixed variables" and the "seven-step pipeline" will likely become the standard for any organization serious about its long-term AI strategy.
Ultimately, the goal is not to "hack" the AI, but to understand its probabilistic nature. Brands that adopt these rigorous measurement protocols will be better positioned to identify when their visibility drops due to technical issues—such as poor entity mapping—versus when it drops due to a genuine decline in market authority. In an era of AI-generated content, evidence-based measurement is the only path to sustainable competitive advantage.
Alexandre Caramaschi is the Chief Strategy Officer at Nuvini (Nasdaq: NVNI). The findings and methodologies presented here are issued in his capacity as the Founder of Brasil GEO and do not necessarily represent the official position of Nuvini.







