Search Console AI Reports: How to Measure GEO Performance

Search Console AI Reports: How to Measure GEO Performance
What changed with Search Console’s generative AI reports?
Search Console’s generative AI performance reports create a dedicated view of how often site URLs appear in Google AI Overviews, AI Mode and generative experiences in Discover. Google announced the experiment on 3 June 2026 and is initially testing it with a subset of sites, so the reports will not yet appear in every property.
The Search report includes impressions, pages, countries, devices and dates. The Discover version focuses on impressions, pages and dates, while both support hourly, daily, weekly and monthly views. Their underlying data remains part of the main Performance report; adding the dedicated and overall totals together would therefore double count the same visibility.
This release resolves one long-standing measurement gap: teams can separate at least part of Google’s generative search exposure from conventional result activity. It does not reveal the quoted sentence, the visual prominence of a citation or the other sources used in an answer. An impression is evidence of eligibility and exposure, not proof that a person noticed, trusted or clicked a brand.
Google’s official announcement makes the limited rollout explicit. A team without access should not postpone measurement. It can establish a baseline from the standard Performance report, analytics referrals and a controlled prompt audit, then add the new report when it becomes available.
The country and device dimensions also make the data useful for geo-targeted PR planning. If a source page earns AI impressions on mobile in the United Kingdom but remains absent in another priority market, the next action may be local evidence and market-specific framing rather than another generic translation.
What belongs in an AI citation audit?
An AI citation audit combines a fixed set of buyer questions with answer captures, cited URLs, referral data and business outcomes in one repeatable record. A practical first audit uses 30–50 prompts across three markets, at least three answer engines and four weekly observations.
The prompt set should be built around decisions rather than the company name. Questions such as “How do I measure brand visibility in AI search?”, “What evidence should a GEO agency provide?” and “How does earned media affect answer-engine citations?” expose different needs. Each prompt needs a language, market, funnel stage, service theme and expected source type.
Each observation should record:
- The exact prompt, date, market, language and platform.
- Whether the brand appears and whether the context is neutral, positive, inaccurate or explicitly recommendatory.
- Every cited URL, its publisher, date, author and main supporting claim.
- Whether the link leads to the brand, an independent editorial source or a third-party directory.
- Any unsupported, outdated or misleading statement that requires correction.
- The change from the previous observation and the content, technical or PR action that preceded it.
One screenshot is not a trend. Answers can change with model version, location, session context, retrieval timing and the wording of a follow-up. Run the core sample in clean sessions, within a defined time window and under the same market conditions. Use four observations before classifying movement as a directional signal.
Google’s technical guidance for AI features also prevents wasted work. There is no special AI schema or separate text file required for inclusion. Pages still need to be crawlable, indexable and eligible to show a snippet, while robots controls and preview settings continue to apply.
Which metrics should define GEO performance?
GEO performance should be measured through four layers: exposure, citation, referral and business outcome. Keeping those layers separate prevents a visibility increase from being reported as commercial impact before the evidence exists.
- Exposure: Generative AI impressions, the number of surfaced URLs, countries, devices and movement over time.
- Citation: Brand mention rate, linked citation rate, citation position, source diversity and the pages selected as evidence.
- Referral: Sessions from answer engines, engaged time, new-user share, assisted journeys and progress to a priority page.
- Outcome: Qualified forms, meetings, sales opportunities and revenue influence by service or market.
Suppose Google AI impressions rise by 25 per cent while the prompt audit shows no change in brand mentions. More pages may be entering the generative result environment, but the brand has not necessarily become the preferred source. If mentions rise while referral traffic stays flat, the answer may satisfy the user without a visit or may place the citation where few people act on it.
OpenAI’s publisher FAQ states that ChatGPT search referrals include utm_source=chatgpt.com, which gives analytics teams a consistent acquisition marker. It also asks publishers to allow OAI-SearchBot for summaries and snippets. That access creates the possibility of retrieval; it does not guarantee placement or rank.
Perplexity describes its answer engine as a system that searches the web, identifies trusted sources and synthesises a response. Its crawler documentation says robots.txt changes can take up to 24 hours to be reflected. A complete audit therefore checks crawl access, but never treats an allowed crawler as evidence that a source has already been selected.
This layered model extends the logic behind integrated SEO and PR measurement. Organic search, an AI answer and an editorial article can be separate steps in one decision journey; a last-click report will miss much of that relationship.
Why can Search Console not measure every answer engine?
Search Console cannot provide a complete GEO view because it measures Google properties, not citations and referrals generated by ChatGPT, Perplexity or other independent answer engines. It also does not show the sentence attributed to a page, the nearby narrative or whether an external editorial source reinforced the brand.
Google currently groups traffic from its AI features within the Web search type in Search Console. The dedicated experiment improves impression analysis, but a researcher must still inspect representative answers to understand how the source is used. The manual layer should record accuracy, citation prominence, competing evidence and the relationship between the cited page and the brand claim.
ChatGPT offers a clearer referral marker, yet many citations produce no click. Perplexity exposes its own source links and uses a separate crawler. A brand can consequently improve in one environment while remaining static in another. Combining the platforms into a single “AI visibility score” can hide that difference and encourage the wrong intervention.
The foundational Generative Engine Optimization research, published at KDD 2024, also treats visibility as a multidimensional problem because citations differ in position, length and presentation. Its controlled benchmark found that some interventions could improve visibility substantially, but those findings are not a universal performance promise for every query or category.
Measurement should preserve provenance. Every chart must state which platform, market, prompt set, time period and sample size it represents. When a product launch, content revision, technical repair and media campaign happen together, the report should avoid assigning the whole movement to one channel without a comparison group.
How does a 30-day GEO measurement sprint work?
A 30-day GEO measurement sprint moves through baseline, source improvement, controlled retesting and an action report in four weekly stages.
- Days 1–7: Define 30–50 decision prompts and capture results from Google, ChatGPT and Perplexity. Export the previous 28 days from the dedicated Search Console report where available; otherwise retain the standard Web performance baseline.
- Days 8–14: Repair crawl barriers and improve the weakest source pages. Add a direct first answer, dated evidence, an accountable expert, clear definitions and links to primary sources.
- Days 15–21: Repeat the prompt sample under the same conditions. Mark changes in mentions, cited URLs, citation prominence and factual accuracy.
- Days 22–30: Join exposure, citation, referral and outcome data. Assign each prompt group a decision, an owner, a review date and a measurable success threshold.
Log every publication and technical change. If a team changes the headline, body copy, structured data, crawler access and external evidence in the same week, it loses the ability to learn which intervention mattered. Split the prompt set into clusters and change one primary variable for each test whenever operationally possible.
The reporting interval should match the decision. A crisis or regulatory update can justify daily monitoring; durable category authority is better assessed weekly. The monthly management view should replace screenshot collections with an explanation of what moved, how confident the team is, what evidence supports the diagnosis and what happens next.
A useful quality gate is simple: no metric reaches the executive summary unless it leads to a decision. Impression growth can justify protecting a source page; a declining linked citation rate can trigger a clarity review; inaccurate summaries can require stronger first-party evidence and spokesperson commentary; strong referrals without qualified action can signal a landing-page problem.
How should findings change content and PR investment?
GEO findings should place every query cluster into one of four actions: protect, strengthen, correct or create independent evidence. This converts monitoring into a prioritised programme for editorial, technical, analytics and media-relations teams.
High Google AI exposure with weak brand context calls for sharper definitions, comparison criteria and evidence on the source page. Strong technical eligibility with repeated citations to independent publishers signals a different gap: publishing more owned articles may not change the evidence environment. The brand may need credible third-party reporting that connects its experts with a specific subject.
This is where a successful PR and SEO strategy becomes operational. Earned media created through real journalist demand records an independently evaluated relationship among a person, an organisation and an area of expertise. The audit can then test whether those articles begin to support answers in the intended markets.
FL PR & Communications uses three decision questions: Does the brand appear for the right buyer question, is it connected to credible evidence, and does the exposure move closer to qualified demand? A negative answer does not automatically justify more content. The constraint may be crawl access, answer clarity, source freshness, market relevance or a lack of independent editorial proof.
The most useful executive output is a short decision register. Each row contains the prompt cluster, present visibility, evidence gap, recommended intervention, responsible team, next measurement date and success threshold. That format makes GEO comparable with other investments without pretending that every mention can be reduced to one universal score.
A mature programme also records what it will not claim. It does not promise inclusion, treat correlation as attribution or label every model response as stable. It uses platform-native data where available, preserves the limits of each source and keeps the measurement method consistent enough for the next decision to be better than the last.
Frequently Asked Questions
These answers address the four implementation questions that most often arise when teams build a GEO measurement and AI citation audit programme.
Are Search Console generative AI reports available to every site?
No. As of June 2026, Google is testing the reports with a subset of sites. Properties without access can use standard Search Console data, answer-engine referrals and a controlled manual prompt audit to establish a baseline.
Does an AI impression prove that the brand was cited?
No. An impression shows that a URL appeared in a generative AI surface; it does not prove that the brand name appeared, the citation was prominent or a person clicked it. Citation and referral evidence must be measured separately.
How many prompts should an AI citation audit include?
A first audit should usually include 30–50 decision prompts, tagged by language, market, funnel stage and service theme. Run them across at least three answer engines for four weekly observations before treating movement as directional.
How long does it take to measure GEO progress?
A 30-day sprint can reveal an initial direction, while a dependable trend normally requires 8–12 weeks. Crawling, indexing, editorial publication and model refreshes run on different schedules, so a single weekly increase is not conclusive.