Blog Posts

How to Measure AI Search Visibility with a Prompt Test Set

Two researchers reviewing AI search prompt cards in a library

How to Measure AI Search Visibility with a Prompt Test Set

What is a prompt test set for AI search?

A prompt test set is a fixed group of real user questions used to measure when a brand appears in AI answers, which sources support that answer and how stable the result is over time. It is not a keyword rank tracker; it is a repeatable evaluation system for answer engines.

Classic SEO reporting can show whether pages receive impressions and clicks, but it cannot fully explain how an AI answer chooses, paraphrases or cites a source. Google says traffic from AI features is included in Search Console’s Web search type, yet the report does not isolate every AI answer or citation. That makes a separate prompt-level testing layer necessary for serious GEO work.

A practical first set contains 30-50 prompts. The set should not ask only branded questions. It should simulate the moments where a buyer, patient, founder, marketing director or communications lead asks for a definition, a shortlist, a risk check, a process or proof.

GEO-targeted PR starts from that question map. The aim is not to publish more owned pages for their own sake, but to understand which independent and owned sources help an answer engine connect a brand to a topic with confidence.

Which prompts belong in a representative test set?

Representative prompts are selected from decision intent, not from a flat keyword list. A good set covers definition, vendor selection, proof, process, comparison, objection and market-specific questions.

The difference matters. “GEO agency” is a search term; “How does a GEO agency help a healthcare brand get cited in AI search?” is a decision prompt. The second version reveals whether the answer can explain the process, mention credible source types and connect the topic to a service category.

  • Definition prompts: Questions such as “What is GEO?” or “How is AEO different from GEO?” test conceptual ownership.
  • Selection prompts: Questions such as “What should a brand check before hiring a GEO agency?” test buying criteria.
  • Proof prompts: Questions such as “Which types of media evidence influence AI citations?” test trust signals.
  • Process prompts: Questions such as “How can a brand prepare an AI source profile in 90 days?” test operational clarity.
  • Comparison prompts: Questions such as “How does earned media differ from distribution-only PR for GEO?” test strategic framing.

The prompt list should include neutral category questions and practical problem questions. If every prompt is written to force a brand mention, the report becomes weak. The useful question is not “Did the brand appear once?” but “In which answer contexts does the brand appear with the right evidence?”

successful PR and SEO strategy depends on shared language between search intent, media relations and owned content. A prompt test set checks whether that shared language is actually visible in AI answers.

How should market, language and session conditions be controlled?

Market, language and session conditions must be controlled because AI answers can change when the country, language, interface, model or browsing context changes. Without those controls, two test runs cannot be compared with confidence.

Every row in the test sheet should record the exact prompt, target market, language, date, tool, answer surface, session state and source URLs shown in the response. If a prompt is adapted for the United States, the United Kingdom or Turkey, that version should be stored as a separate market variant. Small wording changes can produce meaningful differences, so the core prompt wording should stay stable between reporting periods.

The first 90 days should use a simple cadence: one full baseline, one 30-day repeat, one 60-day repeat and one 90-day readout. Additional checks can follow major media placements, technical changes or new service pages, but the official comparison should remain tied to the fixed set.

A card layout used to group AI search visibility test prompts
A prompt test set becomes useful when questions, markets, source types and answer quality are recorded in the same repeatable table.

Repeating a test does not remove AI variability. It makes that variability visible. If the same prompt is run three times on the same day, the results should not be averaged away; each answer should be stored so the team can calculate whether a citation is stable, occasional or absent.

How should citations and answer quality be scored?

Citations and answer quality should be scored by brand presence, source inclusion, source type, factual accuracy, answer prominence and relevance to the decision question. A brand mention without a credible source is a weak GEO signal.

A workable scorecard can use a 0-3 scale. Zero means the brand is absent. One means the brand appears but with weak or unclear relevance. Two means the brand appears in the right context. Three means the brand appears in the right context and is supported by a credible source that a user can open.

  • Entity accuracy: Is the brand, person, service area and market relationship described correctly?
  • Source type: Does the answer rely on owned media, earned media, an official profile, a list page or a news article?
  • Source availability: Does the cited URL open, match the topic and remain indexable?
  • Answer prominence: Is the brand part of the main recommendation, a secondary example or an unrelated mention?
  • Claim control: Are dates, results, credentials and service claims supported by live sources?
  • Category context: Are alternatives compared by process, evidence and measurement criteria rather than vague claims?

This scoring system brings media evidence and search reporting into one language. the transformative role of SEO in PR strategy is clearest when technical discoverability, editorial authority and trust signals are measured together.

OpenAI’s evaluation guidance points in the same operational direction: define the task, the data source and the criteria before judging model output. OpenAI’s Evals guide describes structured evaluation for model behavior; in GEO reporting, the equivalent is a maintained prompt set with answer and citation records.

How do repeated runs separate signal from variation?

Repeated runs separate signal from variation by testing the same prompt set under the same conditions across multiple reporting periods. One positive answer is not a trend; a brand that appears with the right source across three cycles has a stronger visibility signal.

AI answers are naturally variable. A tool may cite a brand in one session, omit it in another and reintroduce it after stronger editorial evidence becomes visible. The report should treat that movement as data, not noise. Stability, source quality and accuracy should be tracked alongside simple presence.

A useful reporting table has five columns for each prompt: presence, citation, source quality, answer accuracy and next action. The next action column matters because GEO measurement should lead to work: a profile update, a clearer source page, a stronger internal link, a media evidence gap or a pitch angle for earned media.

Google’s 2026 AI search optimization resource reinforces the same principle: helpful, accessible and distinctive content remains the base layer. Google Search Central’s AI search optimization guidance supports focusing on useful pages and clear evidence rather than creating repetitive pages for every variation of a query.

How does a GEO agency run a 90-day evaluation programme?

A GEO agency runs a 90-day evaluation programme through source inventory, prompt design, evidence improvement, earned media integration and repeat testing. The deliverable is a decision dashboard that shows where the brand is cited, where it is absent and which sources need to be strengthened.

The first 15 days are used to map the current source surface. Owned pages, expert profiles, case studies, press coverage, structured data, service pages and social proof are reviewed separately. Gaps are not labelled as “write more content”; they are mapped to the decision questions they fail to answer.

The second phase builds the prompt set and scorecard. The third phase improves the source layer: internal links, expert pages, evidence-led service content and editorial proof. brand authority in the AI era depends on making independent evidence and expertise visible, not merely repeating a brand’s own claims.

The 90-day workflow is simple enough to manage and strict enough to compare:

  • Days 0-15: Map owned pages, expert profiles, live media evidence and existing AI answer behavior.
  • Days 15-30: Lock the 30-50 prompt set, market variants and scoring criteria.
  • Days 30-60: Improve missing source pages, internal links, expert proof and earned media evidence.
  • Days 60-90: Repeat the fixed prompt set and report source quality, accuracy and visibility movement.
  • Day 90: Keep high-signal prompts, remove weak prompts and add new market or service variants.

FL PR & Communications uses this model to connect digital PR, global earned media and GEO measurement in one operating system. The point is not to count published articles; it is to show which credible sources help answer engines verify a brand’s expertise.

Frequently Asked Questions

The answers below cover the most common questions about AI search visibility measurement and prompt test sets.

How many prompts should a test set include?

A first test set should include 30-50 prompts. That range is enough to cover definition, selection, proof, process and comparison questions without creating a reporting system that is too large to repeat.

Does Search Console separately report AI search visibility?

Google states that traffic from AI features is included under the Web search type in Search Console performance reports. Search Console is useful, but it does not replace prompt-level answer and citation testing.

Is a brand mention enough to count as GEO success?

No. The brand should appear in the right decision context and, ideally, with a credible source that supports the claim. A mention without context or citation should receive a low score.

How often should prompt tests be repeated?

During the first 90 days, a full test should be repeated monthly. Extra checks can follow major earned media placements or technical updates, but the main comparison should use the fixed monthly set.