Skip to content
All articles
AI Search

AI Search Prompt Coverage Auditor: Find Blind Spots Before an AEO Study

By the AEOeye editorial team·Updated Sep 30, 2026·9 min read
Research notes and a laptop used to audit an AI search prompt register.
Photo by Pexels on Pexels

Before collecting AI answers, run the prompt register through a declared coverage check. The downloadable AI Search Prompt Coverage Auditor counts only eligible rows across buyer stage, persona, topic, locale, and engine, reports unmet minimums and unknown values, and produces deterministic JSON. It does not prove representativeness, demand, provider quality, or that a prompt will receive an answer.

This is a documentation-based measurement asset with a synthetic fixture, not a live provider benchmark. Its governance framing follows the NIST AI Risk Management Framework, its Generative AI Profile, and provenance concepts in W3C PROV-O. The Python implementation uses only the standard library.

Table of contents

A researcher comparing a structured checklist with a laptop, representing prompt coverage review. Editorial image: Pexels photo 3769021; it is not a study observation.

What does prompt coverage auditing check?

It checks whether the register you declared is internally coherent before collection starts. The tool validates required columns, unique prompt_id values, allowed statuses (active, disabled, and excluded), requirement dimensions, integer minimums, and the active-row counts for every required dimension/value pair.

An “eligible” prompt is an active row. Disabled and excluded rows remain visible in the input so a reviewer can explain why they were not sampled, but they never inflate coverage. The report also lists active values that were not declared in the requirements file. That catches a spelling drift such as UnknownEngine, rather than silently creating a new category.

The result is a sampling-frame diagnostic. It is not a model score or a quality grade. A passing register can still contain biased wording, duplicate intent, weak buyer questions, or too few repetitions. A failing register can be useful during study design when gaps are expected and being resolved.

Why does coverage change visibility claims?

Coverage changes the denominator behind every mention, recommendation, or citation rate. If pricing prompts are missing, a claim such as “the brand is recommended in 40% of buyer answers” may describe only discovery questions. If one locale is absent, an overall rate can hide a market where the brand is never named.

Keep the dimensions separate because they represent different decisions:

  • Stage: discovery, consideration, and decision questions test different moments in a buyer journey.
  • Persona: a founder, marketer, analyst, or procurement reader may request different evidence.
  • Topic: category, comparison, alternatives, pricing, and implementation expose different competitive frames.
  • Locale: language-region combinations should be explicit; fr-FR is not interchangeable with en-US.
  • Engine: engine labels identify the declared systems being sampled, not a claim about their internal routing.

NIST’s risk-management guidance emphasizes context, measurement, and documented limitations. In practice, that means publishing the frame alongside the rate. Pair this asset with AEOeye’s AI search audit methodology template, AI visibility experiment checklist, and AI answer volatility metrics so readers can connect coverage to collection and repeatability.

What is the file contract?

The package uses two CSV files. prompts.csv must include prompt_id, status, and at least one dimension column. Every row should preserve the exact prompt elsewhere in a study’s evidence store; this small fixture focuses on coverage metadata. requirements.csv has dimension, value, and non-negative integer minimum columns.

The auditor emits sorted JSON with active and excluded counts, each requirement’s count/minimum/met flag, unmet requirements, unrecognized values, declared dimensions, and basic validation counts. Sorting makes a review diff meaningful when input row order changes. Python’s CSV module documentation explains the standard parsing behavior used here.

The package includes the Python auditor, prompt register, requirements, expected report, and README. It requires no API key, package install, network request, or provider login, so it is safe to run in CI before collection.

What does the populated example find?

The fixture has 20 synthetic rows: 18 active, one disabled, and one excluded. Its requirements intentionally contain three minimum-count gaps: consideration has 6 active prompts against a minimum of 8, pricing has 4 against 5, and fr-FR has 1 against 2. One active row uses the undeclared UnknownEngine value. These are deliberate test conditions, not observations about any real engine or market.

The deterministic report therefore says gaps_detected: true, lists the three unmet requirements, and lists the unknown value. It does not delete or “repair” rows. A human decides whether to add prompts, lower a defensible minimum, split a dimension, or document why the gap remains. The expected-report.json file is the committed contract for this exact fixture.

How do you run the exact test?

From the resource directory, run:

python3 audit_prompt_coverage.py --prompts prompts.csv --requirements requirements.csv --output actual-report.json

The normal command exits 1 because gaps are detected. Its JSON must match expected-report.json byte for byte. The documented exploratory mode is:

python3 audit_prompt_coverage.py --prompts prompts.csv --requirements requirements.csv --output allow-gaps-report.json --allow-gaps

That command exits 0 and emits the same report. The README also specifies a failure test: changing a status to pending must exit 2, because an unknown status is a register validation error rather than a coverage gap. These exit codes make the tool useful in a preflight job without pretending that a green process means a valid study.

How should you adapt requirements?

Start with the decision the study must support, then define the smallest defensible frame. For a product-comparison study, a team might require every stage and topic, with two or more active prompts per cell. For a multilingual study, define locales from the markets actually in scope and record language, region, and interface conditions separately.

Keep the requirement file versioned with the prompt register. When the frame changes, record why: a new market, a product launch, a changed engine set, or an explicit scope reduction. Store collection date, model label, account state, and source evidence under your provenance process; W3C PROV-O is a useful vocabulary for describing entities, activities, and agents without claiming hidden provider details.

Do not “balance” a frame by copying one question into many rows. Distinct prompts should represent distinct buyer intents, and repeated runs should be represented as runs in the study data rather than disguised as new prompt coverage. After the preflight passes, inspect prompt wording and deduplicate semantic intent manually.

Once the frame is fixed, preserve the answer and source context with the AI citation evidence preservation protocol so later reviewers can distinguish a coverage decision from a changed answer.

What are the limitations?

This tool audits declared coverage, not the world outside the register. It cannot estimate search volume, discover missing buyer questions, test whether an engine actually receives a prompt, or establish that an answer is accurate. It also does not infer registrable domains, model identity, user demand, causal effects, or statistical power.

The fixture is synthetic and intentionally small. Its counts should never be reported as market rates. Requirements are only as good as the analyst’s frame, and a minimum of two does not solve repeated-prompt dependence or clustered engine behavior. Record exclusions and missing answers explicitly, preserve raw responses under an approved retention policy, and use a separate coding protocol for mentions, recommendations, and citations. A coverage pass is the start of an auditable AEO study, not its conclusion.

Frequently asked questions

What does this prompt coverage auditor prove?

It proves only that active rows meet the dimension/value minimums in requirements.csv and pass structural validation. It does not prove representativeness, real search demand, prompt popularity, provider accuracy, or causal visibility.

Why do disabled and excluded prompts not count?

They remain in the register for auditability but are outside the eligible sampling frame. Counting them would make coverage look healthier than the set that will actually be collected or analyzed.

What should I do when the normal run exits 1?

Read the deterministic report, decide whether each gap is material, and add or revise eligible prompts or requirements. Use --allow-gaps only when you intentionally want a successful exploratory command; it does not hide findings.

Can passing coverage make my AEO study representative?

No. It means only that the declared cells meet declared minimums. Define the population, selection rationale, collection conditions, provenance, missing-data rules, and review process separately.

Sources

FAQ

What does this prompt coverage auditor prove?+

It proves only that an active prompt register meets the dimension/value minimums declared in requirements.csv and that its IDs, statuses, and columns pass validation. It does not prove representativeness, real search demand, prompt popularity, provider accuracy, or causal visibility.

Why do disabled and excluded prompts not count?+

They are retained for auditability but are outside the eligible sampling frame. Counting them would make a register look covered even when those prompts will not be collected or analyzed.

What should I do when the normal run exits 1?+

Open the deterministic JSON report, decide whether each gap is a real sampling problem, and add or revise eligible prompts or requirements. Use --allow-gaps only for exploratory reporting; it changes the exit status, not the findings.

Can passing coverage make my AEO study representative?+

No. Passing means the declared cells meet declared minimums. You still need a defensible population definition, prompt-selection rationale, collection protocol, provenance, and review of missing or ambiguous answers.

Sources

Is AI recommending you?

Run a free AI visibility audit and find out in under a minute.

Keep reading