Skip to content
All articles
AEO Operations

AI Answer Uncertainty Codebook: 10 Labels for What the Model Does Not Know

By the AEOeye editorial team·Updated Sep 16, 2026·9 min read
Laptop displaying charts beside notes for an AI answer uncertainty review.
Photo by Lukas Blazek on Pexels

An AI answer can sound cautious, decisive, or polished while its evidence state remains unknown. The practical fix is to annotate the observable uncertainty separately from factual correctness: preserve the exact span, name the uncertainty pattern, and let a reviewer record confidence against cited evidence. This AEOeye operating proposal supplies ten labels for that job; it is not an industry or provider standard.

Uncertainty matters in an AI visibility audit because “not sure,” “probably,” and an unsupported definite statement lead to different decisions. NIST’s AI Risk Management Framework asks teams to define context, measure risks, and document limitations; its Generative AI Profile also emphasizes that evaluation must account for the distinctive risks of generated content (NIST AI RMF, NIST AI 600-1).

Table of Contents

A team reviewing a documented workflow on a laptop, representing evidence review and annotation. Photo by Yan Krukau on Pexels. Photo by Yan Krukau on Pexels; used as an editorial visual for research workflow.

What exactly should an uncertainty annotation capture?

Capture four separate fields: the exact answer span, the model’s expressed uncertainty, the reviewer’s evidence-based confidence, and the factual verdict. A single label should describe the dominant uncertainty pattern in that span, not the reviewer’s emotional reaction to its tone.

Use a record such as: answer_span, label, reviewer_confidence, evidence_urls, factual_status, scope, timestamp, and notes. Copy the answer verbatim, including qualifiers such as “may,” “typically,” or “I cannot verify.” The AI citation data schema and evidence preservation protocol provide useful surrounding fields for keeping the record auditable.

The distinction between expression and truth is essential. Research on linguistic calibration studies whether expressed doubt tracks correctness, while citation-evaluation work shows that a cited answer still needs claim-level support (Kadavath et al.; ALCE). A hedge may accompany a true statement; a confident sentence may be false.

What are the 10 uncertainty labels?

The following ten labels are AEOeye’s proposed top-level codebook. Examples are fictional and hypothetical; they are not claims about any named model, company, or product. Assign one primary label per span and add a note when another label is also plausible.

Code Label Operational definition Fictional example
U1 Explicit uncertainty The answer directly states doubt, lack of confidence, or inability to verify. “I’m not certain whether Northstar Widget launched in May.”
U2 Probability or range The answer expresses uncertainty numerically or with a bounded interval, without asserting that the estimate is a measured fact. “There is a 40–60% chance the fictional launch occurred in May.”
U3 Source conflict The answer identifies materially inconsistent sources or claims and does not silently choose one. “Record A says April; record B says June, so the date is disputed.”
U4 Temporal uncertainty The answer signals that a fact may have changed, is stale, or cannot be dated confidently. “The fictional store’s current hours may differ from this older listing.”
U5 Scope limitation The answer limits a claim to a population, place, dataset, version, or condition instead of generalizing it. “This fictional result applies only to the three stores in the sample.”
U6 Conditional answer The answer depends on an explicit assumption, scenario, or missing condition. “If the fictional plan includes shipping, the total is higher.”
U7 Insufficient evidence The answer says the available material is not enough to support a conclusion, without necessarily expressing personal doubt. “The supplied fictional records do not establish which vendor caused the delay.”
U8 Refusal or abstention The system declines to answer, or abstains, because it cannot safely or responsibly make the requested claim. “I can’t identify the fictional person from that private description.”
U9 Unmarked uncertainty The reviewer finds an uncertainty-bearing claim—such as a forecast, estimate, or disputed proposition—expressed without a qualifier that would help a reader interpret it. “Northstar will certainly dominate the fictional market next quarter.”
U10 False certainty The answer uses definite, high-confidence language that conflicts with available evidence or omits a material unresolved limitation. “Northstar launched on May 3,” when the fictional records explicitly disagree.

U1–U8 describe visible qualifications or evidence states. U9 and U10 are reviewer flags for missing calibration or overstatement; they should not be assigned merely because a claim later turns out to be wrong. The IPCC’s uncertainty guidance demonstrates why calibrated terms need shared meanings and explicit context, while research on communicating uncertainty cautions that readers can interpret the same word differently (IPCC AR6 WGI Chapter 1; Ho et al.).

How can reviewers separate uncertainty from factual error?

First label what the answer says about its own knowledge; then check the proposition against evidence. Record both fields. For example, “I may be wrong, but the fictional date is June” can be U1 plus factual_status: supported, contradicted, or unresolved. Conversely, an unqualified statement can be factual_status: supported without becoming well-calibrated.

Use a three-pass review:

  1. Span pass: mark the smallest complete clause carrying uncertainty or certainty. Preserve nearby qualifiers and the claim it modifies.
  2. Evidence pass: follow each citation, compare the claim with the source, and note whether the source is accessible and in scope. The AI citation URL normalization rules help prevent a changed URL from being mistaken for changed evidence.
  3. Calibration pass: compare language with the evidence state. Do not convert reviewer confidence into a claim about the model’s hidden probabilities.

If sources disagree, use U3 even when one source seems more authoritative, then record the adjudication rationale. If a source is inaccessible, mark that access limitation rather than guessing that the claim is false. Citation research such as RAGTruth treats unsupported or contradicted spans as annotation problems requiring explicit evidence, not as tone judgments (RAGTruth).

How should a team report uncertainty in an AI audit?

Report label counts alongside valid-answer and factuality measures, never as one blended “confidence score.” Break results down by engine, model label, prompt family, locale, retrieval condition, and collection date; include the denominator and missing-data rule. The AI search experiment reporting checklist is a good companion for that protocol.

An audit row might say: U4 | temporal uncertainty | reviewer confidence: high | factual status: unresolved | evidence: two dated pages | collected: 2026-09-16. At aggregate level, publish the number of annotated spans and the number of independently reviewed spans. If two reviewers disagree, retain both decisions and an adjudicated result instead of overwriting the disagreement.

For AEOeye, a proposed dashboard convention is to show U1–U8 as expression/evidence states and U9–U10 as calibration-risk flags. This is an operating choice, not a standard. Teams may rename or split labels, but should version the guide and preserve historical annotations; the AI search audit methodology template can hold those protocol details.

What are the limitations of this codebook?

This codebook cannot reveal hidden model confidence, prove that a hedge was causally generated by uncertainty, or guarantee that a cited source supports every clause. It is designed for observable text and available trace evidence. Provider interfaces, retrieval access, languages, and answer formats can change, so labels require a dated protocol.

The ten labels also do not replace domain review. Medical, legal, financial, safety, and identity questions may need specialist adjudication and stricter escalation. A reviewer should mark unresolved rather than force a verdict when the evidence set is incomplete. Calibration language can reduce overstatement without making the conclusion accurate; conversely, a correct conclusion can be expressed too confidently.

Finally, fictional examples are training aids only. Before using production data, define privacy handling, redaction, reviewer access, and retention rules. Keep the raw answer and source snapshots under the team’s approved governance process, then use AEOeye’s audit to test repeatable buyer questions with those limits visible.

Frequently asked questions

Is U9 the same as a hallucination?

No. U9 flags missing uncertainty language around a claim whose uncertainty matters. Factuality requires a separate evidence check; the claim may be true, false, or unresolved.

Can one span receive two labels?

Use one primary label for consistent counting, then record secondary labels in notes. For example, a conditional answer can also contain a temporal limitation.

Why preserve the exact wording?

Small qualifiers change meaning. Retaining the exact span lets another reviewer inspect whether “may,” “likely,” or “cannot verify” was present and what it modified.

Should probabilities be treated as facts?

No. A probability or range is an expressed estimate. Record its label, assumptions, method, and evidence separately; do not invent calibration data that the answer does not provide.

FAQ

Is this an official uncertainty standard?+

No. This is AEOeye's operational proposal for making observable AI-answer uncertainty easier to annotate and compare. It is not a standard issued by NIST, the IPCC, or any AI provider.

Is an uncertain answer automatically wrong?+

No. Uncertainty is a property of expression or evidence state, while factual error is a claim that conflicts with the best available evidence. Annotate both dimensions separately.

Should reviewers score the model's confidence?+

Reviewers should record the model's expressed uncertainty and their own evidence-based confidence as separate fields. Do not treat a confident tone as proof, or a hedge as proof of error.

How should a team use this codebook?+

Preserve the exact answer span, assign one primary uncertainty label, record reviewer confidence and evidence links, then report results by engine, prompt, locale, and date. Keep fictional test examples separate from production observations.

Sources

Is AI recommending you?

Run a free AI visibility audit and find out in under a minute.

Keep reading