AI Citation Source Concentration Calculator: HHI, Top-Source Share, and Effective Domains

This package calculates descriptive source concentration from answer-level citation records. It does not prove that an engine prefers a domain, that a cited page is correct, or that a concentrated source set is bad. The included 30-row CSV is synthetic and exists to make the method inspectable and reproducible, not to report live provider behavior.
Table of contents
- What does source concentration reveal?
- Which metrics does it calculate?
- How does host normalization work?
- What is in the package?
- What does the 30-row fixture produce?
- How should teams compare snapshots?
- Limitations
- FAQ
- Sources
What does source concentration reveal?
Source concentration answers a narrow question: how much of a citation set is supplied by the same normalized host? A report with 30 URLs can still have only a few underlying hosts, so URL count alone can make evidence look more diverse than it is.
That matters in AI-search visibility work. If a brand audit sees citations from one publisher repeatedly, the finding may be useful but fragile: a page change, access change, or retrieval shift could affect many answers at once. Conversely, a concentrated set is not automatically suspicious. A regulator, standards body, or product documentation site may be the right primary source for a specific question.
Use concentration beside claim-level review, not instead of it. A citation failure taxonomy can classify whether a source supports a claim, while this calculator describes the shape of the source pool. The AI citation parser test corpus helps preserve extraction boundaries before counting anything.
Editorial image: Pexels photo 5905445; it is not a test screenshot.
Which metrics does it calculate?
The calculator counts each valid citation row once, then computes the following summary. Shares use total citation rows as the denominator.
| Metric | Meaning |
|---|---|
| Total citations | Number of accepted CSV rows |
| Unique hosts | Number of distinct normalized hostnames |
| Host share | A host's citation count divided by total citations |
| Top-one / top-three share | Share held by the largest one or three hosts |
| HHI (0–1) | Sum of squared host shares |
| HHI (0–10,000) | The same HHI multiplied by 10,000 |
| Effective domain count | 1 / HHI; a concentration-adjusted diversity lens |
The HHI is a mathematical summary, not an AEO grade. The 2023 US DOJ/FTC Merger Guidelines discuss concentration in a competition context; that context should not be imported as a universal threshold for citation quality. This page does not label a score “healthy,” “unsafe,” or “good.”
How does host normalization work?
The script accepts only absolute http and https URLs. It lower-cases the parsed hostname, removes a trailing dot, and removes one leading www.. Thus HTTPS://WWW.Example.org/page counts as example.org.
It intentionally does not implement a public-suffix list. news.example.com and example.com remain separate hosts, as do prov.w3.org and w3.org. That boundary is explicit because silently guessing registrable domains can create a different measurement. URL syntax is grounded in RFC 3986, and parsing uses Python's standard urllib.parse.
Raw URLs remain in citations.csv for review; only the normalized host enters the counts. The script also rejects embedded URL credentials, missing hosts, non-HTTP schemes, blank IDs, duplicate IDs, missing columns, and empty input. prompt_id and engine are required provenance fields even though they are not grouped in this simple report. Preserve these relationships in a broader evidence model; W3C PROV-O is a useful vocabulary reference.
What is in the package?
Download the calculator source, 30-row synthetic citations CSV, deterministic expected report, and run instructions. The package uses only Python 3's standard library and needs no API key or network request.
Run this from that directory:
python3 calculate_concentration.py citations.csv > actual.json
cmp -s actual.json expected-report.json
The second command is a byte-for-byte check, not a visual approximation. The documented malformed-URL test changes one temporary copy to ftp://...; the calculator must reject it with exit status 2. That failure test is important because a concentration number from an invalid URL set is worse than no number.
What does the 30-row fixture produce?
The synthetic fixture contains ten citations from nist.gov, five each from docs.python.org, justice.gov, and rfc-editor.org, three from w3.org, and one each from prov.w3.org and standards.w3.org. After normalization, the exact report is:
| Output | Value |
|---|---|
| Total citations | 30 |
| Unique hosts | 7 |
| Top-one share | 0.333333 |
| Top-three share | 0.666667 |
| HHI (0–1) | 0.206667 |
| HHI (0–10,000) | 2066.666667 |
| Effective domain count | 4.838710 |
These numbers are arithmetic checks for the fixture, not observations about NIST, Python, RFC Editor, DOJ, or W3C citation frequency. A real report should retain the query, engine, collection time, answer ID, raw citation, and URL access result so a reviewer can explain the denominator.
How should teams compare snapshots?
Freeze the collection protocol before comparing periods. Keep the same prompt set, engines, answer-availability rule, citation extraction version, URL normalization rule, and host grouping boundary. If any changes, report them beside the metric rather than presenting a clean time series.
Break out results by engine or prompt family when the sample supports it. A pooled top-source share can rise simply because one engine contributed more rows. Also report row counts and missing or inaccessible URLs; otherwise a lower HHI may merely reflect a changed parser or denominator.
Pair the summary with the AI visibility metrics dictionary and an AI search audit. The practical question is not “is HHI low?” but “which source dependencies should we inspect, diversify, or explain for this buyer-question set?”
For a complementary view of breadth across answer sources, compare the results with the AI search source diversity metrics guide and keep both denominators visible.
Limitations
This is a host-level descriptive calculator. It does not resolve redirects, follow canonical links, identify page ownership, collapse subdomains, assess source authority, test factual entailment, detect duplicate page content, or infer why an engine selected a URL. It cannot establish causality or provider ranking behavior.
The fixture is synthetic and deliberately small. Its expected output is a regression artifact, not a benchmark, population estimate, antitrust analysis, or claim about live AI systems. Effective domain count is a transformation of HHI, not a literal count of independent publishers. For consequential research, preserve raw evidence and have a reviewer inspect edge cases and the sampling frame.
FAQ
What does this citation concentration calculator measure?
It summarizes host-level concentration in an answer citation CSV using total citations, unique hosts, top-source shares, HHI, and effective domain count. The included rows are synthetic.
Does a high HHI mean an AI answer is wrong?
No. HHI describes source concentration, not correctness, authority, independence, or recommendation quality. A focused answer may legitimately rely on one primary source.
Does the script combine subdomains into registrable domains?
No. It deliberately counts normalized hosts. It lower-cases the hostname and removes one leading www., but does not guess public-suffix boundaries or ownership.
How do I run the calculator?
From the resource directory, run python3 calculate_concentration.py citations.csv and compare the output byte-for-byte with expected-report.json using cmp -s.
Sources
FAQ
What does this citation concentration calculator measure?+
It summarizes host-level concentration in an answer citation CSV using total citations, unique hosts, top-source shares, HHI, and effective domain count. The included rows are synthetic.
Does a high HHI mean an AI answer is wrong?+
No. HHI describes source concentration, not correctness, authority, independence, or recommendation quality. A focused answer may legitimately rely on one primary source.
Does the script combine subdomains into registrable domains?+
No. It deliberately counts normalized hosts. It lower-cases the hostname and removes one leading www., but does not guess public-suffix boundaries or ownership.
How do I run the calculator?+
From the resource directory, run python3 calculate_concentration.py citations.csv and compare the output byte-for-byte with expected-report.json using cmp -s.
Sources
Is AI recommending you?
Run a free AI visibility audit and find out in under a minute.