Retrieved Often vs. Cited Hard: The AEO Metric Nobody’s Talking About
Retrieval rate and citation rate get treated as the same metric. They’re not, and real Peec.ai client data shows why small, specific sites can out-cite the giants that out-rank them.
Jul 21, 2026 · ~14 min readMost AEO advice treats retrieval as the finish line. Get your page pulled into as many AI answers as possible, and you’ve won. I used to think the same thing.
Then I started looking at the same account from a different angle: not how often a page gets pulled in, but once it’s pulled in, how often the AI actually quotes it. Those turned out to be two different games, with two different winners, and the winners weren’t who I expected.
I’ve had this feeling for a while. Working with clients, I’d notice AI engines citing small, oddly specific websites over the big names you’d assume would win by default. A regional news site over Vogue. A personal blog over a major retailer. I didn’t have a tidy explanation. I still don’t, entirely. But I went and checked, properly, instead of guessing. Here’s what I found.
Two numbers that aren’t the same thing
Peec.ai (and most AEO trackers) report two metrics that get conflated constantly.
Retrieval rate is how often a page gets pulled into an AI’s answer at all. Citation rate is, once retrieved, how many times it actually gets quoted inline, as a source, in the answer text. (See our full breakdown of AI visibility metrics if these terms are new to you.)
These sit on different axes. A page can get retrieved constantly and cited rarely, because the AI keeps looking at it but doesn’t trust it enough to quote it directly. Or a page can get retrieved occasionally and get cited nearly every single time it shows up. The second pattern is the interesting one, and it’s the one nobody optimises for, because nobody’s watching it.

The pattern: small sites getting cited harder than giants
Here’s the data from one consumer beauty brand’s project I track: 50 tracked prompts, running since March 2026, across ChatGPT, Perplexity, and Google AI Overviews.
The domains with the highest citation rate in the whole dataset aren’t the ones you’d guess.
| Domain type | Retrievals | Citation rate |
|---|---|---|
| Small personal beauty blog | 18 | 4.44 |
| Random lifestyle blog (wrong category, more on that below) | 25 | 4.32 |
| UK grooming/affiliate blog | 52 | 3.38 |
| Local hair salon’s blog | 50 | 3.38 |
| A DTC skincare brand’s own blog | 54 | 2.89 |
Now compare that to the sites carrying the actual volume.
| Domain | Retrievals | Citation rate |
|---|---|---|
| Vogue.com | 3,715 | 1.26 |
| Allure.com | 3,180 | 1.12 |
| Google.com (SERP snippets) | 6,233 | 1.25 |
Vogue gets pulled into an AI answer roughly 200x more often than that small personal blog. But per retrieval, the small blog gets quoted 3.5x more intensely. A citation rate above 1.0 means a domain often earns more than one citation per retrieved chat: multiple passages from the same page, quoted more than once in a single answer. That isn’t a rounding error. It’s the AI treating one source as noticeably more worth quoting than the other.
I didn’t want to leave this at eight hand-picked examples, so I pulled the retrieval count and citation rate for 1,000 domains in this project in a single query, ordered by volume, and ran a proper check rather than eyeballing a chart.
Here’s what the full dataset actually shows: retrieval volume and citation rate do have a relationship, but it’s the opposite of what the two tables above suggest at a glance, and it’s weak.
Retrieval volume explains only about 5% of the variation in citation rate

Each dot in the chart above is one of the 1,000 domains pulled directly from this project’s own domain report, plotted by retrieval count (log scale) against citation rate. The line is a linear regression fitted across all of them. On a log scale, the correlation comes out to r=0.22, R²=0.048. In plain terms: domain volume explains under 5% of why citation rates differ from one domain to the next. The other 95% is coming from somewhere else, which is what the rest of this piece is about.
If anything, average citation rate rises with volume, not falls. By retrieval-count tier: under 100 retrievals averages 0.47, 100 to 300 averages 0.66, 300 to 1,000 averages 0.72, and over 1,000 averages 0.84 once you set the brand’s own domain aside (0.93 with it included). Bigger domains are, on average, the safer bet, not the worse one.
What volume actually controls is the ceiling, not the average, but that ceiling needs one correction first. The single highest score among domains above 1,000 retrievals belongs to the tracked brand’s own site, at 2.26. That’s not a fair comparison: a brand’s own domain naturally shows up more in answers about itself. Strip it out, and the real ceiling for independent domains above 1,000 retrievals drops to about 1.26.
The regression itself barely moves when you remove that domain: r=0.21 instead of 0.22. So the “5% of the variation” finding holds either way. Only the specific 2.26 ceiling number was skewed by one data point.
The real outliers, the 2.5-to-4.4 range in the small-domains table earlier, only ever show up among genuinely independent, low-to-moderate-volume domains. And most small domains never get close: in a sample of just over 300 domains retrieved exactly once, nearly nine in ten were never cited at all.
So small isn’t a shortcut to a high citation rate. Most small domains are cited unreliably or not at all. It’s just the only place a high citation rate is possible. That reframes the question: not “why do small sites beat big ones,” but “why do a handful of small sites do something almost none of their peers manage.”
Does this look the same on every engine?
I’d been treating ChatGPT, Perplexity, and Google AI Overviews as one pooled dataset up to that point, which is a real gap, because they don’t behave the same way. I split the domain-level data by engine and ran the same regression three times, on roughly 300 to 400 domains per engine.
Google AI Overview shows the clearest relationship of the three: r=0.33, R²=0.11. Still weak by normal standards, but three times stronger than the pooled figure. Citation rate climbs fairly steadily with volume, from a 0.49 average under 100 retrievals to 1.02 average above 1,000. Google rewards volume more consistently than the other two engines, though the intensity stays modest throughout. Nothing in this sample scores above roughly 1.3.
ChatGPT sits in the middle: r=0.27, R²=0.07. The volume-reward pattern is there but noisier, and the ceiling runs higher: a few domains above 1,000 retrievals sit above 2.0, including the tracked brand’s own site (2,027 retrievals, 2.64 citation rate) and reddit.com (1,028 retrievals, 2.02 citation rate). The brand’s own domain shouldn’t really count here either, for the same reason as above: it naturally shows up more in answers about itself. Strip it out, and this engine’s R² drops from 0.07 to roughly 0.06, a small but real difference given only 400 domains in this slice.
Perplexity is the real outlier, and not in the direction I expected: r=0.07, R²=0.004, essentially no relationship between volume and citation rate at all. A domain retrieved 20 times has close to the same odds of a 5.0 citation rate as a domain retrieved 500 times. Perplexity also runs a noticeably higher citation rate across every volume tier than the other two engines: 0.63 average even under 100 retrievals, against ChatGPT’s 0.48 and Google’s 0.49. That fits how Perplexity is built: it shows inline numbered citations for nearly everything it references, more aggressively than a ChatGPT answer or a Google AI Overview snippet tends to.
One caveat worth flagging: the top-volume buckets are thin, 8 domains above 1,000 retrievals for ChatGPT, 4 for Google, only 2 for Perplexity, so treat the ceiling figures as directional rather than precise. The overall pattern (Google clearest, Perplexity flattest, ChatGPT in between) is the finding I’d stand behind; the exact decimal ceilings less so.
The practical read: if you’re optimising across all three engines with one strategy, you’re optimising for an average that doesn’t describe any of them particularly well. Google rewards being a big, established domain more than the other two do. Perplexity barely cares about volume at all, small and specific has close to the same shot as big and established there. ChatGPT sits between the two, with more room at the top for domains that combine real size with genuine authority.
It’s not just my client, though I want to be careful about the sources
I looked for outside confirmation on two different things: whether this general pattern, size not determining citation strength, shows up beyond one project, and whether there’s anything more rigorous than a vendor’s marketing blog behind it.
There is one real academic source in this territory: Aggarwal et al., “GEO: Generative Engine Optimization” (Princeton, accepted to KDD 2024, a genuine peer-reviewed venue). It doesn’t test my specific question, retrieval volume against citation intensity, but it does show, with a proper benchmark and controlled evaluation, that content-level changes can shift visibility in generative engines by up to 40%, independent of a site’s existing authority. That’s real, checkable support for the broader idea that something other than domain size is driving these outcomes, even if it’s not a direct test of my numbers.
Search Engine Land, an independently edited trade publication, reported in March 2026 on a 30-million-source analysis (run by Peec AI, the same tool behind this data) showing Reddit, YouTube, and LinkedIn as the most-cited domains across ChatGPT, Google AI Mode, Gemini, and Perplexity, ahead of the traditional big editorial names. That’s about total citation share, not the retrieval-to-citation ratio I’ve measured, so it doesn’t confirm my regression directly. But it’s the same underlying theme, from an outlet with actual editorial oversight rather than a vendor writing about its own product: conventional size and authority aren’t running this the way they used to.
Worth being precise about what “Reddit and YouTube dominate AI citations” actually means, though, because it’s not quite what my regression is measuring, and the pooled number actually hides the more interesting story. YouTube (9,293 retrievals pooled) sits at a citation rate around 0.85, close to what its volume tier predicts, with a modest per-engine spread: 0.38 on ChatGPT, 0.49 on Perplexity, 1.12 on Google. Reddit is the one that really breaks the “it’s just volume” story. Pooled, it looks unremarkable at 0.91, almost exactly where its 2,785 retrievals predict. Split by engine, it swings from 0.18 on Perplexity, essentially never converting relative to how often it’s retrieved there, to 2.02 on ChatGPT, matching the intensity of the small niche blogs earlier in this piece, with Google in between at 0.53. The pooled average isn’t describing Reddit’s citation behaviour so much as averaging over three quite different ones. Both domains still rack up huge total citation counts regardless, 4,006 for YouTube, 2,132 for Reddit pooled, because they’re retrieved constantly, not because each individual retrieval converts unusually well. That total-citation-count figure is what “most-cited domain” rankings like the Search Engine Land piece above are usually reporting. It’s a real number, but it’s citation share, not citation rate, and the two aren’t the same question. Wikipedia does something closer to what Harvard and the American Academy of Dermatology do below: barely retrieved in this niche at 51 times in four months, but cited at 2.53 when it does show up, the same extreme range as the small blogs in the table earlier.
Past those two, most of what’s published on “which sites get cited by AI” in 2026 comes from companies selling AEO and GEO monitoring tools: Clairon, Digital Applied, Contently, and others. They publish specific-sounding numbers, correlation coefficients, citation-share percentages, as marketing content for their own products, without public raw datasets or independent audit, and some of their figures conflict with each other. I’m flagging those as background colour only, not evidence, and the regression above, run on this project’s own data, is doing the actual work in this piece.
So I went and actually looked at these sites myself
This is the part that mattered most, and the part I got wrong on the first pass. My first instinct was to explain the pattern as the AI rewarding the cleanest keyword match rather than quality. It’s a clean, slightly cynical story. It’s also not what I found when I opened roughly 20 of these small, high-citation-rate URLs and actually read them.
Most of them aren’t “uncurated” in any meaningful sense. They’re small, not bad. A DTC haircare brand writing about its own product on its own blog: real brand, real product, on topic. A local hair salon’s blog post about styling, genuinely written for a real, named business, thin but accurate. An independent hair transplant clinic’s page, authored by a named, credentialed doctor, covering actual clinical treatments. A personal lifestyle blogger in her sixties, writing in the first person about her own fifteen-plus years of going grey, with specific detail and a genuine comment section of real readers.
None of that is low quality. It’s narrow and small-scale, and often more specific to the exact question than a 26-tips listicle on a major glossy site trying to cover a dozen questions at once. That lines up with the regression: since volume explains only about 5% of the variation in citation rate, the other 95% has to come from somewhere, and on the evidence of these pages, a lot of it is specificity and a direct answer to one narrow question, not domain size at all.
There’s also a more mechanical explanation worth considering, one I can’t fully confirm or rule out with what I have in front of me. AI systems built on retrieval-augmented generation typically chunk a page into passages before deciding what to pull into an answer. A long, script-heavy editorial page might contribute only one chunk to a given answer, while a short, focused blog post can fit almost entirely into the retrieved context and get quoted from more than once in the same response. If that’s part of what pushes citation rate above 1.0 for small pages, it’s a property of how the retrieval pipeline handles document length, not necessarily a judgment about specificity or quality. I don’t have visibility into Peec’s or any individual model’s chunking logic, so I can’t separate this from the specificity story with the data available, only flag it as a real alternative, or additional, mechanism.
But two of the roughly 20 sites I checked were something else entirely, and this is the part that should worry anyone taking AEO seriously.
One domain in the top “citation concentration” table (retrieved 35 times, cited 104 times, a 3.15 citation rate) is a B2B wholesale marketplace: a trade-sourcing platform for factory buyers, not a beauty or wellness site at all. It has, or had, a blog post about shampoo for thinning hair sitting on a domain that otherwise sells wholesale sourcing services. That isn’t organic topical authority. It’s a parasite page riding an unrelated domain’s crawl footprint, and it’s getting cited by AI engines as a legitimate hair-care source dozens of times over four months.
A second domain (retrieved 52 times, cited 141 times, a 2.71 citation rate) no longer exists as the site the AI is citing. The URL now redirects to a live Indonesian gambling operation with nothing to do with hair care. Whatever content used to sit there when it was crawled and cited, the domain has since been compromised or repurposed, and the citation trail still points at it.
That’s two of roughly 20: not a majority, not proof the whole pattern is fake, but a real, uncomfortable data point. AI engines are citing hijacked or entirely unrelated domains as sources for consumer health and beauty questions, at a rate that puts them near the top of the citation-intensity table, indistinguishable in the data from a legitimate small niche blog. Nothing in the retrieval and citation numbers alone tells you which is which. You have to go and look.
The exception that complicates the pattern
Not every part of the pattern is small beats big, and four real examples show why size alone was never the right axis. Harvard’s own site sits at 185 retrievals and a 2.46 citation rate, squarely inside the small-domain range despite being about as institutionally authoritative as a domain gets. The American Academy of Dermatology does something similar at a much bigger scale: 972 retrievals, properly high-volume territory, and still a 1.89 citation rate, well above where any comparably-sized generalist editorial site sits. Woman & Home, a mainstream lifestyle title with no institutional authority at all, hits 2.02 at 181 retrievals, close to Harvard’s volume from a completely different kind of domain. Wikipedia rounds this out from the reference-source side: barely retrieved in this niche at 51 times across four months, but cited at 2.53 when it does show up, general-purpose encyclopedic trust doing the same job here that medical-institution and lifestyle-title trust do above.
So here’s the finding in full, three parts. Small and specific can get cited hard, though most small domains don’t. Big and generalist gets cited lightly per retrieval, consistently. Big and genuinely authoritative on the topic, or trusted as a reference regardless of niche, gets cited hard and often, at any volume. Domain size on its own predicts almost nothing, about 5% of the variation. Topical fit and real authority predict most of the rest, and a hijacked domain can occasionally fake the first one.
What this means if you’re the one deciding where to invest
This isn’t for small Amazon sellers chasing a quick win. If that’s you, this piece isn’t really aimed at your situation. It’s for the people running an actual marketing programme, a Head of Marketing, a CMO, an agency owner, who assume their domain’s existing authority or Google rankings will carry over into AI citations.
They won’t, automatically. The practical takeaway is to audit your best-performing pages for whether they answer one question completely, or partially answer ten. A single-purpose, tightly scoped page, even on a smaller property, is competing on far more even footing against a big generalist competitor in AI citations than it ever was in classic Google rankings. For the detailed playbook on structuring content for this, see How to get your brand cited by AI search engines.
There’s a channel-allocation version of this too. Most AEO advice defaults to “get present on Wikipedia, Reddit, and YouTube” as if that’s channel-agnostic advice that applies to every market equally. It isn’t. In this market, Wikipedia was retrieved only 51 times across four months of tracking, against 9,293 for YouTube and 2,785 for Reddit, out of tens of thousands of retrievals overall. If your own category’s prompts don’t send AI engines to a given platform in meaningful volume in the first place, no amount of getting cited well there moves anything. Before committing budget to a specific channel, pull your own project’s domain report and check whether AI engines actually retrieve from it for your category at all. It’s a five-minute check against actually spending the money to find out the hard way.
It’s worth weighting that by engine too, not just by market. If a meaningful share of your visibility runs through Perplexity, raw size matters less there than almost anywhere else, so a well-targeted small property is a comparatively efficient bet. If Google AI Overview carries more of your traffic, size and established volume matter more, so building genuine scale is better spent there than on Perplexity.
And check who else is getting cited alongside you. If a hijacked domain or an off-topic parasite page is sitting in the citation set for your category’s key prompts, that’s worth flagging to whoever runs your AEO programme, and worth asking your tracking tool whether it can catch it automatically, because right now most can’t.
How to check whether this holds in your market
Everything above is one project, one vertical, four months of tracking. Before you build a channel strategy on it, run the same check on your own account. Here’s the order I’d do it in:
- 1Pull your own domain report. Sort by retrieval count descending and export retrieval count and citation rate for as many domains as your tracker gives you.
- 2Set your own brand’s domain aside before judging any ceiling. A brand’s own domain will always look artificially strong in answers about itself; exclude it before drawing conclusions about what an unrelated domain can achieve.
- 3Bucket by volume before you eyeball the top of the list. Group domains into rough tiers (under 100, 100 to 300, 300 to 1,000, over 1,000 retrievals) and compare average citation rate per tier, not just the single highest number.
- 4Split by engine before you conclude anything. A pooled average can hide a huge per-engine swing, the way it did for Reddit in this dataset. Run the same check separately for each engine you track.
- 5Adjust for how heavily your vertical is guardrailed. Categories like finance, health, and legal lean much harder on institutional and government sources; expect a lower ceiling for small independent domains there than in a category like beauty, though I haven’t tested that directly, only reasoned toward it from how these models are generally described to behave.
None of this requires anything you don’t already have if you’re using an AEO or AI-visibility tracker. It’s the same domain report, just read in the right order.
Conclusion
The mystery I started with, why the AI sometimes picks small, seemingly random sites, turns out to be less mysterious than I assumed, and uncomfortable in a different way than I expected.
Most of it isn’t random. Specificity and topical fit are doing most of the work, and once you see it, that’s a fairly reassuring mechanism. What’s less reassuring is that the same signal rewarding a genuine small expert also rewards a compromised domain serving gambling spam, if it happens to still carry the citation trail from before it was hijacked. The system can’t yet tell the difference between small and genuinely useful, and small and quietly broken. Until it can, that’s worth checking for yourself, one domain at a time.
Methodology note
Single client project, roughly 50 tracked prompts, from 31 March to 21 July 2026, US market, hair and beauty vertical only. Directional, not a universal law. The roughly 20-site manual review was a targeted look at the highest citation-rate domains specifically, not a random sample, so the two flagged domains represent about 10% of that targeted sample, not an estimate of prevalence across the wider web. The pooled regression (r=0.22, R²=0.048) is run on 1,000 domains pulled directly from this project’s own domain report in a single query ordered by volume, not a curated subset. The “close to nine in ten single-retrieval domains were never cited” figure comes from a separate, targeted pull of just over 300 domains at exactly one retrieval each, since the volume-ordered pull above doesn’t reach that far down the tail. The per-engine regressions were run separately on roughly 300 to 400 domains per engine (ChatGPT, Perplexity, Google AI Overview), each ordered by retrieval count and pulled directly, not curated; the top-volume buckets in that breakdown are thin (2 to 8 domains above 1,000 retrievals depending on engine), so treat the exact ceiling figures per engine as directional. Outside vendor research is cited for directional colour only; the specific decimal figures from those sources are unverified marketing claims, flagged as such in the text. The two compromised domains are described generically, with no domain name given, so readers can’t navigate to the gambling site through this post.
If you want to sanity-check this on your own data: if you use Peec.ai or another AI visibility tracker, open your own domain report, sort by retrieval count descending, and compare the citation rate at the top of the list (thousands of retrievals) against domains near the bottom of the “meaningful volume” range (15 to 70 retrievals). If this pattern holds for your market too, you’ll likely see something similar: few if any domains above roughly 1,000 retrievals clearing a 2.3 citation rate, while several domains under 70 retrievals sit at 2.5 and above.