---
title: "Why AI Answers Cite Reddit and YouTube More Than Your Blog"
url: https://hostmy.blog/why-ai-cites-reddit-youtube/
date: 2026-09-20
modified: 2026-09-06
lang: en
author: "Aditya Sharma"
description: "YouTube and Reddit are the two largest single sources in a 153,425 citation set. The arithmetic underneath that is more useful than the headline."
categories:
  - "AI Search"
image: https://hostmy.blog/wp-content/uploads/2026/09/hmb-card-1883-1024x538.jpg
word_count: 1549
---

# Why AI Answers Cite Reddit and YouTube More Than Your Blog

YouTube supplied 9,868 citations and Reddit 6,595 across a set of 153,425 from six AI platforms. They are the two largest single sources in the data.

Now do the division, because the headline hides the useful part. YouTube is 6.4 percent of that set. Reddit is 4.3 percent. Together the two account for 9,868 plus 6,595 citations out of 153,425, which leaves roughly nine citations in ten going somewhere else entirely.

So the finding is real and the conclusion people draw from it is wrong. Two enormous domains lead a very long tail. Your blog is competing in the tail, not against the leaders.

## The comparison is unfair in a way that matters

Reddit and YouTube are each a single domain holding hundreds of millions of pages. Your blog is a single domain holding a few hundred.

Counting citations per domain will always favour aggregators, in the same way counting library visits favours libraries over bookshops. Per-domain totals measure the size of the container.

| Source | Citations in set | Share of 153,425 |
| ------ | ---------------- | ---------------- |
| YouTube | 9,868 | 6.4 percent |
| Reddit | 6,595 | 4.3 percent |
| Everything else | 153,425 less those two, 136,962 | roughly 89.3 percent |

That bottom row is the market you are actually in. It is enormous and it is fragmented, which is a better situation than the headline suggests.

The two most-cited sources, against everything else
The two most-cited sources, against everything else
Everything else

89.3% of the record
YouTube

9,868
Reddit

6,595
Out of 153,425 citations. No other single source came close to these two.

YouTube and Reddit lead every other single source and together account for roughly 11% of the record. The other 89% is ordinary sites.

## Their pages are shaped the way the citation data rewards

The unfair comparison does not mean there is nothing to learn. Reddit threads and YouTube descriptions share four structural properties, and every one of them lines up with the numbers.

One question per page. A thread asks one thing. A video answers one thing. Your average blog post covers a topic, which is a different job. Homepages, the least focused page type of all, are cited about 4 percent of the time.

The answer sits at the top. The accepted answer on a thread is near the beginning. 41.9 percent of citations came from the first 30 percent of the page, with the mean cited position in that set of 9,968 cited sentences sitting 37 percent down.

Sentences are short. Forum replies and spoken transcripts run short by nature. Mean cited sentence length was 9.27 words and the median was 10. Across all 11,346 cited sentences a study could extract, nothing longer than 18 words was cited once, and the 6 to 10 word band carried 45.2 percent.

The register is plain. Readability in the cited set splits into two camps: 22.9 percent very easy and 20.5 percent very confusing, with only 2.6 percent in the Flesch 50 to 59 middle and a median of 66.4. Forum writing sits firmly in the easy camp. Polished marketing prose sits in the thin middle.

None of that requires you to write like a forum. It requires the answer to be early, single, short and plainly worded, and [how credibility is judged sentence by sentence](/eeat-seo-ai-answers/) follows the same shape.

## Ranking is a weaker route than it used to be

Only 23.05 percent of cited URLs ranked in the organic top ten. For ChatGPT the figure was 4.2 percent. Perplexity is the outlier at 50.3 percent overlap with organic top-10.

Ahrefs found the same movement from a different angle. Across 863,000 keywords, AI Overview citations coming from top-10 pages fell from 76 percent to 38 percent.

So the old route, rank first and get everything downstream, has weakened. A page can be cited without ranking and can rank without being cited. Those are separate outcomes now, and planning that assumes one delivers the other will be wrong more often than it used to be.

## The platforms disagree with each other more than they agree

Chasing "AI visibility" as one thing runs into a hard problem. There is no one thing.

Citation volume was AI Mode 88,392, Grok 30,676, Gemini 13,487, Copilot 8,779, Perplexity 8,562 and ChatGPT 3,529. Grok returned 35.79 citations per query. Gemini returned 7.06.

AI Mode and Gemini, both Google products, shared only 4.66 percent of cited domains.

A tool that reports a single AI visibility score across platforms this different is compressing away the thing you needed to know. Checking one platform tells you about that platform. Gemini is the one that will still show you its working, and how to read it is in [reading a Gemini citation URL](/gemini-citation-url/).

## Zero-click is a measurement problem before it is a traffic problem

The click was always a proxy. It stood in for a person getting an answer, and it was easy to count.

When an assistant answers in place, the person still got the answer and the proxy stopped counting it. Your analytics records nothing, because a crawler does not run JavaScript and a reader who never arrived leaves no session. The arithmetic for that, [what the missing click actually costs](/zero-click-search-cost/), is worth running on your own numbers.

That means the honest response is to change what you watch. Server-side crawler request counts, not sessions. Which pages get fetched by retrieval agents. Whether your best pages return 200 to those agents at all. The commands for all of that are in [the path a citation takes through your infrastructure](/ai-crawlers-request-by-request/).

Do not replace one bad proxy with a worse one. Nobody can promise you citations, and no measurement here predicts them.

## What to actually do with your own pages

Five changes, in the order they pay off.

Give each real question its own page. Take the topics currently living as sections inside long posts and let the important ones stand alone with a title that matches the question a person would ask.

Put the answer in the opening two paragraphs, then explain underneath. The explanation is what keeps a human reading and the answer is what can be lifted.

Get the claim sentence under the ceiling. Split a 28 word opener into a short claim and a separate qualification. Both survive and one becomes quotable.

Order your headings properly. Sequential headings showed a 2.8x lift. Fix the order, resist adding more.

Keep the page maintained and dated. The median cited page was 298 days old and 61.9 percent of dated cited pages came from 2025 or 2026, while stale pages were roughly 3 times more likely to lose citations. The sorting method for an old archive is in [pruning that is not busywork](/content-pruning-for-ai/).

## Being present where the discussions happen is a separate job

Reddit and YouTube being large sources does raise an obvious tactic, and it needs one caveat.

Participate where you genuinely have something to say, under your own name, answering the question asked. That is worth doing on its own merits and it occasionally produces a page that gets cited.

Manufacturing threads to be cited is a different activity, and it fails for the ordinary reason: communities are good at spotting it, and the cost of being caught is larger than the upside. This is not a growth channel to be gamed.

## Make sure your pages can be fetched before you rewrite them

All of the above assumes an agent can reach your page.

Measured on our own infrastructure, a site returned 429 to GPTBot while its robots.txt said Allow. The refusal reproduced 3 times out of 3, and static assets were refused too while robots.txt and the sitemap returned 200, placing the block at the web server layer rather than in WordPress.

Payload matters as well. Another page returned 4.7 MB of HTML against 18 KB of the same content in Markdown, roughly 260 times the size. Markdown negotiation is honoured by coding agents only, so Claude Code, Copilot Chat and CLI, Cursor, Microsoft Copilot, OpenClaw and OpenCode ask for it, while ChatGPT browse, Claude.ai, Perplexity, Gemini and Grok do not.

Check reachability before you spend a month on writing. The four-command sequence is in [what gets checked before a page is used as a source](/ai-seo-checker-curl/).

## The plugin side of this list

RankReady is a WordPress AI SEO plugin covering the machine-readable half: Article, Speakable and FAQPage schema in the raw HTML, a Markdown version of each post, robots.txt rules including Content Signals, freshness signals, an author box and llms.txt if you want one. Setup takes about five minutes.

It runs alongside Rank Math, Yoast, AIOSEO and SEOPress. Titles, meta descriptions and the XML sitemap stay with your existing plugin.

The writing half stays yours. A plugin can make a page cheap to fetch, clearly typed and honestly dated. It cannot make the answer good, and it cannot promise a citation from systems that agree with each other 4.66 percent of the time. Sorting [which claims in this category actually hold](/what-ai-seo-means/) from the ones that do not is worth doing before you buy anything.

The plugin is on [the WordPress.org directory](https://wordpress.org/plugins/rankready-ai-llm-seo/).

## The question hiding behind the headline

That bottom row is the interesting number in this whole post.

If roughly nine citations in ten go to sites that are neither YouTube nor Reddit, what is stopping one of your pages from being one of them, and is that reason a writing problem or a 429?