---
title: "Gemini Still Tells You Which Sentence It Quoted. Nothing Else Does."
url: https://hostmy.blog/gemini-text-fragments/
date: 2026-09-07
modified: 2026-09-03
lang: en
author: "Aditya Sharma"
description: "Gemini exposes the exact cited sentence on 84.1 percent of citations. Every other platform exposes zero. Google AI Mode went from 70.9 percent to 0."
categories:
  - "Data Studies"
image: https://hostmy.blog/wp-content/uploads/2026/09/hmb-card-1843-1024x538.jpg
word_count: 1945
---

# Gemini Still Tells You Which Sentence It Quoted. Nothing Else Does.

Gemini links to the precise sentence it took from your page, on 84.1 percent of its citations. Every other platform in the study exposes nothing at all.

That figure comes from an analysis of 153,425 citations across six AI platforms in May 2026. Gemini's coverage was 84.1 percent. Google AI Mode, Grok, Copilot, Perplexity and ChatGPT were all at zero.

Google AI Mode used to do it. In March 2026 it exposed the cited sentence on 70.9 percent of citations. By May the figure was zero.

So one window into what an AI actually lifted from your page is open, one closed within two months, and four were never open. If you want ground truth about which of your sentences gets quoted, Gemini is currently the only place to get it.

## What the link actually contains

Gemini's citations use a text fragment, which is a URL feature that tells the browser to scroll to and highlight a specific string of text on the destination page.

The URL looks like this:

`https://example.com/your-post/#:~:text=Stale%20pages%20lose%20citations`

Everything after `#:~:text=` is the quoted material, URL encoded. Click it and the browser jumps to that exact sentence and highlights it, and [decoding that fragment back into plain text](/gemini-citation-url/) takes one command if you would rather read fifty at once.

That string is a direct readout of what the model selected. Not a guess from correlation. Not an inference from which page ranked. The actual words it lifted, published in the link.

| Platform | Text fragment coverage |
| -------- | ---------------------- |
| Gemini | 84.1 percent |
| Google AI Mode, March 2026 | 70.9 percent |
| Google AI Mode, May 2026 | 0 percent |
| Grok | 0 percent |
| Copilot | 0 percent |
| Perplexity | 0 percent |
| ChatGPT | 0 percent |

Share of citations that reveal the exact sentence quoted
Share of citations that reveal the exact sentence quoted
Gemini

84.1%
Google AI Mode, previously

70.9%
Google AI Mode, now

0%
Everything else reports the URL and nothing about which sentence was used.

Gemini writes the quoted sentence into the citation link itself. Google AI Mode used to and stopped, which removes the only direct evidence of what was lifted.

## AI Mode dropping from 70.9 to zero is the part to worry about

Two months, one product, a feature that went from covering seven citations in ten to covering none.

AI Mode matters more than the others by volume. In the same citation set it accounted for 88,392 citations, which is 57.6 percent of the total. Gemini accounted for 13,487, or 8.8 percent.

So the surface that produces the most citations by a wide margin is now completely opaque, and the surface that still tells you anything is a fraction of its size.

No reason for the change has been published. Text fragments make it easy to audit an AI's quoting, which is not universally welcome, and they add length to every link. Whatever the motivation, the effect on publishers is the same: measurement got harder, and it got harder on the biggest surface first.

There is a second Google oddity worth putting next to this. AI Mode and Gemini, both Google products, share only 4.66 percent of their cited domains. Two systems from one company, picking almost entirely different sources, and only one of them shows its working. Anyone treating "Google AI" as a single thing to optimise for is working from a model the data does not support. More on that in [why the rank tracker stopped answering the question](/ranking-vs-being-the-answer/).

## Why this is worth more than any dashboard

Everything else in AI search measurement is inference.

You see a ranking and infer something about citation, which the data says is a weak connection at best, since only 23.05 percent of cited URLs appear in the organic top ten at all. You see referral traffic and infer which page it came from. You see a mention in an answer and infer which passage produced it.

A text fragment removes the inference. The sentence is in the URL.

That matters most for the sentence-level work, because it is the one place you can check the guidance against your own content. The mean cited sentence runs 9.27 words, the median is 10, and nothing above 18 words appeared in the set of 11,346 cited sentences. Those numbers are from other people's pages. A fragment from your own page tells you which of your sentences cleared that bar.

Read ten of them and you will change how you write the opening of a post. Reading your own quoted sentences is a different experience from reading a statistic about sentence length. The full argument is in [the length ceiling, and how it was arrived at](/18-word-ceiling/), and text fragments are how you verify it rather than taking it on faith.

## How to actually collect them

Here is the part most articles about text fragments skip, because it is inconvenient.

You cannot capture these server side. The part after the hash is never sent to your server, so it is not in your access logs and it cannot be. Browsers also keep the fragment directive out of ordinary page scripts, so there is no JavaScript workaround waiting to be written.

Collection is manual. That is the honest answer, and it is why almost nobody does it.

The routine that works:

- **Pick 10 to 15 questions your best pages should answer.** Real questions, phrased the way a person would type them, not keywords.

- **Ask Gemini each one.** Plain queries, no operators, no prompting tricks.

- **When one of your pages is cited, copy the citation URL.** The whole thing, including everything after the hash.

- **Record four fields in a sheet.** The query, your URL, the extracted sentence, and its word count.

- **Repeat monthly on the same question set.** The value is in the comparison over time, not in any single run.

Twenty minutes a month per site. After three months you have a real dataset about your own content, which is more than almost any competitor of yours will have. To grade the pages you have not been cited on yet, [scoring a post sentence by sentence](/measure-quotable-writing/) does the same counting locally.

## What reading ours changed

Three things tend to come out of doing this on your own pages, and none of them are on the list beforehand.

**The quoted sentences are almost never the ones you would have picked.** The polished thesis sentence is rarely the one selected. What gets taken tends to be a plain declarative statement of fact sitting a paragraph later, often one written quickly.

**Word counts cluster tightly and low.** Consistent with the wider finding of a 9.27 word mean, the sentences that get quoted are short, and the ones that felt too blunt while writing are disproportionately the ones selected. [Drafting deliberately under the ceiling](/write-sentences-ai-can-quote/) is mostly a matter of resisting the urge to qualify.

**Position is earlier than expected.** Consistent with the finding that 41.9 percent of citations come from the first 30 percent of a page, with the mean cited sentence sitting 37% down. The sentences chosen sit early. Material below the halfway mark rarely appears.

None of that proves a mechanism, and a handful of fragments off one site is a small sample. It does line up with a 153,425 citation dataset collected independently, which is the kind of agreement worth acting on while acknowledging what it is.

## Gemini is stingy, so plan for a small sample

One practical warning before you start. Gemini returns 7.06 citations per query in that dataset. Grok, at the other end, returns 35.79.

Five times fewer sources per answer means five times fewer chances of appearing in any given one. Ten queries may produce two of your pages, or none. That is normal and it is not a signal about your content.

Two adjustments make the exercise survive it. Use twenty questions rather than ten, on topics where your pages are genuinely the best answer available. And run the same set monthly, so a thin month is a data point rather than a conclusion.

The upside of a stingy platform is that when it does cite you, the selection was competitive. Seven sources tells you more about the sentence than thirty six does.

## The window may close, so use it now

AI Mode's drop from 70.9 percent to zero in two months is the warning. Features that expose model behaviour are not stable, and they are not promised to anyone.

Gemini at 84.1 percent could go the same way, without notice or explanation. If that happens, publishers lose the last direct readout of what these systems actually take, and everything goes back to inference.

Which argues for building the habit now. Three months of records stay useful after the feature disappears, because the selected sentences are still sentences you wrote.

Worth being clear about what this is and is not. A text fragment tells you that a sentence was quoted, not why, and it is not a metric to report as performance. It is evidence about your own writing, and it should change your editing rather than your reporting.

## Where the plumbing fits

None of this works if the page cannot be fetched in the first place, which is the first of [five layers worth auditing with curl](/five-layer-ai-seo-audit/).

RankReady, our [WordPress AI SEO plugin](https://wordpress.org/plugins/rankready-ai-llm-seo/), handles that side inside WordPress: crawler access rules, Markdown endpoints, schema, freshness dates. About five minutes to set up, and it runs alongside Rank Math, Yoast, AIOSEO or SEOPress without conflicting with any of them. It does not capture text fragments, because nothing on the server can, and any plugin claiming to do it is describing something the web platform does not permit.

Two prerequisites sit under everything above. Your server has to answer the crawler, which fails silently more often than people expect, and that failure is invisible from the dashboard. Your pages have to be fresh enough to keep being fetched, since stale pages lose citations at roughly three times the rate in the same dataset.

Check the first one before anything else. A month of collecting text fragments on a site that returns 429 to crawlers is a month producing an empty spreadsheet, and that failure is diagnosed in four commands in [an upstream rule nobody mentioned](/host-blocking-ai-crawlers/).

## What to record, and what not to

Keep the sheet boring. Query, URL, extracted sentence, word count, date. Five columns.

Two things do not belong in it. Skip estimated traffic, because a text fragment tells you a citation happened and nothing about what it was worth. Skip share of voice, because twenty queries against a platform returning around seven citations per answer cannot support a percentage anyone should act on.

The sheet is for pattern recognition in your own writing. After three months it answers questions no dashboard can: which page types get quoted, how long those sentences run, and how far down the page they sit. Almost nobody in this field has that, and it costs twenty minutes a month.

## The measurement problem in one line

Five of six platforms tell you nothing about what they took from your page. The sixth tells you almost everything, and the biggest one stopped telling you two months ago.

That is the state of measurement in AI search, and it explains why the field is so full of confident claims. When almost nothing is observable, anything can be asserted, and very little can be checked.

Gemini's 84.1 percent is a rare exception. It is a real, free, verifiable readout of model behaviour on your own content, available today, requiring nothing but a browser and twenty minutes.

If you ran those ten queries this afternoon and read the sentences that came back, which of them would you be surprised to see quoted, and what would that tell you about the way you open a post?