THE LEDGER ABD SHANTI
CITATION STRENGTH0%
SOURCE #071 VERIFIED · LIVE OCT 03, 20266 MIN READDATA PART OF MEASUREMENT

Indexed, Not Quoted: The Gap Between Stored and Used

¶ DATATHE GAP

01Read every day, quoted never

For fifteen days this blog was read constantly. Bing’s crawler requested its entries 172 times, Amazon’s 94, OpenAI’s training crawler 20, Google’s 15. Every one of those visits ended with a copy of a page stored somewhere. In the same fifteen days, the number of times an assistant fetched one of these entries because a person had just asked it a question was zero. Being stored is not being used. A crawler that copies your page has not quoted it.

Those are two different events that look identical in a log line count, and most people who check their logs add them together. They should be kept apart, because they answer different questions. One asks whether a machine knows the page exists. The other asks whether the page was useful to someone today.

// THREE STAGES OF BEING READTHIS BLOG · FIFTEEN DAYS
01STORED
Training crawlers took a copyGPTBot, ClaudeBot
20 + 7
02INDEXED
Search crawlers can find itbingbot, OAI-SearchBot, Googlebot
172 + 11 + 15
03QUOTED
A live answer fetched it for a personChatGPT-User, Perplexity-User, Claude-User
0181 on the other site
Stored and indexed, every day. Quoted, not once. Same server, same author: the other site was fetched live 181 times.

The three stages

The way I read it now is as three stages. Stored means a training crawler copied the page, which might shape what a future model knows. Indexed means a search crawler can find it, which makes it a candidate for live answers. Quoted means an assistant actually fetched it while answering a person, which is the closest thing to a citation you can see from your own server. OpenAI’s own documentation draws the same line: its live agent visits a page when a user asks a question and does not crawl the web on its own. A page can be stored and indexed every day for months and still never be the answer to anything.

¶ DATATHE CONTROL

02The same writer, a different address

The comparison I did not plan

The useful part of this data is that I have a control. Another site I run sits on the same server, follows the same writing rules and is written by the same person. In the same fifteen days, ChatGPT’s live agent fetched its pages 181 times, after removing requests that borrowed the agent’s name to probe for configuration files. OpenAI’s search crawler made 153 successful page fetches there. Same author, same rules, same server software. The writing is not the variable.

What actually differs

Three things differ, and none of them is on the page. This blog lives on a country domain that search engines geotarget to one small region. It has almost no links pointing at it. And it is young in search. The other site has a neutral domain, links, and a longer history. A live answer starts with a search, and if the search does not return your page, the answer cannot fetch it, however good it is. That is the point ranking was a location made months ago, now with a control group. An assistant cannot quote a page its search step never returned.

EXTRACTED — THE SENTENCE THIS ENTRY EXISTS FOR
Crawls prove you exist. Live fetches prove you were useful.
¶ DATAHOW TO CHECK YOURS

03Find out which stage you are stuck at

The diagnosis

Split your log by stage rather than by bot. If you see no training crawlers at all, check your robots rules and your firewall first, as in the guide to who is reading your site. If training crawlers come but search crawlers rarely do, the problem is discovery: sitemaps, internal links, and submitting pages to the engines. If search crawlers come every day and live fetches never do, you are indexed but not chosen, and the fix is upstream, in the things a search ranks on: links, the domain, and whether your page answers a question people actually ask. When the crawlers come and the live fetches do not, stop editing the page and start working on how it is found.

Count the real fetches only

One caution on reading live fetches. Scanners borrow these names to look for exposed files, so count only successful requests for real pages, and verify the address against the ranges each company publishes when it matters. A request for a page that does not exist is not a citation, whatever name it carries. More often it is somebody looking for your passwords. On the other site, 274 requests carried ChatGPT’s live agent name and 181 were real page fetches. If I had counted every line, I would have overstated it by half. The same discipline from the second scoreboard applies: the totals hide the change that matters.

Tomorrow: how to add a paywall without losing the citations a page has already earned.

// QUICK ANSWERS
>What is the difference between being indexed and being cited by AI?+
Indexed means a search crawler can find your page; cited means an assistant used it to answer someone. In fifteen days this blog was crawled hundreds of times by Bing, OpenAI, Google and others, and fetched by an assistant answering a person zero times, while another site on the same server was fetched live 181 times. Only the second kind of visit means you were the answer.
>How can I tell if ChatGPT is citing my pages?+
Look for ChatGPT-User in your server log. It is the agent OpenAI uses when it visits a page because a user asked a question, as opposed to GPTBot, which collects training data, and OAI-SearchBot, which builds the search index. Count only successful requests for real pages, because scanners borrow the name to probe for files.
>Why is my site crawled by AI bots but never cited?+
Usually because the search step of the answer never returns your page. Live answers start with a search, so a page that ranks poorly, sits on a geotargeted domain or has few links can be crawled every day and still never be fetched for a person. Work on how the page is found before rewriting the page itself.
Abd Shanti, author of CITED
VERIFIED HUMAN
Abd Shanti
GEO EXPERT · THE AUTHOR
// CITE THIS ENTRY
Abd Shanti. "Indexed, Not Quoted: The Gap Between Stored and Used" CITED, Entry 071, Oct 3 2026. unknown.ps/blog/indexed-not-quoted/
// RELATED ENTRIES
QUOTE COPIED — CITE FREELY