Indexed, Not Quoted: The Gap Between Stored and Used
01Read every day, quoted never
For fifteen days this blog was read constantly. Bing’s crawler requested its entries 172 times, Amazon’s 94, OpenAI’s training crawler 20, Google’s 15. Every one of those visits ended with a copy of a page stored somewhere. In the same fifteen days, the number of times an assistant fetched one of these entries because a person had just asked it a question was zero. Being stored is not being used. A crawler that copies your page has not quoted it.
Those are two different events that look identical in a log line count, and most people who check their logs add them together. They should be kept apart, because they answer different questions. One asks whether a machine knows the page exists. The other asks whether the page was useful to someone today.
The three stages
The way I read it now is as three stages. Stored means a training crawler copied the page, which might shape what a future model knows. Indexed means a search crawler can find it, which makes it a candidate for live answers. Quoted means an assistant actually fetched it while answering a person, which is the closest thing to a citation you can see from your own server. OpenAI’s own documentation draws the same line: its live agent visits a page when a user asks a question and does not crawl the web on its own. A page can be stored and indexed every day for months and still never be the answer to anything.
02The same writer, a different address
The comparison I did not plan
The useful part of this data is that I have a control. Another site I run sits on the same server, follows the same writing rules and is written by the same person. In the same fifteen days, ChatGPT’s live agent fetched its pages 181 times, after removing requests that borrowed the agent’s name to probe for configuration files. OpenAI’s search crawler made 153 successful page fetches there. Same author, same rules, same server software. The writing is not the variable.
What actually differs
Three things differ, and none of them is on the page. This blog lives on a country domain that search engines geotarget to one small region. It has almost no links pointing at it. And it is young in search. The other site has a neutral domain, links, and a longer history. A live answer starts with a search, and if the search does not return your page, the answer cannot fetch it, however good it is. That is the point ranking was a location made months ago, now with a control group. An assistant cannot quote a page its search step never returned.
Crawls prove you exist. Live fetches prove you were useful.
03Find out which stage you are stuck at
The diagnosis
Split your log by stage rather than by bot. If you see no training crawlers at all, check your robots rules and your firewall first, as in the guide to who is reading your site. If training crawlers come but search crawlers rarely do, the problem is discovery: sitemaps, internal links, and submitting pages to the engines. If search crawlers come every day and live fetches never do, you are indexed but not chosen, and the fix is upstream, in the things a search ranks on: links, the domain, and whether your page answers a question people actually ask. When the crawlers come and the live fetches do not, stop editing the page and start working on how it is found.
Count the real fetches only
One caution on reading live fetches. Scanners borrow these names to look for exposed files, so count only successful requests for real pages, and verify the address against the ranges each company publishes when it matters. A request for a page that does not exist is not a citation, whatever name it carries. More often it is somebody looking for your passwords. On the other site, 274 requests carried ChatGPT’s live agent name and 181 were real page fetches. If I had counted every line, I would have overstated it by half. The same discipline from the second scoreboard applies: the totals hide the change that matters.
Tomorrow: how to add a paywall without losing the citations a page has already earned.
>What is the difference between being indexed and being cited by AI?+
>How can I tell if ChatGPT is citing my pages?+
>Why is my site crawled by AI bots but never cited?+
Abd Shanti. "Indexed, Not Quoted: The Gap Between Stored and Used" CITED, Entry 071, Oct 3 2026. unknown.ps/blog/indexed-not-quoted/
