Your llms.txt Is Already Out of Date
01The one file nobody updates
In August I audited how my sites publish, file by file, to find out which parts could fall behind without anyone noticing. Most of them could not. The sitemap is generated from the files on disk every time something changes, and the search engines are pinged automatically after every deploy. There is nothing to forget. Then I got to the llms.txt file, the plain text map of a site written for language models, and it was the only important file that a person had to remember to edit. Everything you generate stays current. Everything you edit by hand starts going stale the day you publish.
What the check found
On one site the file was eight pages behind the sitemap. One of those pages was a real article that had been live for two weeks and had never been added. The others were section pages that had been built since the file was last touched. On the second site the file was missing the index page of its topic hubs, the single page that links to everything else. Nobody had made a mistake on purpose. Hand edited files fall behind quietly, and nobody notices until something checks.
import re, sys, urllib.request as u
site = sys.argv[1].rstrip("/")
get = lambda p: u.urlopen(site + p).read().decode("utf-8", "ignore")
urls = set(re.findall(r"<loc>(.*?)</loc>", get("/sitemap.xml")))
llms = get("/llms.txt")
missing = sorted(x for x in urls if x not in llms)
for x in missing: print("missing", x)
print(len(missing), "pages in the sitemap but not in llms.txt")
sys.exit(len(missing))02An out of date map is worse than none
Who actually reads it
I want to be honest about the size of the stakes. In a fourteen day window, one site’s llms.txt was fetched nine times and the fuller version twice, and not one of those requests came from an AI company. The file is mostly read by directories, audit tools and curious people. That is why I argued in the first entry about this file that it earns almost nothing when it is right and costs you when it is wrong.
The cost of a stale map
A stale map does a specific kind of damage. It tells whoever reads it that your newest work does not exist, and it describes the site as it was, not as it is. If a reader or a tool trusts the file, your best recent page is invisible to them. If a model ever does lean on it, you have handed it an old version of yourself, which is the problem described in the cost of being wrong in public. A file written for machines should never be the least maintained file on your site.
If a person has to remember it, a script should check it.
03Make the deploy refuse to forget
The rule the guard enforces
The fix was not to try harder. It was a small script that runs after every publish, reads the sitemap, reads llms.txt and the fuller version, and prints every page that is in the first and missing from the others. Its exit code is the number of gaps, so a deploy can refuse to finish while the number is above zero. Since August it has run after every publish on both sites, and it has caught the same kind of omission more than once before anyone saw it. The minimal version above does the core comparison in nine lines. An exit code that counts the gaps turns a habit into a rule nobody can skip.
The better long term answer, if your site is built from data, is to generate the file from the same source as the sitemap, so the two can never disagree. I keep mine curated by design, because each line carries a description chosen for that page, and a generated description is usually the page title said twice. That choice is exactly why the check exists. A curated file needs a guard that a generated one does not.
Two refinements
Two refinements make it usable rather than noisy. First, skip pages that do not belong in a curated map on purpose, such as privacy, terms and accessibility pages, or the check will nag about them forever. Second, check the descriptions as well as the links when you can, because a page listed under last month’s description is only half fixed. The goal is not a file that lists everything. It is a file that never leaves out something that matters. And if your live fetches look like the ones in indexed, not quoted, this is still worth doing, because the day a model starts reading the file, you want it to read the current version.
Tomorrow: how to read your Google and Bing search data directly through their programming interfaces.
>How do I keep llms.txt up to date?+
>Do AI assistants read llms.txt?+
>What should llms.txt include?+
Abd Shanti. "Your llms.txt Is Already Out of Date" CITED, Entry 073, Oct 5 2026. unknown.ps/blog/llms-txt-out-of-date/
