Under the hood
Crawl, extract, generate, verify. Here's exactly what happens in the 30 seconds between pasting your URL and downloading a ZIP — including the step most generators skip entirely.
Enter any publicly accessible website URL. Lab451 accepts root domains, subdomains, and specific path prefixes. No authentication, no API key, no installs required on your end.
We send a polite crawl bot (identifiable by its user-agent Lab451Bot/1.0) to walk your site's link graph. It respects your existing robots.txt — even the one it's about to replace.
Follows internal links to discover all public pages
Strips nav/footer/boilerplate, extracts main content
Reads title tags, meta descriptions, Open Graph data
Finds publish dates for accurate sitemap lastmod values
With your site fully mapped, Lab451 compiles each file individually. Each is built to a specific spec:
llms.txt
llmstxt.org spec
A concise, Markdown-formatted document describing your site: its purpose, main sections, key topics, and preferred AI interaction guidelines. Think of it as a README for AI models.
llms-full.txt
Extended spec
A full-content companion that includes the actual text of every page (cleaned and structured). Used by RAG pipelines and AI systems that do deep document retrieval rather than just crawling.
sitemap.xml
Sitemaps.org protocol
A standards-compliant XML sitemap with accurate lastmod dates, appropriate changefreq hints, and canonical URLs. Submitted-ready for Google Search Console.
robots.txt
RFC 9309
Configured to welcome major AI crawlers (GPTBot, ClaudeBot, Googlebot, etc.) while blocking known scrapers. Auto-references your new sitemap.xml location.
Your docs are full of things that point somewhere — install commands, links, deploy URLs. Any of them can quietly stop resolving, and an AI assistant can invent a plausible package name that lands in a README. Once that's sitting in a file agents treat as authoritative, it stops being a typo and starts being a security problem. So we resolve what's in your files before you get them.
Here's what gets checked:
Checked against npm, PyPI, crates.io, RubyGems, NuGet, Packagist and the Go proxy
Resolved, with dangling CNAMEs to unclaimed deploy targets flagged
Content invisible to human readers is stripped and reported, never published
Findings graded critical to info, with the source page for each
Download all four files in a single ZIP, then upload them to your web server's root directory. That's literally it. No ongoing maintenance required on the free plan.
When significant content changes, such as new product features, updated API documentation, or updated FAQs. For active sites, a weekly or bi-weekly update cadence is recommended.
The honest version
Fair question, and you'll find plenty of people online arguing both sides. Here's where things genuinely stand as of mid 2026.
No ambiguity here. Every search engine and essentially every well-behaved crawler reads both.
These two have been load-bearing infrastructure for two decades, and a sitemap with real
lastmod dates is worth having whatever happens with the newer formats.
The clearest real-world use today is developer tooling. Point Cursor, Claude Code, Copilot,
Windsurf, Cline or Aider at a documentation URL and it fetches the file on demand. That's why
developer-docs sites adopted it far faster than everyone else — and why Anthropic publishes one
at code.claude.com/llms.txt. Retrieval crawlers from Microsoft and OpenAI have also
been observed fetching both files.
The consumer search side is a different story. While Chrome's Lighthouse includes basic checks regarding agentic readiness (llms.txt) and Googlebot may technically fetch the file if it encounters the URL on a standard crawl, Google Search treats it as non-actionable; having or omitting the file will neither help nor hurt your visibility.
Because it's cheap, the agent-tooling use is real today, and the format will most likely broaden. That's a hedge, and we think it's a sensible one — we've argued exactly that on our blog.
Additionally, generating verified files (all four) is a security bonus no one should miss in todays unpredictable and malicious web environment.
And verification isn't a hedge at all — an unregistered package name in your docs is a live risk today, whether or not a single agent ever fetches your llms.txt.
Easiest honest test we know: put a URL inside your llms.txt that appears nowhere else
on your site — something like /agent-check-8f2a — and watch your server logs. Nothing
reaches that path by browsing, so any hit came from something that read the file. Give it a few
weeks. You'll know more than most of the internet does.