⚡ Generate My Files

From URL to four verified files

Crawl, extract, generate, verify. Here's exactly what happens in the 30 seconds between pasting your URL and downloading a ZIP — including the step most generators skip entirely.

01

You paste your URL

Enter any publicly accessible website URL. Lab451 accepts root domains, subdomains, and specific path prefixes. No authentication, no API key, no installs required on your end.

02

Our crawler maps your site structure

We send a polite crawl bot (identifiable by its user-agent Lab451Bot/1.0) to walk your site's link graph. It respects your existing robots.txt — even the one it's about to replace.

add_link
Link graph walking

Follows internal links to discover all public pages

folder_code
Content extraction

Strips nav/footer/boilerplate, extracts main content

page_info
Metadata parsing

Reads title tags, meta descriptions, Open Graph data

edit_calendar
Date detection

Finds publish dates for accurate sitemap lastmod values

03

We generate your four files

With your site fully mapped, Lab451 compiles each file individually. Each is built to a specific spec:

llms.txt llmstxt.org spec

A concise, Markdown-formatted document describing your site: its purpose, main sections, key topics, and preferred AI interaction guidelines. Think of it as a README for AI models.

llms-full.txt Extended spec

A full-content companion that includes the actual text of every page (cleaned and structured). Used by RAG pipelines and AI systems that do deep document retrieval rather than just crawling.

sitemap.xml Sitemaps.org protocol

A standards-compliant XML sitemap with accurate lastmod dates, appropriate changefreq hints, and canonical URLs. Submitted-ready for Google Search Console.

robots.txt RFC 9309

Configured to welcome major AI crawlers (GPTBot, ClaudeBot, Googlebot, etc.) while blocking known scrapers. Auto-references your new sitemap.xml location.

04

This is the part most generators skip

Your docs are full of things that point somewhere — install commands, links, deploy URLs. Any of them can quietly stop resolving, and an AI assistant can invent a plausible package name that lands in a README. Once that's sitting in a file agents treat as authoritative, it stops being a typo and starts being a security problem. So we resolve what's in your files before you get them.

Here's what gets checked:

inventory_2
Package names

Checked against npm, PyPI, crates.io, RubyGems, NuGet, Packagist and the Go proxy

travel_explore
Hostnames

Resolved, with dangling CNAMEs to unclaimed deploy targets flagged

visibility_off
Hidden text

Content invisible to human readers is stripped and reported, never published

rule
Severity report

Findings graded critical to info, with the source page for each

We report, we don't rewrite. Your docs are yours. If we find a package name nobody has registered, we tell you where it is and why it matters — we don't silently edit your content.
05

You download, upload, and you're done

Download all four files in a single ZIP, then upload them to your web server's root directory. That's literally it. No ongoing maintenance required on the free plan.

06

How often should I re-generate my files?

When significant content changes, such as new product features, updated API documentation, or updated FAQs. For active sites, a weekly or bi-weekly update cadence is recommended.

Tip: The point of these files is to stop agents describing your product the way it worked eight months ago. That's the win you can actually measure, and it only holds if the files keep up with the site. Re-running also re-checks every reference, which matters more than it sounds — packages get deprecated and domains expire on their own schedule, so a file that verified clean in March isn't necessarily clean in September.

Will anything actually read this?

Fair question, and you'll find plenty of people online arguing both sides. Here's where things genuinely stand as of mid 2026.

sitemap.xml and robots.txt: yes, definitively

No ambiguity here. Every search engine and essentially every well-behaved crawler reads both. These two have been load-bearing infrastructure for two decades, and a sitemap with real lastmod dates is worth having whatever happens with the newer formats.

llms.txt: Coding agents, yes. Search models: Not everyone is onboard.

The clearest real-world use today is developer tooling. Point Cursor, Claude Code, Copilot, Windsurf, Cline or Aider at a documentation URL and it fetches the file on demand. That's why developer-docs sites adopted it far faster than everyone else — and why Anthropic publishes one at code.claude.com/llms.txt. Retrieval crawlers from Microsoft and OpenAI have also been observed fetching both files.

The consumer search side is a different story. While Chrome's Lighthouse includes basic checks regarding agentic readiness (llms.txt) and Googlebot may technically fetch the file if it encounters the URL on a standard crawl, Google Search treats it as non-actionable; having or omitting the file will neither help nor hurt your visibility.

So why publish one?

Because it's cheap, the agent-tooling use is real today, and the format will most likely broaden. That's a hedge, and we think it's a sensible one — we've argued exactly that on our blog.

Additionally, generating verified files (all four) is a security bonus no one should miss in todays unpredictable and malicious web environment.

And verification isn't a hedge at all — an unregistered package name in your docs is a live risk today, whether or not a single agent ever fetches your llms.txt.

Want to measure it yourself?

Easiest honest test we know: put a URL inside your llms.txt that appears nowhere else on your site — something like /agent-check-8f2a — and watch your server logs. Nothing reaches that path by browsing, so any hit came from something that read the file. Give it a few weeks. You'll know more than most of the internet does.

We cannot guarantee: that we will get you into a model's training data or fast-track you into anyone's index. Nobody can. What we can do is make sure that when something does read your site, it finds a clean, current, verified description of it.