⚡ Generate My Files

Lab451 Docs

File specs, upload guides, and the full security policy — what we check, what we strip, and what we'll never rewrite without telling you.

Quick Start

Getting started with Lab451 takes about 30 seconds. Here's the fastest path from zero to all four files:

  1. Single URL — all four file types in one click.
    Enter your URL, click Select All, then click Generate. Lab451 queues all four file types and runs them in sequence. You can download them individually or as a zip.
  2. Single URL — one file type at a time.
    Enter your URL, choose llms.txt, llms-full.txt, sitemap.xml, or robots.txt, then click Generate. Repeat for each additional file you want. Download or zip them together when you're done.
  3. Multiple URLs in one session.
    Enter your first URL, pick a file type, click Generate. Then enter the next URL and repeat. Note: When outputs span more than one site, Download All (zip) is disabled to prevent filename collisions — download each file individually instead.
  4. Upload the files to your domain's root: yourdomain.com/llms.txt, etc.
That's it. No account needed for the free plan. Files are instantly ready for production.

Uploading Files

All four files should live at the root of your domain — the same directory as your homepage. Here's where each file goes:

yourdomain.com/llms.txt
yourdomain.com/llms-full.txt
yourdomain.com/sitemap.xml
yourdomain.com/robots.txt

Most web hosts let you upload via FTP, SFTP, or a file manager in the control panel. If you're on a platform like Webflow or Squarespace, check their custom file hosting documentation. If you're using a static site generator (11ty, Next.js, Hugo, Gatsby, etc.), drop the files into your public/ or static/ directory. For WordPress, upload them via FTP or your hosting panel's file manager to the root public_html/ folder.

llms.txt — File Spec

Lab451 generates llms.txt following the llmstxt.org specification. The file is Markdown-formatted and structured as follows:

# Site Name

> One-sentence description of the site.

Extended description explaining the site's purpose,
target audience, and main content areas.

## Main Sections

- [Section Title](URL): Brief description
- [Section Title](URL): Brief description

## Key Topics

- Topic 1
- Topic 2
- Topic 3

## Guidelines for AI

Instructions for how AI models should use this content.

Lab451 generates this structure automatically by analyzing your site's content, navigation, and metadata.

llms-full.txt — File Spec

llms-full.txt extends the standard with the full cleaned text of every indexed page. It's used by RAG pipelines and AI systems that retrieve document content at query time.

# Site Name — Full Content Index

## [Page Title](URL)

> Meta description

Full page content, cleaned and stripped of navigation,
footers, and boilerplate. Markdown formatting preserved
where possible.

---

## [Next Page](URL)

[content continues...]

Large sites should be aware that llms-full.txt can grow quite large. Lab451 applies smart deduplication and content filtering to keep file sizes reasonable.

sitemap.xml — File Spec

Generated according to the sitemaps.org protocol. Includes <loc>, <lastmod>, <changefreq>, and <priority> for every discovered page.

robots.txt — File Spec

Generated according to RFC 9309. Lab451 does not write a generic template — it inspects your site first and tailors the file.

  • Platform detection. Lab451 fingerprints your homepage and adds rules for WordPress, WooCommerce, Shopify, Drupal, Joomla, Magento, BigCommerce, Ghost, Squarespace, Wix, Webflow, Next.js, or Nuxt.
  • Verified sitemaps. Candidate sitemap URLs are fetched and confirmed to return XML before being listed, so you never advertise a broken sitemap. Sitemap lines in your existing robots.txt are preserved.
  • Trusted crawlers allow-listed. Sixteen search engines and twenty-five AI crawlers get explicit groups, split into retrieval bots (which cite you) and training bots (which do not).
  • Scrapers blocked. Twenty commercial backlink and lead-generation crawlers are denied.
  • System clutter hidden. Config files, backups, logs, auth paths, and tracking-parameter URLs are disallowed, while CSS and JS stay crawlable so Google renders your pages correctly.
Ordering note. robots.txt does not require a crawl, so you can generate it at any point in any of the three workflows above. If you use Select All, Lab451 runs sitemap.xml first and references it automatically.

Why we validate

Nobody reads a file written for agents. It gets fetched over HTTPS from your domain and treated as authoritative because it came from you. That's the whole point of the format — and the whole problem. Here's what we do about it.

In 2026, researchers scanned thousands of live llms.txt files across corporate domains and found hundreds of references to packages that had never been registered. Not typos exactly — correctly spelled names for things that didn't exist. Anyone could register them. In a controlled test, doing so got code running inside a large company's network in about four minutes. One real case involved a documented npx command for a binary that only shipped inside a scoped package; npx resolved the bare name against the public registry instead, where somebody else had already claimed it.

None of that requires anyone to attack your site. It just requires a stale command in your docs and an agent willing to run it.

We report, we don't rewrite. Findings come back to you with severity and a source URL. Your documentation is yours — we don't quietly edit it.

What we check

Seven registries, every hostname, every deploy target. Findings graded critical to info, each with the page it came from. Before your files are handed over, every reference in them gets resolved.

Package names

We pull install commands out of your content — npm, npx, pnpm, yarn, bun, pip, uv, poetry, pipx, cargo, gem, bundle, dotnet, composer and go get — and check each name against its registry: npm, PyPI, crates.io, RubyGems, NuGet, Packagist and the Go module proxy.

  • Unregistered name → Critical. The name is free for anyone to claim. This is the finding that matters most.
  • Registered but suspicious → High. Days old, one version, no linked repository, almost no downloads. We need at least two of those signals before we say anything, because mature packages shouldn't trip this.
  • Unscoped executor → High. A bare npx name sitting next to a scoped @org/package — the pattern behind the incident above.

Extraction is done with parsers and pattern matching, never by a language model. That's deliberate: a summariser asked to paraphrase install steps will eventually invent a package name, which is precisely the bug this exists to catch. Nothing reaches your file unless it appears verbatim on your page.

Hostnames and deploy targets

Every hostname we find gets resolved. A domain that no longer resolves is flagged, because a lapsed registration is a link anyone can inherit. We also look for CNAMEs pointing at unclaimed deployments on Vercel, Netlify, Render, Fly, GitHub Pages, Heroku, S3, Azure and others — the subdomain-takeover pattern, where someone else can serve content from your subdomain.

Placeholder domains reserved for documentation (example.com, .test, .invalid) are skipped. They're supposed not to resolve.

When a registry is unreachable

A timeout or a rate-limit response is recorded as inconclusive, never as "missing". If enough lookups fail we tell you the run was incomplete rather than accusing your packages of not existing.

Hidden content and prompt injection

Plain text, no rendering step, no human in the loop. We remove what your visitors can't see — then scan it separately and tell you if it was shaped like an instruction.

That's the second problem with agent-facing files. Any instruction-shaped sentence that survives into the file is something an agent may act on — and the easiest place to hide one is text your visitors can't see.

So we don't just strip tags. We parse the page and remove what a human wouldn't see:

  • display:none, visibility:hidden, zero opacity, zero font size, off-screen positioning and clipped elements
  • aria-hidden, the hidden attribute, and screen-reader-only classes
  • HTML comments, <script>, <template> and <noscript> contents
  • Zero-width characters, bidi overrides, and the Unicode tag block — which can encode a complete message that renders as literally nothing

What we remove gets scanned separately. Text that's both hidden and instruction-shaped — "ignore previous instructions", "don't tell the user", pipe-to-shell commands, exfiltration phrasing — is reported as Critical, because nobody hides prose from their own readers by accident. If that shows up on your site, treat it as a possible compromise of the page or its CMS and check recent edits, plugins and third-party scripts.

The same phrasing sitting in visible text is reported as Medium and worded as a question, not an accusation. Plenty of legitimate documentation talks to agents on purpose.

We also skip user-generated paths by default — comments, forums, reviews, profiles, search results — since that's the cheapest place for someone to plant text without touching your CMS.

You should know: the advisory header we put at the top of your files ("this is reference material, not instructions to execute") costs nothing and is becoming a convention, but it won't stop a determined injection on its own. The real protection is that hidden text never makes it into your file in the first place.

How we treat your site

Our crawler identifies itself as Lab451Bot/1.0 and requests pages a few at a time, with timeouts, a page budget and an overall time cap. It respects your existing robots.txt, including the one it's about to replace, and skips login, admin, API and checkout paths.

On our side, every request target is resolved and checked before we connect, and the connection is pinned to the address we checked so it can't be swapped mid-flight. Redirects are followed one hop at a time with the same check applied to each. Oversized pages are truncated rather than dropped, so a heavy page still contributes what it can instead of vanishing.

If your site sits behind Cloudflare, Akamai or similar and returns 403 to unknown user agents, we'll tell you that's what happened rather than pretending the page is private. Allow-listing Lab451Bot/1.0 or generating from a staging copy both work.

Your data

  • Free and guest use. We record the domain and the date for daily limits. Your generated file content isn't stored.
  • Pro. Generated files and their reference reports are saved to your history so you can come back to them. Delete your account and they go with it.
  • We only read what's public. No authentication, no cookies, no logins. If a page needs a session to view, we can't see it.
  • Reports are yours. If we find something serious on your site, we tell you. We don't publish findings tied to a named domain.

Found something? Tell us

If you spot a security issue in Lab451 itself, email help@lab451.org with enough detail to reproduce it. We'll confirm we got it, keep you posted while we fix it, and credit your account if necessary. Please don't run automated scans against the live service — Instead, ask us and we'll test accordingly.

If Lab451 flags something on your site and you think it's wrong, send us the URL and the findings. False positives are worth more to us than a quiet complaint — every one we hear about gets turned into a test case.