Quick Start
Getting started with Lab451 takes about 30 seconds. Here's the fastest path from zero to all four files:
-
Single URL — all four file types in one click.
Enter your URL, click Select All, then click Generate. Lab451 queues all four file types and runs them in sequence. You can download them individually or as a zip. -
Single URL — one file type at a time.
Enter your URL, choosellms.txt,llms-full.txt,sitemap.xml, orrobots.txt, then click Generate. Repeat for each additional file you want. Download or zip them together when you're done. -
Multiple URLs in one session.
Enter your first URL, pick a file type, click Generate. Then enter the next URL and repeat. Note: When outputs span more than one site, Download All (zip) is disabled to prevent filename collisions — download each file individually instead. - Upload the files to your domain's root:
yourdomain.com/llms.txt, etc.
Uploading Files
All four files should live at the root of your domain — the same directory as your homepage. Here's where each file goes:
yourdomain.com/llms.txt yourdomain.com/llms-full.txt yourdomain.com/sitemap.xml yourdomain.com/robots.txt
Most web hosts let you upload via FTP, SFTP, or a file manager in the control panel. If you're on a platform like Webflow or Squarespace, check their custom file hosting documentation. If you're using a static site generator (11ty, Next.js, Hugo, Gatsby, etc.), drop the files into your public/ or static/ directory. For WordPress, upload them via FTP or your hosting panel's file manager to the root public_html/ folder.
llms.txt — File Spec
Lab451 generates llms.txt following the llmstxt.org specification. The file is Markdown-formatted and structured as follows:
# Site Name > One-sentence description of the site. Extended description explaining the site's purpose, target audience, and main content areas. ## Main Sections - [Section Title](URL): Brief description - [Section Title](URL): Brief description ## Key Topics - Topic 1 - Topic 2 - Topic 3 ## Guidelines for AI Instructions for how AI models should use this content.
Lab451 generates this structure automatically by analyzing your site's content, navigation, and metadata.
llms-full.txt — File Spec
llms-full.txt extends the standard with the full cleaned text of every indexed page. It's used by RAG pipelines and AI systems that retrieve document content at query time.
# Site Name — Full Content Index ## [Page Title](URL) > Meta description Full page content, cleaned and stripped of navigation, footers, and boilerplate. Markdown formatting preserved where possible. --- ## [Next Page](URL) [content continues...]
Large sites should be aware that llms-full.txt can grow quite large. Lab451 applies smart deduplication and content filtering to keep file sizes reasonable.
sitemap.xml — File Spec
Generated according to the sitemaps.org protocol. Includes <loc>, <lastmod>, <changefreq>, and <priority> for every discovered page.
robots.txt — File Spec
Generated according to RFC 9309. Lab451 does not write a generic template — it inspects your site first and tailors the file.
- Platform detection. Lab451 fingerprints your homepage and adds rules for WordPress, WooCommerce, Shopify, Drupal, Joomla, Magento, BigCommerce, Ghost, Squarespace, Wix, Webflow, Next.js, or Nuxt.
- Verified sitemaps. Candidate sitemap URLs are fetched and confirmed to return XML before being listed, so you never advertise a broken sitemap. Sitemap lines in your existing robots.txt are preserved.
- Trusted crawlers allow-listed. Sixteen search engines and twenty-five AI crawlers get explicit groups, split into retrieval bots (which cite you) and training bots (which do not).
- Scrapers blocked. Twenty commercial backlink and lead-generation crawlers are denied.
- System clutter hidden. Config files, backups, logs, auth paths, and tracking-parameter URLs are disallowed, while CSS and JS stay crawlable so Google renders your pages correctly.
Why we validate
Nobody reads a file written for agents. It gets fetched over HTTPS from your domain and treated as authoritative because it came from you. That's the whole point of the format — and the whole problem. Here's what we do about it.
In 2026, researchers scanned thousands of live llms.txt files across corporate
domains and found hundreds of references to packages that had never been registered. Not typos
exactly — correctly spelled names for things that didn't exist. Anyone could register them. In a
controlled test, doing so got code running inside a large company's network in about four
minutes. One real case involved a documented npx command for a binary that only
shipped inside a scoped package; npx resolved the bare name against the public
registry instead, where somebody else had already claimed it.
None of that requires anyone to attack your site. It just requires a stale command in your docs and an agent willing to run it.
What we check
Seven registries, every hostname, every deploy target. Findings graded critical to info, each with the page it came from. Before your files are handed over, every reference in them gets resolved.
Package names
We pull install commands out of your content — npm, npx,
pnpm, yarn, bun, pip, uv,
poetry, pipx, cargo, gem,
bundle, dotnet, composer and go get — and
check each name against its registry: npm, PyPI, crates.io, RubyGems, NuGet, Packagist and the
Go module proxy.
- Unregistered name → Critical. The name is free for anyone to claim. This is the finding that matters most.
- Registered but suspicious → High. Days old, one version, no linked repository, almost no downloads. We need at least two of those signals before we say anything, because mature packages shouldn't trip this.
- Unscoped executor → High. A bare
npx namesitting next to a scoped@org/package— the pattern behind the incident above.
Extraction is done with parsers and pattern matching, never by a language model. That's deliberate: a summariser asked to paraphrase install steps will eventually invent a package name, which is precisely the bug this exists to catch. Nothing reaches your file unless it appears verbatim on your page.
Hostnames and deploy targets
Every hostname we find gets resolved. A domain that no longer resolves is flagged, because a lapsed registration is a link anyone can inherit. We also look for CNAMEs pointing at unclaimed deployments on Vercel, Netlify, Render, Fly, GitHub Pages, Heroku, S3, Azure and others — the subdomain-takeover pattern, where someone else can serve content from your subdomain.
Placeholder domains reserved for documentation (example.com, .test,
.invalid) are skipped. They're supposed not to resolve.
When a registry is unreachable
A timeout or a rate-limit response is recorded as inconclusive, never as "missing". If enough lookups fail we tell you the run was incomplete rather than accusing your packages of not existing.
Hidden content and prompt injection
Plain text, no rendering step, no human in the loop. We remove what your visitors can't see — then scan it separately and tell you if it was shaped like an instruction.
That's the second problem with agent-facing files. Any instruction-shaped sentence that survives into the file is something an agent may act on — and the easiest place to hide one is text your visitors can't see.
So we don't just strip tags. We parse the page and remove what a human wouldn't see:
display:none,visibility:hidden, zero opacity, zero font size, off-screen positioning and clipped elementsaria-hidden, thehiddenattribute, and screen-reader-only classes- HTML comments,
<script>,<template>and<noscript>contents - Zero-width characters, bidi overrides, and the Unicode tag block — which can encode a complete message that renders as literally nothing
What we remove gets scanned separately. Text that's both hidden and instruction-shaped — "ignore previous instructions", "don't tell the user", pipe-to-shell commands, exfiltration phrasing — is reported as Critical, because nobody hides prose from their own readers by accident. If that shows up on your site, treat it as a possible compromise of the page or its CMS and check recent edits, plugins and third-party scripts.
The same phrasing sitting in visible text is reported as Medium and worded as a question, not an accusation. Plenty of legitimate documentation talks to agents on purpose.
We also skip user-generated paths by default — comments, forums, reviews, profiles, search results — since that's the cheapest place for someone to plant text without touching your CMS.
How we treat your site
Our crawler identifies itself as Lab451Bot/1.0 and requests pages a few at a time,
with timeouts, a page budget and an overall time cap. It respects your existing
robots.txt, including the one it's about to replace, and skips login, admin, API and
checkout paths.
On our side, every request target is resolved and checked before we connect, and the connection is pinned to the address we checked so it can't be swapped mid-flight. Redirects are followed one hop at a time with the same check applied to each. Oversized pages are truncated rather than dropped, so a heavy page still contributes what it can instead of vanishing.
If your site sits behind Cloudflare, Akamai or similar and returns 403 to unknown user agents,
we'll tell you that's what happened rather than pretending the page is private. Allow-listing
Lab451Bot/1.0 or generating from a staging copy both work.
Your data
- Free and guest use. We record the domain and the date for daily limits. Your generated file content isn't stored.
- Pro. Generated files and their reference reports are saved to your history so you can come back to them. Delete your account and they go with it.
- We only read what's public. No authentication, no cookies, no logins. If a page needs a session to view, we can't see it.
- Reports are yours. If we find something serious on your site, we tell you. We don't publish findings tied to a named domain.
Found something? Tell us
If you spot a security issue in Lab451 itself, email help@lab451.org with enough detail to reproduce it. We'll confirm we got it, keep you posted while we fix it, and credit your account if necessary. Please don't run automated scans against the live service — Instead, ask us and we'll test accordingly.
If Lab451 flags something on your site and you think it's wrong, send us the URL and the findings. False positives are worth more to us than a quiet complaint — every one we hear about gets turned into a test case.