A robots.txt generator example is the literal plain-text output a standards-aligned tool produces for a chosen website origin and policy: one User-agent: * group, an Allow or Disallow rule, and a Sitemap line formatted under RFC 9309. Reviewing a generator example matters because the file is short, syntax-sensitive, and read by every compliant crawler at the root of your origin, so a single misplaced slash, anchor, or wildcard changes which paths are requested. The output of a standards-aligned robots.txt generator follows a strict shape: one or more groups beginning with User-agent, one rule per line using Allow or Disallow, an optional Sitemap line with an absolute URL, and plain UTF-8 text served at the exact scheme, host, and port it governs. Looking at a real example also lets you verify capitalization and prefix scope before publishing, since path matching begins at the first octet and RFC 9309 specifies that crawlers should compare paths case-sensitively. The result is that /Private and /private can mean different things, and Disallow: /draft does not necessarily express the same intent as Disallow: /draft$. Seeing each example output lets you review these details before any crawler fetches the file.

What a Robots.txt Generator Actually Writes
The text a robots.txt generator writes follows the group-and-rule structure standardized by RFC 9309. Every compliant file starts with one or more groups; each group begins with a User-agent line naming a product token (or the wildcard *), and each subsequent line in the same group is a single rule, almost always an Allow or Disallow directive applied to a path pattern. A focused single-site generator example for one website includes exactly one wildcard group, one rule that reflects the chosen policy, and a final Sitemap line pointing crawlers to the site's sitemap.xml. Because the wildcard token applies when a crawler has no more specific matching group, a generator that intentionally emits one general group covers most production sites without pretending to manage every search engine's private extensions. Lines that begin with a hash sign are treated as comments and are rejected by a strict generator to avoid silently turning part of a user-entered rule into a comment, while wildcards such as * and the trailing $ anchor are preserved because the protocol defines those special matching characters.
The output file is plain UTF-8 text, served with the filename robots.txt (lowercase) and the MIME type text/plain. The file governs only the exact scheme, host, and port where it is served: a file on a subdirectory, another subdomain, or a different protocol does not govern the intended origin. Everything the user types and everything the generator produces stays in the current browser tab, since the file is built locally without sending configuration data to a server.
Three Policy Modes and Their Example Output
A standards-aligned generator offers three well-defined policies and produces an example output for each. The table below shows the exact text written by the Robots.txt Generator when the site origin is https://example.com and the user has not added extra paths or changed the sitemap location.
| Policy mode | Rule line written | Example output |
|---|---|---|
| Allow all crawlers | Allow: / | User-agent: *Allow: /Sitemap: https://example.com/sitemap.xml |
| Block all crawlers | Disallow: / | User-agent: *Disallow: /Sitemap: https://example.com/sitemap.xml |
| Selective disallow | One Disallow line per entered path | User-agent: *Disallow: /admin/Disallow: /private/Disallow: /draftSitemap: https://example.com/sitemap.xml |
Beyond these baseline modes, selective mode adds Disallow lines for each slash-prefixed path the user enters, removes exact duplicates in first-seen order, and caps the list at 50 rules. Any non-empty rule that does not start with / is rejected before the file is generated, so the example output stays parseable by every compliant crawler. A URL containing a page path, query, or fragment is reduced to its origin before the sitemap line is written, which keeps the default sitemap from accidentally inheriting an unrelated route, while explicit ports are preserved.
How to Generate a robots.txt File: An Example Walkthrough
Working through a generator example in the Robots.txt Generator is the fastest way to see how each policy mode translates into actual file text.
- Open the tool and enter the site origin. Type the complete HTTP or HTTPS URL for the website, such as https://example.com, without any path, query, or fragment. The generator normalizes the URL to its origin so the default sitemap line does not accidentally inherit an unrelated route, while explicit ports are preserved exactly as you typed them.
- Choose the overall crawler policy. Pick allow-all for a baseline that lets compliant crawlers request every path, block-all to write Disallow: / when the site must not be crawled yet, or selective-disallow to enter one path per line for the directories and files that should stay out of crawler queues.
- In selective mode, type one disallowed path per line. Each non-empty rule must start with a slash. Use wildcards (*) and trailing anchors ($) where the protocol calls for them, but avoid hash signs so nothing is mistaken for a comment. Exact duplicates are removed automatically and the list is capped at 50 rules.
- Generate the file and inspect the text. The output is plain UTF-8 with one User-agent: * group, the matching Allow or Disallow rules, and a Sitemap: line that points to the normalized origin plus /sitemap.xml. If your real sitemap lives at another location, edit that line before publishing.
- Download the result and compare it with the current production file. Preserve any intentional crawler-specific groups already in production, keep the old file for rollback, and only then publish the new text as /robots.txt at the root of the exact scheme, host, and port it governs.
Reading the Example: Case, Wildcards, and Anchors
Each line in a generator example follows RFC 9309 matching rules, which compare paths against the start of a URL path and use the longest, most specific match when multiple rules apply. Crawlers should compare paths case-sensitively, which means Disallow: /Private and Disallow: /private are not equivalent rules: a crawler honouring both would treat /private-area and /Private-Docs differently even though they share a visible prefix.
The trailing dollar sign anchors a rule to the end of the path, so Disallow: /draft$ matches only the literal /draft page, while Disallow: /draft matches /draft, /drafts, /draft/anything, and every longer path that begins with those characters. Wildcard asterisks add partial matching, so Disallow: /*.pdf blocks PDF files at any directory depth. Before publishing any example output, review capitalization, prefixes, wildcards, and anchors, then test important public and blocked URLs with a search-engine URL tester where one is available. A useful companion reference for retrieval and comparison is the how to get a robots.txt file: retrieve or generate one walkthrough.
What a Robots.txt Example Cannot Do
A robots.txt generator produces a public request, not a security boundary. RFC 9309 explicitly states that these rules are not access authorization: compliant crawlers may honour the requests, but a malicious client can ignore them, and every listed path is publicly visible in the file. Never list a sensitive route as a substitute for login checks, server-side access control, or authentication, because the blocked path is still downloadable by anyone who fetches /robots.txt.
Blocking crawling also does not guarantee that a URL disappears from search results: a search engine can discover the URL through inbound links and retain limited information without ever fetching the blocked page. Use page-level indexing controls such as the robots meta tag or the X-Robots-Tag header, the relevant search-engine removal workflow, or authentication that protects the underlying resource, depending on the actual goal. The generator does not add Crawl-delay because it is not part of RFC 9309 and is not supported consistently across crawlers; do not combine it with the generated file without understanding how each crawler handles the directive. A clean generator example is the starting point, not the finishing line.
Comparing the Generated File With Your Live robots.txt
Before replacing a production robots.txt file, download the generator output and place it next to the current file. The two documents often look almost identical, but production files sometimes include crawler-specific groups that a focused single-group generator intentionally does not produce, along with legacy rules left over from earlier audits. Preserve those intentional groups by adding them back into the new text after the generator output, and keep a copy of the old file in a safe place for rollback in case a crawler interprets the new rules differently than expected.
The generator does not fetch the existing live file, validate server responses, submit the file to a search engine, or confirm that a crawler has refreshed its cache; those steps require access to the deployed site and the search engine's own tools. Google's documentation on creating a robots.txt file describes the production-side checks that no browser tool can replace. Once the comparison is complete and the file is published at the root as plain text with a lowercase filename, the new robots.txt governs only that exact scheme, host, and port, never a subdirectory, another subdomain, or a different protocol.