A sitemap generator for a website is a tool that converts a reviewed list of page URLs into a standards-based sitemap.xml file. The XML Sitemap Generator is the list-first option that runs entirely in your browser tab, so no URLs and no generated XML are uploaded. When you already know which pages belong on your site map, you do not need a crawler — you need a tool that validates each URL, applies the protocol's required XML structure, deduplicates entries in the order you supplied them, and hands back a downloadable file. The list-to-XML approach suits small static sites, freshly launched blogs, hand-curated landing-page sets, and any situation where the maintainer wants explicit control over what gets submitted to search engines. The output is only as accurate as the input, which is exactly why the review-first workflow fits cleanly alongside a careful pre-publish checklist for redirects, duplicate content, blocked pages, and private pages that should never reach a search engine.

sitemap generator for website
Sitemap Generator for Website: Build From a Reviewed List

When a List-First Generator Fits Your Website

Most websites start small and stay small enough that the maintainer already knows every URL worth indexing. A list-first sitemap generator is built for that case. Instead of asking software to crawl the site, fetch responses, follow links, and infer which pages deserve a slot, the maintainer supplies the URL list directly. The tool then validates, normalizes, deduplicates, and serializes the result.

Three situations fit the list-first workflow especially well. A static site with a few dozen pages where the maintainer can list the public URLs by hand. A recently relaunched blog or portfolio where the previous crawl produced stale or broken URLs that should be excluded rather than discovered. And any site where indexability is decided by editorial judgment — thin pages, tag archives, and internal search results that the maintainer wants to keep out of the sitemap on purpose.

What the maintainer controls

  • Which URLs appear in the output, in the order typed
  • Whether to add shared optional metadata (lastmod, changefreq, priority)
  • Which URLs to leave out — the tool never invents missing pages

A crawler-style generator could surface the same URLs in the end, but it would also drag in redirects, soft-404 pages, parameterized variants, and any URL that slipped through your noindex rules. The list-first approach keeps that judgment with the person responsible for the site.

What the Tool Handles — and What You Handle

Handled by the toolStays with the maintainer
Validates each nonblank line as an absolute HTTP or HTTPS URLChoosing which URLs belong in the sitemap
Normalizes URLs through the browser URL parser (lowercase host, default-port removal, root slash, percent-encoding)Reviewing redirects, status codes, and canonical tags
Rejects embedded credentials and URL fragmentsDeciding which pages are indexable
Deduplicates entries in first-seen order, preserving your chosen orderRemoving private pages, internal search results, and admin URLs
Applies shared lastmod, changefreq, or priority when suppliedTruthfully assigning one shared metadata value across every entry
Entity-escapes XML data and emits the official UTF-8 urlset namespaceHosting sitemap.xml on the represented site
Prepares a local download without uploading anythingSubmitting the file to search engines and listing it in robots.txt

Reading that table keeps the workflow honest. The generator's job is mechanical and protocol-bound; the editorial decisions remain yours. If a URL does not belong in the sitemap, it does not belong in the input list either.

Prepare Your Reviewed URL List

Every nonblank line in the input must be an absolute HTTP or HTTPS URL for one host. Bare domains like example.com, relative paths like /about, FTP URLs, and any line with leading or trailing whitespace will fail validation rather than being silently trimmed. Whitespace-only lines are ignored, but everything else is held to the standard.

Before you paste anything, build the list the way you would want a search engine to see it:

  • Include only URLs that resolve, return 200, and are intended to be indexed
  • Strip tracking parameters and session IDs that should not become separate sitemap entries
  • Remove URLs blocked by robots.txt or marked noindex
  • Exclude private pages, admin endpoints, internal search results, and tag or archive pages you do not want surfaced
  • Resolve canonicals — the URL you want indexed is the URL that goes in the sitemap

Keep all URLs on the same serialized host (host plus non-default port). Mixing example.com with example.com:8443 fails the single-host check, even though both schemes are technically valid for that host. HTTP and HTTPS entries for the same host are accepted together because the protocol defines one host, not one origin. URL normalization follows the WHATWG URL Standard, which lowercases hosts, removes default ports, and percent-encodes non-ASCII components before any comparison.

Generate the sitemap.xml File

  1. Open the XML Sitemap Generator in your browser tab.
  2. Paste one absolute URL per line into the URL list field. LF, CRLF, and CR line breaks are all accepted.
  3. Decide whether to add the optional shared fields. Leave lastmod, changefreq, and priority blank if you only want loc-only entries.
  4. Click Generate. The tool parses each line, serializes every URL through the browser URL parser, deduplicates in first-seen order, validates any optional metadata, and measures the final XML size.
  5. Review the completed XML in the preview area. The output starts with a UTF-8 declaration and a urlset element using the official http://www.sitemaps.org/schemas/sitemap/0.9 namespace.
  6. Confirm the count and byte totals. A normalized loc is at most 2,047 characters; the whole file is capped at 10,000 unique URLs and 10,485,760 UTF-8 bytes.
  7. Click Download to save sitemap.xml. The download is a temporary Object URL owned by the current result, so editing any input immediately revokes the link and clears the previous XML.

Any single invalid line — a stray space, a fragment, embedded credentials, a non-HTTP scheme — fails the whole generation rather than skipping the bad entry. That all-or-nothing behavior is deliberate: a sitemap that quietly drops malformed lines is harder to debug than one that refuses to produce a file until every line is correct.

Optional Shared Metadata Fields

The protocol allows three optional child elements per URL: lastmod, changefreq, and priority. The list-first generator applies the same value to every entry when you supply one, so the field is only honest if it is true for every URL on the list.

FieldAccepted valuesWhat it means
lastmodA real Gregorian YYYY-MM-DD date, or a complete date-time with seconds and Z or a valid UTC offsetThe page's last significant modification
changefreqalways, hourly, daily, weekly, monthly, yearly, neverA hint about how often the page typically changes
priorityA decimal from 0.0 through 1.0A hint about relative importance within your site

All three are hints, not crawl commands and not ranking guarantees. Search engines may ignore them entirely. lastmod is the one with the most editorial weight, because a wrong value across a whole sitemap can mislead crawlers into revisiting pages that have not actually changed. Use it only when one date is truthful for every URL on the list.

Hard Limits and Validation Behavior

The Sitemap protocol permits up to 50,000 URLs and 52,428,800 uncompressed bytes per file. The browser-based generator uses lower whole-file budgets to bound preview, Blob, and memory: 10,000 unique URLs, 5,000,000 UTF-16 input code units, and 10,485,760 UTF-8 XML bytes. The exact 10,485,760-byte boundary is accepted; one additional byte fails the generation.

Every limit is all-or-nothing. The 10,000th unique URL is accepted and the next unique URL fails. Duplicate lines do not consume output slots, so cleaning up repeats before pasting can move you under the cap without dropping real entries. XML size is measured after URL serialization, XML escaping, optional tags, indentation, and UTF-8 encoding — not on the raw text you paste.

Validation is strict on purpose. The generator does not create shortened XML, ellipsis, capped preview, or partial downloads. If you need more room than one file allows, split the list across multiple sitemaps and link them through a sitemap index file generated separately. For protocol-level detail on schema, limits, and extension modules, see the sitemap protocol rules for list-based builds. The protocol's overall format and submission rules are documented in Google's build and submit a sitemap guide.

Review and Publish Your Sitemap

Before the file goes live, walk through a pre-publish checklist that the generator cannot do for you. Confirm that every URL in the list resolves, returns a 200 status, and is the canonical version of the page. Remove any URL that redirects, returns 404, or is marked noindex. Drop internal search results, admin endpoints, and any private page that should never have been on the list.

Upload sitemap.xml to the root of the represented site, then reference it from robots.txt with a Sitemap: line so crawlers can discover it without you submitting it manually. For very large sites, the protocol allows sitemap indexes — separate XML files that point to multiple child sitemaps — but the list-first generator produces a single sitemap.xml at a time, so an index workflow is a separate step.

Submitting the file to a search engine is a separate action. The generator does not submit the sitemap, ping endpoints, or modify robots.txt. That separation is part of why the file does not need to be uploaded anywhere during generation — your hosting pipeline, your CMS, or your static-site deploy step is responsible for putting the file on the represented host and announcing it.

Related reading: Extract URLs From a Sitemap for Beginners: First Guide.