Extracting URLs from a sitemap means reading the XML file, taking every direct <loc> child inside each <url> or <sitemap> entry, and writing those values out as a clean one-URL-per-line list with duplicates dropped in the order they first appeared. The operation is local: nothing is uploaded, nothing is fetched, and the parser applies the same rules Google describes in the Sitemaps protocol. After a run completes, you see which root type was detected (urlset or sitemapindex), how many unique loc values were accepted, and how many duplicates were removed. The accepted list is copied with a single click for downstream checks such as crawler validation, canonical comparisons, redirect audits, log-file diffing, or pre-migration reviews. Because parsing happens in the browser tab, staging URLs and prelaunch inventories never leave the device unless you copy them out yourself.

extract urls from sitemap explained
Extract URLs From a Sitemap: The Parsing Step Explained

What "Extracting" Means in a Sitemap Context

Extracting URLs from a sitemap is a parsing task, not a crawling one. The reader hands the parser an XML document that follows the Sitemaps protocol, and the parser walks the tree to collect the direct <loc> values inside each container element. Nothing is requested from the live website, no redirects are followed, and no headless browser is launched. The result is a flat list, one URL per line, that you can paste into a spreadsheet, a crawler config, or an audit document.

That distinction is easy to miss. Many SEOs conflate "extracting URLs from a sitemap" with "crawling the site", but the sitemap itself is only a hint file. It announces what the publisher wants crawlers to consider; extraction simply reads that announcement. After extraction, every URL still needs separate checks for HTTP status, canonical tags, robots directives, and actual indexing. The clean list is the starting point of those checks, not the conclusion.

The Two Sitemap Roots the Parser Recognizes

The Sitemaps protocol defines two top-level roots, and a parser that ignores one of them will silently miss a large share of real-world sitemaps. In a urlset, each page is wrapped in a <url> element, and the location sits in its direct <loc> child. In a sitemapindex, each entry is a <sitemap> element, and the direct <loc> child holds the URL of a child sitemap file rather than a page. The two structures are not nested in a single file; an index points at files that contain page lists.

The extractor distinguishes the two by inspecting the root element, then enforces the matching direct-child rule. If the root is a urlset, every <url> must contain its own <loc> directly. If the root is a sitemapindex, every <sitemap> must contain its own <loc> directly. Optional siblings such as <lastmod> are tolerated but ignored. Extension namespaces are skipped: an <image:loc> does not stand in for a missing page <loc> because it describes a different resource role. The tool accepts a consistent namespace prefix such as <sm:urlset>, <sm:url>, and <sm:loc>, as well as the common default-namespace form.

Input ScenarioRoot DetectedWhat the Extractor Returns
Standard page-list sitemap<urlset>One URL per line for every page loc in the file
Directory of sitemap files<sitemapindex>One URL per line for each child sitemap file (not the pages inside)
Namespaced form like <sm:urlset>urlset with prefixSame page-URL output as the default-namespace form
Compressed .xml.gz archiveNot parsedDecompress outside the tool, then paste the resulting XML
Remote URL pasted into the editorNot fetchedPaste XML text instead; the widget makes no network request
Nested markup inside a <loc>RejectedThe run fails with a specific item number pointing at the broken entry

How to Extract URLs From a Sitemap

The whole workflow is local and runs in the browser tab. The Sitemap URL Extractor reads the pasted XML, applies the validation rules described below, and produces a deduplicated list you can copy out. Five steps cover most use cases.

  1. Open the sitemap source in a plain text editor. If you have a .xml.gz file, decompress it first so you can paste the underlying XML.
  2. Copy the complete XML text from a single urlset or sitemapindex document. The widget processes one document per run.
  3. Paste the text into the Sitemap URL Extractor editor and run the extraction.
  4. Review the report panel. It shows which root was detected, how many unique URLs were accepted, how many normalized duplicates were dropped, and the full first-seen-order list.
  5. Copy the one-URL-per-line list to your clipboard and feed it into live status checks, canonical comparisons, robots reviews, indexing audits, or migration diffs. Those checks happen separately.

What the Parser Validates Before Accepting a URL

Before a loc makes it into the output, the parser applies a fixed sequence of checks. The order matters: a URL that passes one stage but fails the next is still rejected, and the entire run fails rather than producing a partial list. Knowing the rules helps you fix a source file quickly instead of guessing. The rules, limits, and output guide goes deeper on each step, but the short version is below.

First, the document must contain a closed root with the right structure. A leading byte-order mark, XML declaration, and comments are stripped. DTD declarations and custom entity declarations cause the run to fail because the widget does not perform entity expansion. Second, every <url> or <sitemap> item must contain a direct <loc>. Nested markup inside the loc, missing loc children, and incomplete items all fail the whole run.

Third, the decoded loc text must be parseable as an absolute URL. The widget uses the browser's WHATWG URL implementation, so host casing, percent-encoding, and default-port handling follow the standard. URLs with credentials, fragments, raw whitespace, malformed percent escapes, backslashes, relative paths, or non-HTTP schemes are rejected. The serialized URL must also be under 2,048 characters, which mirrors the protocol's loc cap.

Fourth, after normalization the widget deduplicates. Two loc values that serialize to the same URL count once; only the first occurrence is kept so the order matches the source. Fifth, the size cap kicks in. The widget accepts at most 50,000 unique URLs and five million UTF-16 input code units. Crossing either bound fails without truncation, so production sitemaps that approach the protocol's uncompressed byte cap should be split into smaller files referenced by an index.

Reading the Counts: Root Type, Duplicates, and Output

The report panel surfaces four pieces of information that together describe what the parser did. The root type tells you whether the document was treated as a page list or a sitemap directory. The unique count is the number of accepted loc values after normalization and deduplication. The duplicate count is the number of loc values that were dropped because they matched an earlier entry under browser-standard serialization. The complete output is the first-seen-order list itself, ready to copy.

The duplicate count is often more interesting than it looks. Default-port variants and host-case variants can collide under browser-standard serialization, which lowercases hosts and removes default ports. A surprisingly high duplicate count is a hint that the source generator is emitting inconsistent loc strings, which is worth flagging in the same audit rather than dismissing as noise.

What Comes After You Have the List

Extraction is the discovery step. The output is evidence that a URL appeared in the sitemap and passed the parser's validation; it is not evidence that the page is crawlable, canonical, indexed, or ranked. Several live-site checks belong in the next phase, and they are out of scope for the widget itself.

Use the list to drive an HTTP status sweep so 4xx and 5xx responses surface in the same report. Compare each entry with the canonical tag served by that page to spot entries the sitemap claims but the page disavows. Cross-reference with robots directives and noindex meta tags to find URLs that the sitemap lists but the page asks search engines to skip. Diff the list against an analytics landing-page report to confirm the URLs that actually receive traffic are present. If a migration is in progress, diff the new sitemap's extraction against the previous release to see exactly which loc values were added, removed, or re-serialized.

Why Browser-Side Parsing Matters for SEO Audits

Many SEO audits include private staging URLs, prelaunch inventories, or client-confidential paths that should not leave the device. Because the widget runs entirely in the browser tab and never uploads the XML, those lists stay local unless the user chooses to copy them out. That property matters when the audit belongs to a regulated environment or when the data is sensitive enough that a server round trip would require sign-off.

The browser-side model also makes the operation reproducible. There is no API key, no rate limit, and no queue to wait in. The same pasted text produces the same output on any machine that loads the widget, which is helpful when an audit needs to be defended later. Combined with the deduplication rules and the strict item-level error reporting, the result is a small, predictable tool that fits inside a larger workflow instead of trying to replace one.

If you're weighing options, Sitemap Generator for Website: Build From a Reviewed List covers this in detail.