An XML sitemap is just a text file that lists every page URL a website wants search engines to know about, and extracting URLs from a sitemap means turning that nested XML into a flat, one-URL-per-line list you can read, sort, paste into a spreadsheet, or hand to another tool. For beginners, the challenge is not the URLs themselves — every page already lives online — but the wrapper around them. Sitemaps use angle-bracket tags, optional namespaces, and two different root structures, so opening a sitemap.xml file in a browser often looks like noise. A focused Sitemap URL Extractor solves this by reading the pasted XML locally, picking out only the direct loc values, normalizing the URLs to a single canonical form, removing duplicates, and presenting the result as a clean text list. Nothing is uploaded, no API key is needed, and the parser follows the public Sitemaps protocol rather than guessing. What you paste in determines what you get out: a urlset produces page URLs, a sitemap index produces sitemap file URLs, and anything malformed fails the whole run instead of returning a partial list.

extract urls from sitemap for beginners
Extract URLs From a Sitemap for Beginners: First Guide

What a sitemap actually contains

If you have never opened a sitemap before, the first peek can feel overwhelming. The file is plain text, but it uses angle-bracket tags similar to HTML, and it follows a strict structure defined by the public Google Search Central sitemap guide. Two root shapes are allowed by the protocol:

  • urlset — the most common form. Each page lives inside a <url> element, and the actual page address sits inside a single <loc> child. Optional fields such as <lastmod>, <changefreq>, and <priority> may appear, but they do not change the address.
  • sitemapindex — a directory of other sitemap files rather than a directory of pages. Each child entry sits inside a <sitemap> element and contains one <loc> child pointing to another sitemap file.

Some sites add a namespace prefix such as sm: in front of every tag. That is allowed by the protocol and behaves identically to the default form once parsed. XML declarations like <?xml version="1.0"?> and HTML-style comments are also normal and should be ignored by any compliant parser.

Why pull the URLs out into a plain list

The XML format is built for search engines, not for humans. Trying to copy page addresses one at a time from a sitemap file is slow and error-prone, and pasting raw XML into a spreadsheet produces broken rows. A clean, deduplicated list is far more useful for the tasks beginners usually have in mind:

  • Importing the pages into a spreadsheet for review or sorting
  • Building a redirect map during a site migration
  • Comparing what the sitemap claims against what a crawler actually finds
  • Handing the list to a status checker, canonical checker, or batch tool
  • Spotting accidental duplicates caused by host-case or default-port variants

The Sitemap URL Extractor gives you that list directly. Because parsing happens in your browser, the XML never leaves your device, so it is safe to use on staging environments, prelaunch inventories, or any URL list you would rather not share with a third party.

Get the sitemap text ready to paste

Before you can extract anything, you need the actual XML as plain text. Common starting points for beginners include viewing /sitemap.xml directly in the browser, following the Sitemap: line in robots.txt, or asking a colleague to export the sitemap from the CMS. If your file ends in .xml.gz, decompress it first with any standard gzip tool — the extractor does not accept compressed input. If you need help locating the file, the guide on finding sitemap URLs in your browser walks through the usual spots.

Two pitfalls deserve attention before you paste:

  • Pasting a URL into the editor does nothing. The widget makes no network request. You must paste the XML text itself, not the address of the file.
  • Pasting partial XML fails the run on purpose. A truncated file, missing closing tag, or stray edit will not produce a partial list. Fix the source and paste the complete document.

Extract the URLs from your sitemap

  1. Open the Sitemap URL Extractor in your browser tab. No account, no install, no upload.
  2. Paste the complete XML text from one urlset or sitemapindex document into the editor. Make sure the entire file is present, from the opening root tag to its closing tag.
  3. Run the extraction and read the result panel. The tool will report which root type it detected, how many unique URLs were accepted, and how many normalized duplicates were removed.
  4. Copy the one-URL-per-line list using the copy button. The first-seen order is preserved, so the output stays easy to compare line-by-line with the source.
  5. Use the list in your next workflow — spreadsheet import, crawler comparison, redirect audit, or migration review. For status, canonical, robots, and indexing checks, run a separate live tool against each page.

Reading the result: root, unique, and duplicate counts

Beginners often see the result panel and wonder which numbers matter. The three figures are tied to the rules described by the Sitemaps protocol and surfaced plainly so you can sanity-check the file you pasted:

Field in the result What it tells you
Root type Whether the parser recognized a urlset or a sitemapindex. A mismatch with what you expected usually means the wrong file was pasted.
Unique count The number of distinct, normalized URLs accepted after duplicate removal. Each entry is an absolute http or https URL under the 2,048-character protocol limit.
Duplicate count The number of entries that serialized to the same value and were collapsed. Duplicates often reveal host-case or default-port variants in the source file.

The parser accepts up to 50,000 unique URLs and five million UTF-16 input code units per run. Hitting either bound causes the entire operation to fail rather than silently truncate, so plan ahead if you are working with very large sites — split a near-limit file into smaller pieces referenced by a sitemap index and run each one separately.

What the extraction does not tell you

An extracted URL only proves that a valid loc appeared in the pasted XML. It does not mean the page is live, crawlable, canonical, indexed, or ranked. The Sitemaps protocol explicitly describes sitemaps as discovery hints, not indexing guarantees, so beginners should treat the list as evidence for a wider audit rather than proof of any individual page's status.

Common checks the extractor does not perform include:

  • HTTP status codes, redirects, or response time
  • Canonical tag selection on the live page
  • Robots directives that might block crawling
  • Whether search engines actually accepted or processed the source file
  • Whether the URLs are useful, unique, or high quality

For those checks, run separate tools against each page after the list is in hand.

Putting your URL list to work

Once the extraction succeeds, the copy button puts the complete list on your clipboard in a format ready for the next step. A few patterns beginners tend to use the list for:

  • Spreadsheet import. Paste into a single column for sorting, filtering, or splitting into batches.
  • Crawler comparison. Run your own crawler against the same site and diff the two lists to find pages the sitemap is missing or claiming that do not exist.
  • Migration review. Pair each URL with its target address in a redirect map, or check whether legacy URLs still resolve.
  • Audit evidence. Save the list alongside a screenshot of the result panel as proof of what the sitemap declared at a point in time.

When you are ready to send the pages to search engines or check them individually, hand the list to whatever status, canonical, or indexing tool fits the question. The extractor does one focused job — turning sitemap XML into a clean list — and the rest of an SEO audit still belongs to specialized checks run page by page or via your crawler of choice.

If you're weighing options, Create a Descending List in Bulk URL Generator covers this in detail.