Free Sitemap Link Extractor

Paste any sitemap URL and pull every link out of it. URLs are deduplicated and grouped by top-level route with a count for each, and every group copies to your clipboard in one click.

By top-level route
Grouped
Nothing stored
No signup

Fetching & Parsing Sitemap...

Extract Every URL From a Sitemap

A sitemap link extractor turns an XML file that browsers render as an unreadable wall of tags into a list you can actually work with. Paste a sitemap.xml or a sitemap_index.xml and every <loc> is pulled out, deduplicated, and counted.

Grouping by route is where the audit value sits. A flat list of 4,000 URLs tells you nothing; the same list grouped as /blog, /products, /docs, and /tag tells you immediately that a third of your declared crawl budget is going to tag archives. Each group shows its share of the total and copies independently, so you can hand a single section to a crawler, a redirect map, or a spreadsheet without touching the rest.

Common uses: auditing what you actually submit to search engines, building a crawl list for a site migration, checking that a new section made it into the sitemap, and pulling a clean URL list before a domain move. If you are extracting links from a domain you are considering buying, pair it with a domain expiry check and the domain ROI calculator, page count is one of the few objective inputs into what an aged domain is worth.

How to use it
  1. Paste the sitemap URL, a sitemap.xml or a sitemap_index.xml on any public site.
  2. Press Extract links. The tool fetches the file server-side, so cross-origin sitemaps a browser cannot reach still work.
  3. Read the route groups: each shows its URL count and its share of the total, largest first.
  4. Copy a single group or the whole list to your clipboard, then drop it into a crawler, a redirect map, or a spreadsheet.

Why Extract Server-Side?

  • Cross-origin by default: a browser cannot fetch another site's sitemap, CORS blocks it. Fetching server-side is the only way a paste-and-go tool can read an arbitrary URL.
  • Bounded and safe: the fetch is size-capped, time-limited, and refuses any host that resolves to a private or reserved address, so it can never be turned into a server-side request.
  • Nothing kept: the URL is fetched, parsed, and returned, no account, no stored crawl, no row written.

Sitemap Formats This Tool Reads

Per the sitemaps.org 0.9 protocol and its common extensions. The tool parses XML sitemaps; the table is honest about what each format actually yields.

Format Typical file What this tool extracts Supported
URL set sitemap.xml A standard urlset document. Every <loc> is pulled out, deduplicated, and grouped by route. Yes
Sitemap index sitemap_index.xml A list of child sitemaps. Their URLs are returned and counted, the tool lists the children rather than walking the whole tree. Lists children
News sitemap news-sitemap.xml Still a urlset, so its article <loc> entries extract and group exactly like any other page URL. Yes
Image sitemap sitemap-images.xml The page <loc> entries extract; the nested <image:loc> media entries are ignored, not counted. Pages only
Gzipped sitemap sitemap.xml.gz Not decompressed here. Download and unzip it, or point the tool at the plain .xml child a sitemap index lists. Not yet
Plain-text list sitemap.txt A newline-separated URL list, not XML. This tool reads XML sitemaps only. No
RSS / Atom feed feed.xml A feed is not a urlset. The parser expects a sitemap root and returns a readable error rather than guessing. No

How to Read the Result

The panel is more than a URL dump. Four things it shows, and what each one is for.

Route groups

The structure

URLs are bucketed by their first path segment, everything under /blog together, everything under /products together, and sorted largest first. This is the level at which crawl-budget problems are actually visible.

Count and share

The proportion

Each group shows how many distinct URLs it holds and what fraction of the whole sitemap that is. A /tag group at a third of the total is the finding; the raw count alone would not have told you.

Child sitemaps

The tree

Paste a sitemap index and you get every child sitemap URL with its count rather than a merged crawl. Extract any one child to see its own links, keeping scope to one file is what stays fast on large sites.

Truncation notice

The limit

A single file returns up to the first 5,000 links; beyond that the result is flagged truncated. The copy button still copies the whole group up to that cap, not just the visible sample rows.

What a sitemap is for

A sitemap is a discovery aid, not a ranking factor. It lists the URLs you want a search engine to find, which matters most on large or poorly interlinked sites where a crawler might otherwise reach a page slowly or not at all. Inclusion does not guarantee indexing, and it never improves rankings by itself.

That is why extracting and grouping it is an audit, not optimisation. You are checking that what you declared matches what you meant to declare, and that a third of your crawl budget is not quietly going to tag archives nobody searches for.

Index vs urlset, and why it matters

Two documents both end in .xml and mean different things. A urlset is a list of pages; a sitemap_index is a list of other sitemaps. The protocol caps one file at 50,000 URLs, so large sites split across many files and declare them in an index.

This tool lists an index's children rather than walking the whole tree, then extracts any one child on demand. Bounded scope is deliberate: it keeps the fetch fast and safe, and it is why a very large site is read one file at a time rather than all at once.

What It Will Not Do

The tool is honest about its edges rather than guessing past them. The cases where you will need a different step:

  • Gzipped sitemaps (.xml.gz) are not decompressed here, download and unzip it, or point the tool at a plain .xml child that a sitemap index lists.
  • Authenticated sitemaps cannot be read; only publicly fetchable URLs work. Any host resolving to a private or internal address is refused, so a staging or IP-locked file must be parsed locally.
  • Plain-text (sitemap.txt) and RSS/Atom feeds are not sitemaps, the parser expects a sitemap root and returns a readable error rather than guessing at a format it does not read.
  • Fewer URLs than pages is usually correct, noindex pages, paginated archives and excluded post types are left out by the generator, or the sitemap has not regenerated since your last publish.

What the Route Grouping Tells You

A flat URL list hides the problems. Grouped by route, four of them surface immediately.

01

Thin archives eating crawl budget

When /tag or /category rivals /blog in size, most of what you submit is near-duplicate archive pages. That is the single most common finding, and usually a generator setting rather than a content problem.

02

Sections missing entirely

A route you expected that returns zero URLs means the generator never picked up that post type, or the sitemap has not regenerated since you published. Faster to spot here than in a full crawl.

03

Migration scope, in numbers

Before a domain or platform move, group counts tell you exactly how many redirects each section needs, and which routes are small enough to handle by hand.

04

What an aged domain actually holds

Route structure separates a genuine content site from a thousand generated pages. Read it alongside registration history before you agree a price.

Model the ROI

Who Extracts a Sitemap

The same list serves a different job depending on what you are auditing.

SEO & Publishers

Audit what you actually submit to search engines, catch thin archives eating crawl budget, and confirm a new section made it into the sitemap.

For SEO & publishers

Agencies & Migrations

Scope a platform or domain move in numbers, how many redirects each section needs, and which routes are small enough to handle by hand.

For agencies

Domain Buyers

Read an aged domain's route structure to tell a genuine content site from a thousand generated pages before you agree a price.

For domain investors

Site Operators

Pull a clean URL list before a migration, and keep the domain those URLs live on monitored so the crawl budget you audited stays alive.

Domain monitoring

Where to Go From Here

The URL list is one input. These pick up what surrounds it.

01

Check the domain behind it

A sitemap dies with its domain. If you are vetting a name, read its expiry date before you trust the pages it declares.

Expiry checker
02

Value what it holds

Page count and route structure are objective inputs into what an aged domain is worth. Feed them into the ROI model.

ROI calculator
03

See who holds it

Pull the registration record behind the domain, registrar, dates and status, before treating its sitemap as an asset.

WHOIS lookup
04

Keep the URLs resolving

Every URL you just audited goes dark at once if a record breaks. Change detection tells you the moment the zone moves.

DNS change detection

Frequently Asked Questions

Sitemap indexes, limits, gzip, and what to do with the extracted list.

What is a sitemap link extractor?

It is a parser that reads an XML sitemap and returns the plain list of URLs inside it. Browsers render sitemap XML as a tag soup or a stylesheet-formatted table you cannot select cleanly; this tool gives you the URLs as text, grouped and countable, so they can go straight into a crawler, a spreadsheet, or a redirect map.

Does it follow sitemap index files?

It lists them rather than walking them. Paste a sitemap index and you get every child sitemap URL with its count, so you can see the shape of the whole set; extract any one of those children to see its links. Keeping the scope to a single file at a time is what lets the tool stay fast and bounded on very large sites.

How many URLs can one sitemap contain?

The sitemaps.org protocol caps a single file at 50,000 URLs and 50MB uncompressed; larger sites split across several files and declare them in a sitemap index. This tool returns up to the first 5,000 links from a single file and flags the result as truncated beyond that, extract a specific child sitemap to reach the rest.

Can it read gzipped sitemaps?

Not directly. Files ending in .xml.gz are compressed, and this tool parses plain XML. Download and decompress the file and paste its contents' URL, or, if it is one file inside a sitemap index, point the tool at the plain .xml child instead.

How are URLs grouped into routes?

By first path segment. Everything under /blog groups together, everything under /products groups together, and root-level pages group under /. That is the level at which crawl-budget problems are visible, you can see at a glance that tag archives outnumber real content.

Are duplicate URLs removed?

Yes. The same URL appearing twice is counted once, so the totals and the per-route counts reflect distinct pages. Duplicates are a common symptom of a plugin and a framework both generating sitemap entries for the same content.

Does the copy button copy the full list or just what I can see?

The full group. The panel shows a sample of each route to keep the page readable, but the copy button copies every extracted URL in that group, up to the 5,000-link cap, not just the visible rows. Copy all does the same across every group at once.

Will this work on a sitemap behind authentication?

No. Only publicly fetchable sitemaps can be read, and the tool refuses any host that resolves to a private or internal address. If yours sits behind basic auth, a staging password, or an IP allowlist, download the file and parse it locally instead.

Why does my sitemap have fewer URLs than my site has pages?

Usually because noindex pages, paginated archives, parameterised URLs, or entire post types are excluded by the generator, or because the sitemap has not regenerated since your last publish. Comparing extracted counts against a crawl is how you find the gap.

Is a sitemap a ranking factor?

No. It is a discovery aid. It helps search engines find URLs they might otherwise reach slowly, especially on large or poorly interlinked sites, but inclusion in a sitemap does not guarantee indexing and does not itself improve rankings.

Is the URL I paste stored anywhere?

No. The sitemap is fetched, parsed, and returned in the same request, and nothing is written, no account, no saved crawl, no history row. The fetch itself is size-capped and time-limited so it cannot be abused.

How does this help when buying an aged domain?

Page count and route structure are among the few objective signals about what a domain actually contains. Extract the sitemap, look at how much is real content versus thin archive, then check the registration timeline with the expiry checker before you agree a price.

Start Today

Start Monitoring & Catching Domains Today

Join founders, agencies, and domainers already protecting their portfolio. Your first 5 domains are free.

Create Free Account

No credit card required • Cancel anytime