SEO TOOLS

Sitemap Parser API

What is a sitemap parser API?

A sitemap parser API fetches a site’s XML sitemap and returns its contents as structured JSON, so you never parse the XML yourself. upAPI’s Parse Sitemap operation takes a sitemap URL and returns each page’s loc, lastmod, changefreq and priority. For a sitemap index, it returns the child sitemap URLs.

Parse Sitemap fetches a sitemap over HTTP and returns it as JSON, so a crawler, an SEO audit or a content pipeline can read a site’s URL list without writing an XML parser. The only required input is the sitemap url, usually /sitemap.xml, and an optional limit caps how many entries come back; it defaults to 500 and accepts up to 5000. The response reports the sitemapUrl you sent, a type of urlset, sitemapindex or txt, how many entries the file holds as totalFound, how many were returned, and a fetchedAt timestamp. For a regular sitemap, urls holds one record per page with its loc, lastmod, changefreq and priority, and a tag the sitemap does not declare comes back as null. A sitemap index is not expanded. The operation returns type sitemapindex with the child sitemap addresses in childSitemaps, and you call it again for each child you want to read, which keeps one call to one file. A plain-text sitemap with one URL per line is read as type txt. Only the standard sitemap tags are read: image and video extension data is ignored, a .xml.gz archive is not unpacked, and the sitemap is not located through robots.txt, so you pass the sitemap address itself. A missing file raises NOT_FOUND and any other error response from the server raises UPSTREAM_ERROR. Every call authenticates with your upAPI key. The exact price and cache lifetime are on this page’s FAQ, generated from the same catalog the calls run against.

What data you get

Parse Sitemap

type
stringrequired
urls
arrayrequired
urls[].loc
stringrequired
urls[].lastmod
stringrequired
urls[].priority
numberrequired
urls[].changefreq
stringrequired
returned
integerrequired
fetchedAt
stringrequired
sitemapUrl
stringrequired
totalFound
integerrequired
childSitemaps
arrayrequired

Calling Parse Sitemap

A schema-derived request and response shape — not a captured production call, since upAPI has none to publish. Every field is real, from the operation’s own published schema.

bash
curl -X POST https://api.upapi.io/sitemap-parse.get \
  -H "X-Api-Key: $UPAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "url": "https://github.com",
  "limit": 500
}'
json — example response shape
{
  "type": "example",
  "urls": [
    {
      "loc": "example",
      "lastmod": "example",
      "priority": 1,
      "changefreq": "example"
    }
  ],
  "returned": 1,
  "fetchedAt": "example",
  "sitemapUrl": "example",
  "totalFound": 1,
  "childSitemaps": [
    "example"
  ]
}

Pricing

Every unit draws from one pooled monthly quota shared across the whole catalog. What a unit costs on each plan is on the pricing page, linked below.

Questions people actually ask

What does the sitemap parser API return for each URL?

Each entry in urls has loc, lastmod, changefreq and priority. loc is always present; the other three come back as null when the sitemap does not declare them, and priority is null when it is not a number. The response also carries type, totalFound, returned and fetchedAt.

Does it follow a sitemap index into the child sitemaps?

No. For an index it returns type sitemapindex and lists the child sitemap addresses in childSitemaps, with urls empty. Call the operation again with each child address to read that file’s URLs. limit applies to the list of children as well.

Can I use it as a sitemap scraper to extract every URL on a site?

It extracts the URLs a site chooses to list in its sitemap. It does not crawl pages or discover URLs the sitemap leaves out. Each call reads one file and returns at most limit entries from the start of it, with no offset, and totalFound tells you how many entries the file holds so you can see when a file was cut off.

Which sitemap formats and sources does it read?

Standard XML sitemaps, sitemap indexes and plain-text lists with one URL per line. A .xml.gz archive is not unpacked, image and video sitemap extension tags are ignored, and the sitemap is not found through robots.txt, so pass its address directly. A response that matches none of those formats comes back as type txt, so check totalFound rather than assuming a sitemap was found.

How much does this cost?

Every operation in this cluster costs 1 weighted unit per call, drawn from your plan’s pooled monthly quota — see upapi.io/pricing for what a unit costs on each tier.

How fresh is the data — is it cached?

every cached operation here returns a response held for 1 hour. A repeat call inside a cached window returns the cached response and is billed nothing.

Related

Start with the free plan

One key and one pooled monthly quota cover Parse Sitemap and every other API in the catalog. See what a unit costs on each plan, or run the operation first.