Sitemap Not Found? How to Tell a Missing Sitemap From a Blocked Crawler
If an audit reports “No XML sitemap found” but your sitemap opens perfectly in your own browser, the sitemap is rarely the problem. Here's how to tell the three causes apart and fix each one.
If an audit reports "No XML sitemap found", or crawls only a single page, but you can open your sitemap in a browser without any trouble, the sitemap is almost never the real problem. Something on the site is answering the crawler differently from how it answers you. This guide shows you how to work out which of three things is happening, and exactly what to do about each.
The three things this message can mean
- Your sitemap exists, but at a URL the audit did not check. The most common cases are /sitemap_index.xml, which Yoast SEO and Rank Math use, and /wp-sitemap.xml, which WordPress generates on its own.
- Your sitemap exists and loads perfectly for you, but your server, CDN, firewall or a security plugin refuses or redirects the crawler. This is the most common cause by far, and the most confusing, because everything looks fine from your own screen.
- You genuinely do not have a sitemap.
They need completely different fixes, so it is worth spending two minutes identifying which one you have before changing anything.
Step 1: Read what the audit actually said
MRK Audit distinguishes the three cases for you. The wording in the Sitemap tab and the Pages Crawled card tells you which one you are looking at.
- "No sitemap found - we checked robots.txt and 11 common locations" means we looked everywhere sensible and there was nothing there. That is cause three.
- "Sitemap unreadable - the site returns a web page, not XML" means we asked for the sitemap and were handed an HTML page instead. That is cause two: the crawler is being turned away, and the message will name the URL it was redirected to.
- "Sitemap blocked - HTTP 403" (or 401, 429, 503) means the server refused outright. Also cause two.
Step 2: Check the obvious URLs yourself
Open each of these in your browser and note which ones return XML rather than a normal web page:
- yoursite.com/sitemap.xml
- yoursite.com/sitemap_index.xml
- yoursite.com/wp-sitemap.xml
- yoursite.com/robots.txt
Your robots.txt should contain a line beginning with "Sitemap:" that points at the real file. If it does not, add one. Every serious crawler reads robots.txt before it guesses at URLs, so that single line fixes discovery for search engines, AI crawlers and audit tools all at once.
Step 3: Prove whether the crawler is being blocked
The giveaway for cause two is that the same URL behaves differently depending on who asks for it. You do not need any technical setup to test this - you just need a second vantage point.
- Open the sitemap URL on your phone using mobile data with Wi-Fi turned off. A different network means a different IP address. If it works on one and not the other, you have a block.
- Ask a colleague in another country to open the same URL and tell you what they see.
- In Google Search Console, submit the sitemap under Sitemaps and see whether Google can fetch it. Google reporting "Couldn't fetch" while it loads fine for you is the same signal.
- Try an online HTTP header checker. If it reports a redirect or a 403 for a URL that loads normally in your browser, the site is treating outside tools differently from you.
What a blocked crawler actually looks like
These are the patterns worth recognising, because each one points at a different kind of rule:
- Every URL except the homepage redirects to the homepage. That is usually a geo-restriction or "block datacentre traffic" plugin, and it is the hardest to spot because nothing ever returns an error.
- The sitemap returns HTTP 403 or 401. That is a firewall, a WAF rule, or bot protection.
- HTML comes back where XML should be. Either a soft 404, or the same redirect behaviour as above.
- Static files such as images, CSS and JavaScript load fine, but anything your CMS generates does not. That narrows the rule down to your CMS or a plugin rather than the web server.
Step 4: Fix it
If your sitemap is simply at a different URL
- Add a Sitemap: line to robots.txt pointing at the real file. This is the single highest-value fix, and it takes one minute.
- Optionally add a redirect from /sitemap.xml to your real sitemap, so the conventional path also works. Yoast does this automatically; some setups lose it after a migration.
- Submit the sitemap in Google Search Console and Bing Webmaster Tools.
If the crawler is being blocked or redirected
- Ask whoever manages the site to allowlist the crawler. MRK Audit fetches from 46.202.158.104. In Cloudflare that is Security, then WAF, then Tools, then IP Access Rules, then Allow.
- Look for a security or geo-restriction plugin. Any rule that sends non-domestic or datacentre traffic to the homepage will produce exactly this behaviour. Exempt /robots.txt and any /sitemap path from it.
- Check server-level rules too - .htaccess, nginx configuration, ModSecurity, and any "bad bot" blocklist. These are often installed once and forgotten.
- If the site is behind Cloudflare, check Bot Fight Mode and any custom WAF rules, and review the Security Events log for your own blocked requests.
One thing worth saying to whoever owns the site: this is not really about one audit tool. A rule that hides your sitemap from us hides it from other crawlers too, and you will usually never see an error message telling you so. If your sitemap cannot be fetched from outside your own network, that is worth fixing on its own merits.
If you genuinely have no sitemap
- WordPress: install Yoast SEO, Rank Math or All in One SEO and one is generated for you. WordPress 5.5 and later also publishes /wp-sitemap.xml with no plugin at all.
- Shopify, Squarespace, Wix and Webflow all generate /sitemap.xml automatically - there is nothing to install.
- Custom or static sites: generate one as part of your build, or use a crawler-based sitemap generator and upload the result.
- Whichever route you take, finish by adding the Sitemap: line to robots.txt and submitting it in Search Console.
Why this matters more than it looks
A sitemap tells search engines and AI crawlers which pages exist without relying on them finding every internal link. Without one that can actually be fetched, discovery falls back to link crawling, which is slower and misses pages that are not well linked - and AI answer engines, which crawl far less thoroughly than Google, tend to miss the most.
It also affects what an audit can tell you. If only your homepage can be read, every site-wide number in the report describes one page rather than your site.
What MRK Audit does about this automatically
- Reads the Sitemap: directives in robots.txt first, then checks eleven conventional locations including /sitemap_index.xml and /wp-sitemap.xml.
- Validates that what comes back is genuinely XML, so an HTML page returned with a 200 status is never mistaken for a sitemap.
- Distinguishes "we could not find one" from "we were refused", and names the URL it was redirected to so you can see the rule in action.
- Automatically retries blocked requests from a United States IP address, which resolves the common case of a site that refuses European or datacentre traffic.
If you have worked through this guide and the audit still cannot read your sitemap, send us the site and the exact message you are seeing, and we will trace the request end to end and tell you what your server is doing.
Audit