robots.txt and Sitemap Check
Your domain's robots.txt is read and laid out in a table, the declared sitemaps are opened and their URLs counted, and the recurring mistakes that hide a site from Google are marked.
- No sign-up needed
- URL count for each sitemap
- Common mistakes marked
Run the tool right here on the page
Determines which parts Googlebot may reach and whether your sitemap is intact.
- What it costs Free
- Data source Our own server
- Access No sign-up
The short answer
robots.txt tells crawlers which parts of the site to skip, and the sitemap tells them which URLs exist and when each last changed. If a page of yours will not index, these two files are the first thing to look at, before anything else.
The worst event for a site fits in one line: Disallow: / . Every staging environment carries that line, and rightly, but when the site goes from staging to the live server it usually travels with it and removes the whole site from Google. No warning shows anywhere; traffic just slides to zero over a few weeks and nobody knows why. After every site migration, check this on day one.
The mistake we keep meeting on Iranian WordPress sites is blocking wp-content. That logic was once sound and is not now: to judge whether your page is mobile friendly, Google must see its CSS and JavaScript, all of which sit inside wp-content. Block that folder and Google sees an unstyled page and judges a shapeless one.
For sitemaps, the number worth reading is the gap between the URL count in the sitemap and the indexed page count in Search Console. If the sitemap holds a thousand URLs and Google has indexed two hundred, the sitemap is not at fault; Google has seen eight hundred pages and judged them not worth indexing. The right question then is the quality of those pages, not the sitemap file.
The data source and method
The file is read by our server straight from your domain; there is no external service and no fee. When robots.txt names no sitemap, the usual paths, sitemap.xml, sitemap_index.xml and wp-sitemap.xml, are attempted as well. A pattern we run into constantly: a large share of Iranian WordPress sites still keep wp-content blocked, which leads Google to render the page without CSS and score it as not mobile friendly.
The situations where it helps
The first item to inspect when a page refuses to index. Also after every site move or host swap, since a staging robots.txt usually travels with Disallow: / and removes the entire site from Google.
Where it falls short
This tool reports what robots.txt permits, not what Google actually did. A page can be fully allowed in robots.txt and still go unindexed, because Google judged it not worth it or its canonical points elsewhere. The reverse also occurs: a page blocked in robots.txt but linked from elsewhere sometimes shows in results with an empty title, because robots withholds permission to crawl, not permission to index. For do-not-index, the right instrument is the noindex tag, not robots.txt. To see what Google did with a URL, the only correct source is that site's own Search Console.
A lesson from working with it
If robots.txt declares no sitemap, this tool also tries the usual paths: sitemap.xml, sitemap_index.xml and wp-sitemap.xml. We added that on purpose, because in practice we have often seen a sitemap that exists and is healthy but is declared nowhere, with the owner assuming there is none. Declaring the sitemap in robots.txt is one line, and that line is what tells a crawler reaching your site for the first time where to begin.
What it costs
This tool executes on our own infrastructure with no purchased data behind it, so it costs nothing and stays that way. Its only limit is a daily cap that stops one bot from taking the whole capacity.
The daily cap is measured per visitor. If you need more, a free account lifts your limit.
Common questions about this tool
Is it a problem if I have no robots.txt?
Not necessarily. No file means everything is allowed, which for most small sites is exactly what they want. The file becomes necessary when you want to hold sections such as internal search pages or shop filters out of crawling.
Can Disallow remove a page from Google?
Not reliably. If the page is linked from elsewhere it can still appear in results with an empty title. For real removal use the noindex tag, and more importantly leave that page unblocked in robots.txt, because Google must be able to crawl it to see the noindex.
How often should a sitemap be updated?
Automatically, with every new piece of content. Almost every WordPress SEO plugin does this. A hand-built sitemap left un-updated for months is worse than none, because it sends the crawler to URLs that no longer exist.
Reads robots.txt and sitemap.xml and says what is accidentally blocked.
One rule throughout: See fuses onto a word and makes a single new one, never two words side by side. Exactly as See and commerce became Seemerce.
Neighbouring names
- SSL check SeeLock
- Rising trends SeeRise
- Keyword overview SeeVolume
- Keyword suggestions SeeIdeas
- Site keywords SeeTerms
- Ranked keywords SeeMatch
- SERP viewer SeeSerp