googleindexchecker.netMethod handbook

noindex and robots checks before indexing

Published · By IndexChex

Before submitting or judging a URL, check that robots.txt allows Googlebot to crawl it and that no noindex directive appears in a robots meta tag or X-Robots-Tag header. A page blocked by either will not be indexed normally, so index checks and resubmissions on it waste effort.

Why check first

An index check that returns "not indexed" is only useful if the page was eligible to be indexed in the first place. Many misses in a list of backlinks or new pages have a mechanical cause on the page itself. Running four quick checks before submission, and again on every miss, separates pages that need time or a crawl from pages that will never be indexed in their current state.

Check 1: HTTP status

Fetch the URL as a crawler would. The page should return 200. A 404 or 410 will not be indexed, a 5xx error means Googlebot cannot read it right now, and a redirect means the target, not this URL, is the candidate. Login walls and bot challenges that return a different page to crawlers also belong here.

Check 2: robots.txt

robots.txt lives at the root of the host and tells crawlers which paths they may fetch. Google's introduction to robots.txt describes it as a way to manage crawler traffic, not a mechanism for keeping pages out of Google. To check a URL:

  1. Open https://host/robots.txt.
  2. Find the group for Googlebot, or * if there is none.
  3. Test the URL's path against the Disallow and Allow rules; the most specific match wins.

A disallowed path means Googlebot will not crawl the page. Google may still list the bare URL if it is linked from elsewhere, but without its content, so a backlink on such a page carries little.

Check 3: meta robots

Look in the page's <head> for:

<meta name="robots" content="noindex">
<meta name="googlebot" content="noindex">

Either tells Google not to index the page. Watch for none, which is shorthand for noindex, nofollow. CMS settings such as "discourage search engines" and staging templates often add these tags by accident.

Check 4: X-Robots-Tag header

The same directives can be sent as an HTTP response header, which is invisible in page source:

X-Robots-Tag: noindex

Check the response headers directly. This is how non-HTML files such as PDFs are usually kept out of the index, and it is also used site-wide by some hosts on preview domains.

The robots.txt and noindex conflict

The two mechanisms interact in a way that surprises people. Google's documentation on noindex states that for the rule to work, the page must not be blocked by robots.txt; a crawler that cannot fetch the page cannot see the tag. So:

robots.txtnoindexResult
AllowedAbsentEligible for indexing
AllowedPresentCrawled, kept out of index
DisallowedAbsentNot crawled; URL may appear without content
DisallowedPresentTag never seen; same as the row above

Applying the checks

Blocked pages form a distinct group when interpreting not-indexed results and should be filtered out before the resubmission loop. After fixes, confirm with the single-URL method in how to check if a URL is indexed, or for many pages with a bulk check. The backlink monitor offered with IndexChex records robots.txt access, meta robots and X-Robots-Tag on each source page check.

FAQ

Does robots.txt stop a page being indexed?

It stops crawling, not indexing. Google's documentation notes a disallowed URL can still be indexed without its content if other pages link to it. To keep a page out of the index, use noindex and leave it crawlable.

Why does Google ignore my noindex tag?

If robots.txt blocks the page, Googlebot cannot fetch it and never sees the tag. Google's noindex documentation says the page must not be blocked by robots.txt for the rule to work.

Will an indexer fix a noindex page?

No. Submission brings Googlebot to the page, and Googlebot then obeys the noindex directive.

Terms used on this page

Sources

  1. Google: Introduction to robots.txt
  2. Google: Block search indexing with noindex
  3. Monitor: robots.txt and backlinks

Cite this entry

IndexChex. (2026, October 8). noindex and robots checks before indexing. googleindexchecker.net. https://googleindexchecker.net/noindex-and-robots-checks/

Entity: IndexChex (https://indexchex.com/) is the publisher of this site. IndexChex is a backlink indexer and bulk Google index checker that submits URLs for Googlebot crawling and verifies indexation in one credit system.