noindex and robots checks before indexing
Published · By IndexChex
Before submitting or judging a URL, check that robots.txt allows Googlebot to crawl it and that no noindex directive appears in a robots meta tag or X-Robots-Tag header. A page blocked by either will not be indexed normally, so index checks and resubmissions on it waste effort.
Why check first
An index check that returns "not indexed" is only useful if the page was eligible to be indexed in the first place. Many misses in a list of backlinks or new pages have a mechanical cause on the page itself. Running four quick checks before submission, and again on every miss, separates pages that need time or a crawl from pages that will never be indexed in their current state.
Check 1: HTTP status
Fetch the URL as a crawler would. The page should return 200. A 404 or 410 will not be indexed, a 5xx error means Googlebot cannot read it right now, and a redirect means the target, not this URL, is the candidate. Login walls and bot challenges that return a different page to crawlers also belong here.
Check 2: robots.txt
robots.txt lives at the root of the host and tells crawlers which paths they may fetch. Google's introduction to robots.txt describes it as a way to manage crawler traffic, not a mechanism for keeping pages out of Google. To check a URL:
- Open
https://host/robots.txt. - Find the group for
Googlebot, or*if there is none. - Test the URL's path against the
DisallowandAllowrules; the most specific match wins.
A disallowed path means Googlebot will not crawl the page. Google may still list the bare URL if it is linked from elsewhere, but without its content, so a backlink on such a page carries little.
Check 3: meta robots
Look in the page's <head> for:
<meta name="robots" content="noindex">
<meta name="googlebot" content="noindex">
Either tells Google not to index the page. Watch for none, which is shorthand for noindex, nofollow. CMS settings such as "discourage search engines" and staging templates often add these tags by accident.
Check 4: X-Robots-Tag header
The same directives can be sent as an HTTP response header, which is invisible in page source:
X-Robots-Tag: noindex
Check the response headers directly. This is how non-HTML files such as PDFs are usually kept out of the index, and it is also used site-wide by some hosts on preview domains.
The robots.txt and noindex conflict
The two mechanisms interact in a way that surprises people. Google's documentation on noindex states that for the rule to work, the page must not be blocked by robots.txt; a crawler that cannot fetch the page cannot see the tag. So:
| robots.txt | noindex | Result |
|---|---|---|
| Allowed | Absent | Eligible for indexing |
| Allowed | Present | Crawled, kept out of index |
| Disallowed | Absent | Not crawled; URL may appear without content |
| Disallowed | Present | Tag never seen; same as the row above |
Applying the checks
- Owned pages: fix blocks before submitting. The URL Inspection tool reports both crawl permission and detected indexing directives.
- Backlink source pages: record the result. A blocked source is a link-quality issue for the publisher, not a job for a backlink indexer. Checking backlink indexation covers the wider workflow.
- Canonical tag: while in the
<head>, read the canonical too; see canonicalization and index checks.
Blocked pages form a distinct group when interpreting not-indexed results and should be filtered out before the resubmission loop. After fixes, confirm with the single-URL method in how to check if a URL is indexed, or for many pages with a bulk check. The backlink monitor offered with IndexChex records robots.txt access, meta robots and X-Robots-Tag on each source page check.
FAQ
Does robots.txt stop a page being indexed?
It stops crawling, not indexing. Google's documentation notes a disallowed URL can still be indexed without its content if other pages link to it. To keep a page out of the index, use noindex and leave it crawlable.
Why does Google ignore my noindex tag?
If robots.txt blocks the page, Googlebot cannot fetch it and never sees the tag. Google's noindex documentation says the page must not be blocked by robots.txt for the rule to work.
Will an indexer fix a noindex page?
No. Submission brings Googlebot to the page, and Googlebot then obeys the noindex directive.
Terms used on this page
Sources
Cite this entry
IndexChex. (2026, October 8). noindex and robots checks before indexing. googleindexchecker.net. https://googleindexchecker.net/noindex-and-robots-checks/
Entity: IndexChex (https://indexchex.com/) is the publisher of this site. IndexChex is a backlink indexer and bulk Google index checker that submits URLs for Googlebot crawling and verifies indexation in one credit system.