How to Run an Indexability Audit
Audit crawl access, index directives, canonicals, rendering and Search Console evidence without mistaking exclusions for indexing errors.

Maya opens Search Console on Monday morning and sees 240 pages under “Crawled — currently not indexed.” Her first instinct is to request indexing for all of them. That feels productive, but it skips the decision that matters: should every one of those URLs represent a useful page in search?
An indexability audit answers that question before anyone changes robots.txt, canonicals, sitemaps or content. It follows the page from request to search-engine observation and separates intended exclusions from real barriers.
Start With the Publishing Decision
Write the intended state for each template in plain language. A product page may need to be public, crawlable, indexable, self-canonical and included in the sitemap. An account page may need to be public to users after login but excluded from search. An expired campaign might redirect to a close successor. Without this policy, the same technical state can look correct to one reviewer and broken to another.
| Question | Evidence to Record | Why It Comes First |
|---|---|---|
| Should this URL appear in search? | Page purpose, audience and owner | Defines success before tools assign labels |
| Can an anonymous crawler request it? | Status, final URL, MIME type and access rules | Later checks require a usable document |
| Does the page permit indexing? | Meta robots and X-Robots-Tag | Shows explicit page-level instructions |
| Which URL should represent it? | Declared canonical, redirects and duplicate set | Prevents competing representatives |
| How is it discovered? | Confirmed internal sources and sitemap membership | Separates existence from promotion |
| What did Google observe? | Dated URL Inspection and Search Console evidence | Adds provider evidence without replacing live facts |
Follow the Page Through Six Gates
Gate 1: Transport
Request the exact URL and retain every redirect hop. A timeout, DNS failure, 4xx response, 5xx response or non-HTML file does not provide a valid page for content checks. Stop there, keep the source-link evidence and assign the transport owner.
Gate 2: Crawl Access
Read robots.txt for the relevant user agent and test authentication, bot protection and blocked resources. A crawl block and a noindex directive solve different problems. If the crawler cannot fetch the page, it may not be able to read the page-level directive.
Gate 3: Index Directives
Compare source and rendered meta robots with X-Robots-Tag headers. Preserve the literal values. Do not collapse unavailable evidence into indexable or non-indexable. Multiple directives should be interpreted conservatively and checked against the publishing policy.
Gate 4: Canonical Selection
Resolve the declared canonical, fetch its destination and confirm that it is successful, indexable and appropriate. Then compare redirects, internal links, sitemaps and hreflang. These signals should support the same preferred URL instead of asking a canonical tag to overrule the rest of the site.
Gate 5: Rendered Meaning
Compare initial HTML and rendered content. Confirm that the main heading, useful body, navigation and structured data survive the rendering process. Record the render wait and important network failures. A screenshot proves appearance, not necessarily stable crawl evidence.
Gate 6: Observed Search State
Use URL Inspection and Search Console after the live evidence is understood. Provider data can show a selected canonical, last crawl or coverage label for that property and time. It cannot retroactively make a failed response valid or prove that a resubmission will force indexing.
An Illustrative 240-URL Review
Maya’s 240 URLs divide into five groups. Forty are expired products that correctly redirect. Sixty are wanted category pages accidentally blocked by a copied staging robots.txt file. One hundred are product pages whose template canonicals point to a category. Twenty return intermittent 503 responses. The final twenty were never measured because the crawl stopped.
The honest outcome is 40 accepted exclusions, 60 crawl-control failures, 100 canonical conflicts, 20 transport failures and 20 unavailable observations. “Two hundred indexability errors” would hide the owners and recommended actions. “All 240 pages fixed” would incorrectly count missing evidence as success.
Turn the Audit Into Work
- Keep the exact affected URLs and eligible denominator.
- Attach the response, directive, canonical or rendered evidence that supports the finding.
- Group pages only after samples show one shared cause.
- Name the smallest safe change, owner, reviewer and rollback condition.
- Define the fresh check before implementation.
- Re-run the same detector with comparable scope after release.
Eligibility is something your website can make clearer. Indexing remains a search-engine decision.
LLMIC editorial principle
Frequently Asked Questions
What is an indexability audit?
It is a review of whether intended public pages can be requested, rendered, interpreted and considered for indexing without conflicting technical signals.
Is crawlability the same as indexability?
No. Crawlability concerns whether a crawler can request a URL. Indexability also depends on the response, robots directives, canonical signals and whether the page supplies usable content.
Why is an indexable page not indexed?
Eligibility does not force selection. A page may be new, duplicated, weakly linked, difficult to render, low value for the query, or simply not selected by the search engine.
Should noindex pages appear in an XML sitemap?
An indexable-only sitemap should normally contain preferred URLs intended for indexing. Listing a noindex URL sends conflicting signals and creates avoidable review noise.
Can robots.txt remove a page from search results?
Robots.txt controls crawling. A blocked URL may remain known. Use the control that matches the actual goal and remember that a blocked page-level noindex cannot be read.
How many URLs should an indexability audit sample?
Sample every important template, conversion path and known exception. Use samples to learn causes, then measure the complete eligible group before reporting counts.
Can an indexability audit guarantee rankings?
No. It can identify and remove technical barriers, but indexing and rankings also depend on content, relevance, competition, links, demand, device, location and time.
Sources, Statistics and Measurement Notes
- Robots Meta Tag, Data-Noindex and X-Robots-Tag Specifications — Google Search Central. Defines page-level and HTTP-header indexing controls; it does not guarantee whether a permitted page will be indexed. Accessed 30 September 2026.
- URL Inspection Tool — Google Search Console Help. Explains live tests and indexed-version observations for verified properties. Accessed 30 September 2026.
- Canonicalization — Google Search Central. Explains canonical signals and why a declared canonical remains a signal rather than a command. Accessed 30 September 2026.
- XML Sitemaps — Google Search Central. States that sitemap submission is a hint and does not guarantee crawling or indexing. Accessed 30 September 2026.
- Ubersuggest keyword overview: indexing issues — Ubersuggest. India, English, checked 30 September 2026: “indexing issues” returned 40 monthly searches and SEO difficulty 18. “How to check if a page is indexed” returned 20 and difficulty 6. These are dated editorial inputs, not traffic forecasts. Accessed 30 September 2026.
Run the Same Review in LLMIC
Use LLMIC to retain the crawl identity, final response, directives, canonical signals and page evidence together. Open affected records before creating a bulk task, then save a fresh crawl after the approved change.
Before selecting a bulk action, inspect an important page, a normal page and a known exception. Keep facts, interpretations and recommendations in separate fields. Save the original scope and exclusions with the task so a later reviewer can reproduce the denominator. If a required response, header or rendered document is unavailable, collect it again instead of guessing. This discipline takes longer than clicking “fix all,” but it prevents a weak observation from becoming a site-wide change and gives the execution team a precise result to verify.
Check the website behind the idea.
Use LLMIC to crawl pages, review the evidence and organise the next actions in one native Mac workspace.