Why is Google not indexing my pages?
Google skips a page for one of three reasons: it can't reach it, it's told not to keep it, or it doesn't think the page is worth keeping. Blocks like noindex, robots.txt rules and server errors are quick to fix. "Crawled - currently not indexed" is the tough one, because it's a quality and duplication call, not a bug. Read the exact status in Search Console's Page indexing report, then fix that cause, not a guessed one.
🎯 The Fast Version
- Open Indexing → Pages in Search Console and read the exact status. Each one points to a different fix
- "Discovered" means Google never fetched the page. "Crawled" means it did, and passed
- Only chase pages that should earn traffic. A zero-unindexed report is not the goal
- Request Indexing is a nudge, not a repair. Change something first
- Unindexed pages can't appear in Google's AI features either, so this now affects more than blue links
- Reading the reports properly: Google Search Console Guide →
- The wider technical picture: Technical SEO Guide →
- How Googlebot spends its time: Crawl Budget Optimisation Guide →
- Running a full site check: SEO Audit Guide →
The most common mistake I see is people treating every unindexed URL the same. Someone exports 4,000 excluded pages, panics, and starts resubmitting all of them. Half of those URLs were filters and old redirects that should never be indexed. Sort them into "should be indexed" and "shouldn't" before you touch anything. That one hour of sorting saves days.
1. What an Indexing Issue Actually Is
Getting into Google is a pipeline: Google finds the URL, crawls it, renders it, and then decides whether to store it. An "indexing issue" is a page that fell out somewhere along that line. Where it fell out tells you what to fix, which is why lumping everything under "not indexed" leads to random fixes.
Two things people mix up: crawling is Googlebot fetching the page, indexing is Google deciding to keep and rank it. A page can be crawled and never indexed. It can also be indexed while blocked from crawling, which sounds odd but happens, and I'll come back to it.
2. Check It's a Real Problem First
Don't trust a quick site:yourdomain.com/page search. It's a rough hint, not a report. Use the URL Inspection tool in Search Console instead. Paste the exact URL and you'll see whether Google has it, when it last crawled it, and which canonical it chose.
Also note that the Page indexing report can lag. If a URL is listed as not indexed but URL Inspection says "URL is on Google", believe the inspection and move on.
3. Every Search Console Status, Decoded
This is the table I keep open when I'm auditing. Severity assumes the page is one you actually want indexed.
| Status | What it means | Risk | First move |
|---|---|---|---|
| Discovered - currently not indexed | Google knows the URL but hasn't fetched it. Last crawl date is empty | MEDIUM | Check server health, internal links and sitemap bloat |
| Crawled - currently not indexed | Google fetched and read it, then chose not to index it | HIGH | Improve unique value, add internal links, check duplicates |
| Duplicate without user-selected canonical | Google found duplicates and you didn't declare a preferred version | MEDIUM | Add self-referencing canonicals |
| Duplicate, Google chose different canonical than user | Google ignored your canonical tag | HIGH | Make redirects, links and sitemap all agree with your canonical |
| Alternate page with proper canonical tag | Working as intended, the page points at its main version | LOW | Usually nothing |
| Excluded by 'noindex' tag | The page tells Google to keep it out | HIGH | Remove noindex if it's unintentional |
| Blocked by robots.txt | Google isn't allowed to fetch it | HIGH | Fix the disallow rule |
| Soft 404 | Returns 200 but looks empty or like an error page | MEDIUM | Add real content or return a proper 404/410 |
| Server error (5xx) | Your server failed when Google asked | HIGH | Fix the server, then validate |
| Not found (404) | The URL doesn't exist | LOW | Fine, unless it's internally linked or has backlinks |
| Page with redirect | The URL redirects somewhere else | LOW | Fine, just link to the final URL |
| Indexed, though blocked by robots.txt | Google indexed it from links without reading it | MEDIUM | Allow crawling and use noindex if you want it out |
4. Discovered - Currently Not Indexed
Here's the detail most guides skip: Google hasn't fetched the page, so it hasn't judged your content at all. Google's own description says it typically held back because crawling was expected to overload the site. So before rewriting a single paragraph, look at the server.
- Check server logs and uptime for 5xx spikes or slow responses around the times Googlebot visits.
- Look at the Crawl stats report under Settings for average response time trends.
- Trim your sitemap. If it lists thousands of thin, parameter or redirected URLs, you're asking Google to spend effort on junk.
- Link to the page from somewhere strong. A URL that only lives in a sitemap looks low priority.
5. Crawled - Currently Not Indexed
This one hurts, because Google read the page and passed. It's a status, not a penalty. The "currently" is real, and pages do move into the index once the reason changes. But nothing moves until you change something.
The honest questions to ask, in order:
If ten pages say the same thing with a city or product name swapped, Google will index one or none. Merge them or give each real, unique substance.
A page that repeats page-one content without a new angle, example or data is easy for Google to skip.
If the page is five clicks deep with one link from a tag archive, it tells Google the page doesn't matter. Link to it from a relevant, already-indexed page.
Some pages should be pruned, not rescued. The Content Pruning Guide covers how to decide.
If you've got a big template-driven site, this status can show up in bulk. That's a different fight, and the Programmatic SEO Guide covers it.
6. Duplicates and Canonical Fights
When Google sees the same content on several URLs, it picks one to index and folds the rest into it. The trouble starts when your signals disagree. Your canonical says A, your internal links point to B, your sitemap lists C, and Google shrugs and picks its own.
The fix is boring but works: choose one URL, then make everything agree. Self-referencing canonical, internal links to that exact version (same protocol, same www choice, same trailing slash habit), a 301 from the alternates, and only that version in the sitemap.
7. Hard Blocks: noindex, robots.txt and Errors
These are the easy wins, and they're also the ones that cause real disasters. A noindex left over from a staging site can quietly wipe out a launch.
noindex. Also note that AI crawlers have their own rules, covered in the Robots.txt & AI Crawlers Guide.- Check the response header too. An
X-Robots-Tag: noindexheader won't show in your page source. - Soft 404s: a "no results" page returning 200 confuses Google. Return a real 404 or 410 for pages that are gone.
- Redirect chains and loops: keep it to one hop. Link straight to the final URL.
- 5xx errors: even brief outages during a crawl can delay indexing. Validate the fix in Search Console once it's stable.
8. JavaScript and Rendering
If your main content only appears after JavaScript runs, Google has to render the page in a second step, and things can go wrong there: blocked resources, content behind clicks, or links that are click handlers instead of real <a href> tags. Use "View crawled page" in URL Inspection and compare the rendered HTML to what a user sees. If your content or links are missing, that's your answer. Details are in the JavaScript SEO Guide, and headless setups have their own traps in the Headless CMS SEO Guide.
9. Speed, Core Web Vitals and Server Health
Let me be straight: Core Web Vitals don't decide whether a page gets indexed. What they do is expose the same problems that do. A server that takes seconds to respond hurts LCP and also makes Googlebot slow down. Heavy scripts hurt INP and also make rendering expensive.
| Metric | Good target | What a bad score often hints at |
|---|---|---|
| LCP (loading) | 2.5 seconds or less | Slow server response, unoptimised hero image, render-blocking CSS |
| INP (responsiveness) | 200 ms or less | Heavy JavaScript that also slows rendering for crawlers |
| CLS (visual stability) | 0.1 or less | Images and embeds without set dimensions, late-loading ads |
Start with time to first byte and caching, then images and scripts. Full walkthrough in the Site Speed & Core Web Vitals Guide.
10. My Fix Workflow
Download the excluded URLs from the Page indexing report. Split them into pages you want indexed and pages you don't. Ignore the second group.
Fix by cause. One canonical fix can clear hundreds of URLs at once, while fixing them one by one can take weeks.
Noindex, robots rules, 5xx and redirect errors come before any content work. They're fast and the payoff is clear.
Add unique value, merge near-duplicates and add contextual links from strong pages. Internal linking is usually the cheapest lever.
Only indexable, canonical, 200-status URLs. Nothing else.
Use Validate fix in Search Console, request indexing once for key pages, then give it a few weeks before judging.
11. Indexing and AI Search Visibility
This is the new part of the story. Google's AI features draw on indexed pages, so a page that's not in the index can't be cited there. Other AI tools lean partly on other indexes, including Bing's. Submit your sitemap in Bing Webmaster Tools, and look at IndexNow if your site updates often. There's a full section on this in the Bing SEO Guide, and the citation side is in Rank in AI Overviews & LLMs.
Structure helps too. Clear headings, a direct answer near the top and valid structured data make a page easier to extract once it's indexed. It won't rescue a page Google chose to skip, but it helps the ones that make it in.
12. Mistakes I Keep Seeing
| Mistake | Why it hurts | Risk | Fix |
|---|---|---|---|
| Resubmitting the same URLs daily | Doesn't change Google's decision, just wastes your time | MEDIUM | Change the page or its links first |
| Blocking a page in robots.txt to deindex it | Google can't see the noindex, so the URL can linger | MEDIUM | Allow crawling and use noindex |
| Sitemap full of junk URLs | Sends Google to redirected, noindexed or duplicate pages | HIGH | List only canonical, indexable 200 URLs |
| Leaving noindex on after launch | The whole site can quietly vanish from results | HIGH | Check headers and meta after every deploy |
| Publishing hundreds of thin pages | Most get skipped, and they can dilute the good ones | HIGH | Publish fewer, better pages and prune the rest |
13. Indexing Checklist
🔍 Diagnose
- Exact status read from Page indexing report, not guessed
- URL Inspection checked for the live and Google-selected canonical
- Excluded URLs sorted into wanted vs unwanted
🛠️ Technical
- No stray noindex in meta tags or X-Robots-Tag headers
- robots.txt doesn't block important pages or resources
- 5xx errors, soft 404s and redirect chains cleared
- Rendered HTML contains main content and real crawlable links
✍️ Quality and Structure
- Every important page has unique value and 2-3 contextual internal links
- Duplicates merged, redirected or canonicalised, with consistent signals
- Sitemap lists only canonical, indexable, 200-status URLs
⚡ Performance
- Server response time stable in Crawl stats
- LCP 2.5s or less, INP 200ms or less, CLS 0.1 or less on key templates
- Sitemap also submitted to Bing Webmaster Tools
14. Frequently Asked Questions
How long does Google take to index a new page?
There's no fixed number. On a healthy site with decent internal links, new pages often get picked up within days. On a new domain, or for a page Google isn't convinced about, it can take weeks or never happen. If a page has sat for a month with no crawl at all, stop waiting and start diagnosing.
Should I keep clicking Request Indexing in Search Console?
No. It's a nudge, not a fix. Google's own help text for "Crawled - currently not indexed" says there's no need to resubmit the URL. If Google already read the page and passed, asking again changes nothing. Change something about the page or its links first, then request once.
Do Core Web Vitals affect indexing?
Not directly. LCP, INP and CLS are page experience signals, and they matter for user experience and ranking. The indirect link is real though: a slow or error-prone server makes Googlebot back off, which shows up as pages stuck in "Discovered - currently not indexed". Fix server response first, then the Web Vitals.
Is it normal to have unindexed URLs in Search Console?
Yes. Filter URLs, tag archives, feeds, old parameter versions and redirected URLs will always show up as excluded, and that's fine. The goal isn't zero unindexed pages. The goal is that every page you want to earn traffic is indexed.
Can an unindexed page show up in AI answers?
Google's AI features pull from pages that are indexed and eligible to show in Search, so an unindexed page is invisible there. Other AI tools lean partly on other indexes such as Bing's, which is why it's worth checking Bing Webmaster Tools too and not only Google.
15. Sources & References
| Source | What it covers |
|---|---|
| Google Search Console Help: Page indexing report | Official definitions of each indexing status, including the Crawled and Discovered descriptions used here. |
| Google Search Central: Consolidate duplicate URLs | How canonicalisation works and which signals Google weighs. |
| Google Search Central: Robots.txt introduction | Why robots.txt controls crawling, not indexing. |
| Google Search Central: JavaScript SEO basics | How Google renders JavaScript pages and common pitfalls. |
| web.dev: Web Vitals | Official LCP, INP and CLS thresholds. |
| IndexNow | Protocol for notifying participating search engines of URL changes. |
Get more from the Page indexing, Crawl stats and URL Inspection tools used throughout this guide.
Read the guide →Why Googlebot skips URLs and how to steer it toward the pages that matter.
Read the guide →Fix LCP, INP and CLS, and the server response times that also affect crawling.
Read the guide →Put indexing checks inside a complete site audit.
Read the guide →(1) Open Indexing → Pages and export every excluded URL. Sort them into "should be indexed" and "shouldn't".
(2) Group the wanted ones by status and fix the cheapest cause first: noindex, robots.txt, server errors.
(3) For "Crawled - currently not indexed", add unique value and contextual internal links, then wait a few weeks before judging.