Crawl Budget: Stop Google Wasting Time on Junk
Google gives your site limited attention. If it spends that on junk URLs, it never reaches the pages you care about.
Why new pages take months to appear
Google works through what it knows about. If most of that is junk, your real pages wait behind it.
-
URLs multiply
Filters, parameters, archives and deleted pages that never fully went away.
-
Google works through them
It knows about five times what you published, and spends its visits on addresses that lead nowhere.
-
The waste gets cut
Dead URLs returned cleanly, chains flattened, parameters handled, sitemaps trimmed to live pages only.
Google gives your site a limited amount of attention. If it spends that on pages you do not care about, it never reaches the ones you do.
That is crawl budget. For most small sites it does not matter at all. For sites with thousands of URLs, most of them junk, it is often the reason new pages take months to appear.
The Honest Version: Most Sites Should Ignore This
Under a few thousand pages, it is not your problem
Google has said this plainly and it is true. If you have 200 pages, crawl budget is not why you are not ranking. Something else is, and I will tell you what rather than sell you this.
When it genuinely bites
Large sites. Ecommerce with filters generating endless combinations. Sites carrying thousands of deleted or redirected URLs. Anywhere the number of addresses Google knows about is far larger than the number of pages you actually have.
The signal that matters
Compare how many URLs Google knows about with how many it has kept. If it knows about five times what you published, it is spending most of its time on things that do not exist or do not matter.
What Eats the Budget
Almost always addresses nobody meant to create.
Deleted pages still being requested
Google keeps checking URLs for years after you removed them, especially if something still redirects to them.
Redirect chains and loops
Every hop is a request. A chain of three costs three times what it should.
Filter and parameter URLs
Colour, size, sort order and page number, multiplying into thousands of near-identical addresses.
Empty archives
Tag and date pages with nothing on them, each one still crawled.
Very slow pages
A slow server means fewer pages crawled per visit. Speed and crawl budget are connected.
Enormous sitemaps of dead URLs
Actively inviting Google to spend time on pages that no longer exist.
How I Work On It
1. Compare known against kept
This single ratio tells you whether you have a problem worth paying to fix. If Google knows about roughly what you published, stop here.
2. Find what it is actually requesting
Server logs where available, Search Console where not. This usually surprises people, because the answer is rarely the pages they were worried about.
3. Cut the waste at the source
Dead URLs returning a clean gone response rather than redirecting into other dead ends. Parameters handled. Empty archives removed. Chains flattened.
4. Point what is left at the right pages
Sitemaps containing only live, indexable URLs, and internal links that lead somewhere.
A Real Example
This site had 642 deleted pages that were still being requested, and 515 redirect rules pointing at them.
So every one of those addresses cost two requests to arrive at nothing. Google was spending a large share of its attention on a loop of dead ends, while 584 real URLs sat unindexed.
The redirects were switched off, the dead addresses were set to return a clean gone response so Google would drop them properly rather than keep checking, and the sitemap shrank as the junk left it.
When Not To Buy This
You have a few hundred pages. Genuinely not your problem. Look at content and internal linking instead.
Your new pages get indexed within days. Then Google is reaching you fine.
You want it done because a tool flagged it. Tools flag crawl budget on sites that will never be affected by it. I will tell you honestly if yours is one.
What You Can Check Yourself First
One number tells you whether this applies to you at all.
1. Compare known against published
In Search Console, the Pages report shows indexed and not-indexed totals. Add them together for roughly what Google knows about, then compare with how many pages you have actually published.
If Google knows about roughly what you published, stop here. Crawl budget is not your problem and nobody should sell it to you.
If it knows about several times more, something is generating addresses you did not intend.
2. Look at the not-indexed reasons
Google groups them. Large counts under alternate pages, redirects or not-found are the signature of wasted crawl, and each group tells you something different about where it is coming from.
3. Check your sitemap count
Open your sitemap and count the URLs. If it lists far more than you have real pages, or includes redirected and dead addresses, you are actively inviting Google to waste its time.
4. Try adding a nonsense parameter to a URL
Load one of your pages and add something like ?test=123 to the end. If it loads normally as a separate page rather than redirecting, your site can generate unlimited addresses, and crawlers will find them.
Price and Turnaround
Scope |
Price |
Turnaround |
|---|---|---|
Under 1,000 URLs |
$349 |
4 – 6 days |
1,000 – 10,000 URLs |
$699 |
1 – 2 weeks |
Over 10,000 URLs |
From $1,199 |
2 – 4 weeks |
I will tell you before quoting whether your site is big enough for this to matter. Often the honest answer is no.
Common Questions
Should I remove old pages to save crawl budget?
Only if they are genuinely worthless. A page earning nothing but costing one crawl a month is not your bottleneck.
Will blocking pages in robots.txt save budget?
It stops them being crawled, which does save budget. It does not remove anything already in the index, and it stops Google seeing a noindex tag if one is there.
Does site speed affect crawling?
Yes. A faster server means more pages fetched per visit. It is one of the few places where speed and crawling are directly connected.
How do I know if I have a crawl budget problem?
Compare the URLs Google knows about with the ones it kept. If the first number is several times the second, it is worth looking.
Does blocking pages in robots.txt help?
Sometimes, and it is easy to get wrong. Blocking a page stops it being crawled but does not remove it from the index if it is already there.
Should I delete old pages?
Better to make them return a clean gone response than to redirect them somewhere irrelevant. Redirecting into a dead end is the worst option.
How long before it improves?
Crawl patterns shift over weeks, not days. Expect four to eight weeks before the index reflects it.
Is this the same as indexing problems?
Related but different. Wasted crawl is one cause of pages not being indexed, and often the one nobody checks.
How many URLs does Google know about?
Send me that number and how many pages you actually have. I will tell you free whether this is worth fixing.
Or email info@shazzseo.com