SEO Crawling Myths: Why Crawl Budget Isn’t Your Problem
Crawl budget is almost never a small site's problem, and crawling cannot be optimised as a lever: crawl frequency, indexing and ranking all flow from one thing, authority. The second silent killer is cannibalisation, where two pages whose slugs target the same keyword or its synonyms get filed into one index and block each other out of the results entirely. Diagnose at the keyword level, earn authority, and stop chasing the busywork the dashboard rewards you for.
Leave a CommentAbstract
An hour and a quarter dismantling the crawl-budget conversation. Crawling cannot be optimised except through authority; a sitemap is not a to-do list; removing pages does not buy anything; and "crawled, not indexed" means the page was reached and found wanting. What replaces the myth is a much smaller and more uncomfortable diagnosis — the absence of authority — plus a long, detailed account of cannibalisation, which is the problem people actually have.
- Crawling happens at page level, in two separate modes, against a triaged web — not at domain level against a queue you control.
- Crawled-not-indexed and indexed-not-ranking are the same diagnosis in two vocabularies.
- Duplicate content is really about the slug, and duplication only matters when the engine cannot decide which page to show.
- Two parsers sit between the page and the result, which is how both pages can rank and neither can be seen.
- Positions five to seven are the only safe zone to add a specialist page.
- Technical work is restorative rather than additive: fixing what broke restores what you had.
Everything below follows the conversation in order, from the origin of the myth through how crawling actually works, into why the guidance drifted, out through pruning and republishing, and then a long central section on cannibalisation — its mechanism, its arithmetic, its diagnostics and its fixes — closing on a true-or-false recap and the sitemap question. For anyone building a personal brand the useful part is the discipline: every claim here comes with the mechanism that would have to exist for it to be true.
Chapter summaries
00:16 - Crawling cannot be optimised, and the myth comes from an oversimplified picture
You cannot optimise crawling; it is a persistent myth born of an oversimplified early "spider" picture.
"Not really. No. Um, it's just one of these old myths I think that"
— 00:16
The picture everybody carries — a spider walking a site, page by page — was a simplification for an early audience, and an entire practice got built on it. The myth persists because the image is memorable and nothing replaced it.
00:56 - Crawling is not a single event where a sitemap gets fetched end to end
Crawling is not a single process event where you submit a sitemap and it fetches every URL.
"childish video right and from that it spawned all these things right like people"
— 00:56
Submitting a file does not cause every URL in it to be fetched. Treating it as a single transaction is where most of the downstream reasoning goes wrong, because everything after that assumes a process that does not exist.
01:13 - "Fix the sitemap" is advice from a point of observation that does not generalise
"Fix your sitemap" advice comes from a point of observation — big high-authority teams where it genuinely helped — and does not generalise.
"observation you have people Like if if you're going to be a person in"
— 01:13
It worked for large, highly authoritative teams, and the people it worked for wrote it down. Advice inherits the conditions it was formed under, and those conditions almost never travel with it.
01:53 - "Crawled, not indexed" means there was no technical impediment at all
"Crawled, not indexed" means the page was found and retrieved, so there is no technical impediment.
"see people asking questions like I can't get crawled, I can't get indexed or"
— 01:53
The page was found and it was retrieved. Whatever the problem is, it is not access — which eliminates the entire category of fixes people reach for when they see the message.
02:56 - The third problem is the absence of authority
The real third problem is the absence of authority; without it you would need magic.
"problem in all of this is the absence of authority. And I did see"
— 02:56
Not access, not structure, but the absence of authority. Without it, there is nothing for a fix to act on, and every technical remedy is being applied to a system that is working correctly.
03:45 - Practitioners who hated the link game hoped a model would save them
Practitioners who hated the backlink game hoped a language-model search engine would save them; brands lost heavy traffic to scaled AI content a competent SEO could have diagnosed in seconds.
"their savior, right? I've seen, for example, a lot of CMOs go on ads"
— 03:45
A model-driven engine was supposed to make links irrelevant and reward better writing instead. What happened was that brands lost heavy traffic to generated content at scale — a diagnosis any competent practitioner could have made in seconds.
04:17 - A broad and authoritative site can rank almost anything it publishes
A very broad, very authoritative site can rank almost anything it publishes.
"be like that if you're writing for Microsoft, right? I can't imagine you can"
— 04:17
Breadth plus authority means almost anything published ranks, which is the clearest demonstration that the page is not what is being judged. It is also why advice from those sites is unusable everywhere else.
04:56 - Nobody is assigned a spider, and crawlers run in two separate modes
You are not "assigned a spider"; crawlers run in two modes — discovery (find URLs and queue them) and fetching — and a discovered URL lacking topical authority is deprioritised.
"the other thing I think people see is that they get assigned a spider"
— 04:56
Nobody is assigned anything. Crawlers run in two distinct modes — discovery, which finds URLs and queues them, and fetching — and a discovered URL without topical authority behind it simply sits lower in the queue.
05:11 - The web is triaged into pools, and crawling happens at page level
The web is triaged into pools: news and high-authority sites crawled roughly hourly, tier-one about every 12 hours, everything else slower; crawling is at page level, not domain level.
"with caffeine. Caffeine was where Google looked at the problem of how do we"
— 05:11
News and high-authority sites get revisited roughly hourly, the next tier about twice a day, and everything else slower. Crucially the triage happens at page level rather than domain level, which is what makes domain-wide fixes pointless.
05:36 - More crawling does not produce better indexing outcomes
More crawling does not equal better indexing outcomes.
"people think more crawling equals better indexing outcomes and I don't know why you"
— 05:35
The assumption is that frequency is a proxy for favour, and it is not. More visits produce more visits and nothing else, which is why the entire budget conversation leads nowhere. Nothing in the fetch changes how the page is scored once it arrives, so the effort spent courting the crawler is effort spent on a step that was never the bottleneck.
08:42 - Page-level crawling can be verified without taking anyone's word for it
Verify page-level crawling yourself: inspect your top-click pages (recent index dates) against your no-click pages (least frequent).
"right? Don't take my word for it. Log into search console look at your"
— 08:42
Inspect the pages that earn clicks and look at their index dates, then inspect the pages that earn none. The difference is visible in the console in a few minutes. Check it rather than accept it.
09:07 - First-hand guidance beats anonymous forum voices
Prefer first-hand guidance from the engine's own spokesperson over anonymous forum voices.
"authority, it doesn't have to pass the other ones, right? So sometimes you might"
— 09:07
An engine's own spokesperson saying something on the record is a different class of evidence from an anonymous account repeating it. The field routinely treats the two as equivalent.
10:06 - Removing paginated pages does not increase the budget
Removing paginated pages does NOT increase crawl budget; click-earning pages sit permanently in a high-priority queue, and internal linking pulls a new page out of the low-priority queue so it is discovered faster.
"Yeah. Um, and so I if if you're so one of the thoughts"
— 10:06
Pages that earn clicks sit permanently in a high-priority queue, and internal linking is what pulls a new page out of the low-priority one so it gets found sooner. Removing pagination changes neither of those things.
10:57 - Crawled-not-indexed and indexed-not-ranking both mean the same thing
Crawled-not-indexed and indexed-not-ranking both mean a lack of authority; adding pages to a giant index without authority is pointless.
"when you look at the lower pages that just are getting called indexed, that's"
— 10:57
Both messages describe a page that was reached and found wanting, which is the same diagnosis in two vocabularies. Adding pages to an enormous index without authority behind them accomplishes nothing at all.
11:20 - The engine is partly to blame for the confusion
The engine is partly to blame: it pulled the old authority and backlink guidance, and its crawling and sitemap docs say "quality" while the real yardstick for quality is backlinks and internal links.
"I I also think Google are partly to blame for the situation we're in,"
— 11:20
It withdrew the old guidance on authority and links, and its crawling documentation talks about quality while the operative measure of quality remains links — internal and external. The gap between the words and the mechanism is where the confusion lives.
12:56 - Downplaying authority created a vacuum that hype filled
Downplaying authority created a vacuum filled by schema and text-file hype, and made the engine's own guidance look inaccurate.
"they've created a vacuum that has allowed schema and llm.txt TXT and and also"
— 12:56
When the real lever stopped being discussed, the space filled with structured markup and text files. The side effect was that the engine's own guidance started to look inaccurate, because it was describing something other than what decides outcomes.
13:19 - The engine may deliberately prefer that nobody know how to win organically
The engine may deliberately not want you to know how to win organically: its starter guide says do not focus on the experience-and-trust signals and do not obsess over the one thing that is fundamental.
"Maybe maybe Google wants that. Google doesn't like I kind of do you"
— 13:19
Its own starter guide advises not focusing on the trust signals and not obsessing over the thing that is actually fundamental. Whether that is caution or strategy, the effect is the same for anyone following it literally.
14:05 - There is no candid successor to the early spokesperson
There is no candid successor to the early engineer-spokesperson; current spokespeople are handed limits and specific things to attack — "accidentally intentional".
"because Matt Cuts, I've been sharing him so much on this show recently."
— 14:05
The early engineer who answered plainly has no equivalent now. Current spokespeople are given limits and specific things to push back on, which makes the omissions look deliberate even when they are simply constrained.
15:46 - Links may be downplayed because engineered farms are hard to catch
The engine may downplay links because it cannot reliably catch well-engineered link farms except through crude heuristics like unnatural link patterns.
"right? they seem to go all in on telling you how it works and"
— 15:46
Well-engineered link networks are hard to detect except through crude signals like unnatural patterns. If the detection is weak, discouraging the whole category is a cheaper policy than policing it.
16:18 - Many who buy links feel forced into it and do not have to play
Many who buy links do not consider themselves black-hat and feel forced — "go bust or get busted in three years" — but they do not have to play that game.
"Yeah. Yeah. Absolutely. And and I'm I'm sure a lot of people who"
— 16:18
Many of them do not think of themselves as operating outside the rules and describe it as forced — go under, or get caught in three years. The forced choice is usually not forced, and the framing is what keeps people in it.
17:27 - Point of observation applies to the person giving the advice too
Be aware of your own point of observation: the guest concedes he is not in a hyper-competitive space where three sites take 90% of clicks, so he cannot universally say people need not buy links.
"want to correct myself because I don't want to say to other people, be"
— 17:27
He concedes he is not working in a space where three sites take almost every click, so he cannot say universally that nobody needs to buy links. Naming that limit is what makes the rest of the advice usable.
18:01 - Most owners overestimate how competitive their niche is
Most owners are not in truly crazy-competitive niches; they overestimate competitiveness and cannot evaluate it.
"but the thing is but the thing is most most business owners are"
— 18:01
They are not in genuinely brutal niches and they cannot evaluate how competitive theirs is. The overestimate is what justifies reaching for the desperate tactics in the first place.
18:25 - When the diagnosis turns mystical, the diagnosis has gone wrong
When diagnosis drifts to "the page looks less alive", you are in tin-foil-hat territory — stop and seek rational answers; the simplest answer is best.
"It it it's interesting. I found myself in a conversation this week with"
— 18:25
When the explanation arrives at the page looking less alive, the reasoning has left the building. The simplest available answer is almost always the right one, and the mystical answer is a signal to stop and restart.
19:25 - Crawling can be optimised — with authority
You CAN optimise crawling — with authority; reducing pages will not help, the last crawl-throttle control is gone, and links to big sites do not force indexing, but crawl frequency is authority-driven.
"Uh can you can you optimize crawling with authority? >> Absolutely. Yeah, for for"
— 19:25
With authority, crawl frequency changes. Reducing pages will not do it, the old throttle control is gone, and links to large sites do not force indexing — but frequency is authority-driven, so the lever exists and it is not a technical one.
20:26 - Crawlers are the parcel courier of the internet
Stop inventing signals (pillar pages "sending signals"); crawlers are the parcel courier of the internet — they fetch pages and report status, nothing more.
"links in them to each other, and I saw someone sharing a page that"
— 20:25
They fetch pages and report status. That is the whole job. Pillar pages do not send signals, structures do not communicate intent, and the courier does not read the parcel — which removes most of the vocabulary the field uses.
20:58 - A crawled page is not necessarily indexed, and indexing does not repeat
A crawled page is not necessarily indexed, and once indexed it is indexed — not re-indexed to "earn trust"; page-level processing runs in milliseconds and spam or scaled-content checks land days later.
"because a page is crawled doesn't mean it gets indexed. And if it gets"
— 20:58
Being crawled does not imply being indexed, and once indexed a page is not re-indexed to earn trust. Page-level processing runs in milliseconds, and the spam and scaled-content checks arrive days later as a separate pass.
21:38 - The engine is not trying to understand the page
The engine is not "trying to understand" your page — an unambiguous product page (brake pads for a specific car and year) is self-evident; vague pages like "About us" or "Services" waste architecture, being less informative, not confusing.
"to understand you. If your page is Hyundai brake pads for a 1998 Elantra,"
— 21:38
A page about brake pads for a specific car and year is self-evident. Vague pages — about us, services — are not confusing, they are uninformative, and the architecture spent on them is spent on nothing.
22:04 - Technical work is restorative rather than additive
"Technical SEO" has degraded from architecting big sites into fixing every error for a gold star; publishing and technical hygiene are restorative, not additive — fixing what broke restores it, it does not make you exceptional.
"the word architecture and technical SEO, right? Technical SEO to me used to mean"
— 22:04
Fixing what broke restores what you had; it does not make the site exceptional. The discipline drifted from architecting large sites into clearing every warning for its own sake. Restoration is not improvement.
23:47 - The tech-stack advice is third-party opinion sold as fact
"Look after your tech stack" — that the engine likes one website builder and dislikes another — is third-party opinion sold as fact; the engine reads only certain things from the HTML, not the whole document, and does not score your stack.
"Uh this is just nonsense right and and there are if you search for"
— 23:47
The engine reads certain things out of the HTML rather than scoring the document or the platform behind it. The claim that it favours one builder over another is opinion circulated until it sounded like fact.
24:53 - The sitemap is not a to-do list the engine works through
The sitemap is not a to-do list the engine follows A-to-Z; high authority means a submitted page is fetched and indexed, no authority means the engine may ignore your sitemap entirely regardless of dates or "urgent" flags.
"do. Um, but this idea that looking after your tech stack, um, I think"
— 24:50
With high authority a submitted page gets fetched and indexed. Without it, the file may be ignored entirely regardless of dates or priority flags, because the file was never an instruction in the first place.
25:55 - Pruning fixes nothing unless the root cause was the slug
On pruning, fix the root cause: off-centre slugs (a kettle site's post about a company retreat) do nothing and belong off the blog, and pruning neither dilutes nor concentrates authority.
"How do we how how do how do you use this um this"
— 25:55
An off-centre post — a kettle retailer writing about a company retreat — does nothing and belongs off the blog. That is a root-cause fix. Pruning itself neither dilutes nor concentrates anything.
27:10 - The console gold star is not worth chasing
Do not chase a search-console gold star; it reports crawl-access errors for everyone and does not differentiate page types, and persistent errors (cannibalisation, duplicates, tracking parameters) reflect the age and basicness of the systems.
"don't worry about trying to get a gold star in Google Search Console. The"
— 27:10
It reports crawl-access errors for everybody and does not distinguish between page types. Persistent errors around duplication, parameters and cannibalisation reflect how old and basic the reporting is, not how broken the site is.
28:08 - On a new site with no links, none of the crawl tricks apply
Crawled or discovered but not indexed on brand-new no-link sites means off-topic, not linked from an authoritative page, no topical authority; ranking first for a huge head term from a new site will not happen, and no crawl trick or pruning fixes it.
"If it's crawled, not discovered, it could be a rendering issue. It's unlikely. It's"
— 28:08
Off topic, not linked from anywhere authoritative, no topical authority — that is the whole diagnosis on a new site with no links. Ranking first for a large head term is not going to happen, and no crawl trick changes it.
28:56 - Big sites stretch authority rather than prune
For big sites, stretch authority instead of pruning: links decay about 85% per hop, so tiers of pages stop ranking, authority dies and they become end points — you need a marketing brain to land traffic and re-power them, treating the link network as a grid, not every room wired in parallel.
"not going to fix it. For big big websites, it's much more important to"
— 28:56
Authority decays around eighty-five per cent per hop, so tiers of pages stop ranking and become end points. Re-powering them needs traffic landed deliberately — treating the link network as a grid rather than wiring every room in parallel.
30:03 - If nothing is indexing, removing pages will probably not help
If nothing is indexing, removing pages likely will not help (unless volume was a targeting vector); he is sceptical of mass "recoveries" and holds that the only way back for a helpful-content-hit domain is to move domain; better than pruning is organising so authority flows to hub pages and then subpages.
"Yes. I it was a great conversation and um I I saw people"
— 30:03
Unless volume itself was the targeting method, removing pages will not start the indexing. He is sceptical of the mass recovery stories, holds that a badly hit domain's real route back is a new domain, and prefers organising authority into hubs over pruning.
31:07 - Republishing a slug can move a page from the fifth page to the first
Republishing a slug can jump a page from around position 50 to 1; it comes back to basic principles, not "special scenarios" — people are certain they are a 100% unique case and they are not; everyone gets the same rotating crawlers.
"You make such a good point. Thanks for for for bringing that up."
— 31:07
A page can move from around the fiftieth position to the first on a republish. It comes back to ordinary principles rather than special circumstances — everyone is certain their case is unique and everyone gets the same rotating crawlers.
31:50 - Many update victims actually lost backlinks
Many "helpful-content victims" actually suffered authority loss (lost backlinks); fixing broken links will not bounce back high-value short-slug pages — in a lower-authority world you must cornerstone, find your watermark and rebuild upward.
"I think a lot of people that think they got hit by HCU"
— 31:50
A great many of them lost links rather than being judged on content, and repairing broken links will not restore a high-value short-slug page. In a lower-authority world the route back is cornerstoning upward from wherever the watermark now sits.
33:08 - "Write good content" is misleading, because content is subjective
"Write good content" is misleading: content is subjective — there is no forum where writers post work for peer review because rivals tear it apart on ego — so there is no "good content", and the same content on a high-authority site will rank.
"think that's part of the problem with the good content thing, right? and and"
— 33:08
Content is subjective, and the absence of any peer-review forum for it is the evidence — writers will not post work for review because rivals attack it. There is no good content in the operative sense, and the same piece on a higher-authority site ranks.
34:22 - Thin content and information gain are post-hoc rationalisations
"Thin content" and "information gain" are usually post-hoc rationalisations; thin pages rank via authority.
"I was going to say a lot of people also create like um like"
— 34:22
Thin pages rank when there is authority behind them, so both concepts are usually explanations applied after the fact to a result that had a different cause. The rationalisation survives because it is more comfortable than the alternative.
35:18 - Duplicate content is really about the document name
Duplicate content's real driver is the document name — the slug — and the weight the engine puts on it; two pages whose slugs target the same keyword, or synonyms and spelling variants like the two spellings of tyres, will cannibalise.
"Uh what about um how do you think about duplicate content in all"
— 35:18
It is the slug, and the weight placed on it. Two pages whose slugs target the same term — or synonyms, or spelling variants of the same word — will compete, which is a naming problem rather than a writing one.
35:59 - Cannibalisation is real and it is not a penalty
Cannibalisation is real — about 30% of his projects are de-cannibalisation — because the engine is weak at semantics and files a page into an index for an implied word it does not even contain; it is not an algorithm or a penalty.
"The problem with duplicate content is that the two pages block each other after"
— 35:59
Around three in ten of his projects are de-cannibalisation work. The cause is weak semantics: a page gets filed into an index for an implied word it does not contain. It is not an algorithm and it is not a penalty.
36:50 - Duplication only matters when the engine cannot decide
Duplicate content only becomes a problem when the engine cannot decide which page to show and the two block each other; roughly 35% of indexed content is duplicative, the engine "doesn't care", and it is not trying to save money.
"It only becomes a problem when when Google can't tell which page to"
— 36:50
Roughly a third of indexed content is duplicative and the engine does not care. It becomes a problem only when it cannot decide which page to show and the two block each other, which is a different failure from duplication itself.
37:08 - Two parsers sit between the page and the result
The mechanism: two parsers — one tailors results for the engine's requirements, one builds the user's results page — and between them each picks a different one of the two cannibalising pages, so neither reaches the user though both "rank" in search console.
"So that's what I was talking about. Um, if you've got two pages"
— 37:08
One parser tailors results for the engine's own requirements and the other builds the page the user sees. Between them, each picks a different one of the two competing pages — so neither reaches the reader while both appear to rank in the console.
38:06 - The engine does not crawl to save money
The engine does not crawl to save money — an hour of 4K video equals roughly 11 billion HTML documents — so "duplicate or thin content costs the engine money" is wrong; cannibalisation, not duplication, is the issue.
"there's a duplicate content issue. A lot of people think that, especially, funny enough,"
— 38:05
An hour of high-resolution video is equivalent to something like eleven billion HTML documents. Storage is not the constraint, so the argument that duplication costs the engine money collapses. What costs you is the cannibalisation, not the duplication.
39:27 - Adjectives in a slug do not differentiate anything
At slug level, adjectives like best and top do not differentiate ("best car parts" equals "car parts") and a subfolder does not change it; even on a high-authority site two near-identical pages can both be withheld from results.
"Can you explain by at a slug level? Um, so would a subfolder"
— 39:27
Best and top do not differentiate a slug — the phrase with the adjective is the phrase without it — and putting one in a subfolder changes nothing. Even on an authoritative site two near-identical pages can both be withheld.
40:22 - The engine auto-tests between two pages about weekly
The engine auto-tests between two pages in a roughly weekly test-auction, which is why the engineering never resolves; a page with longer click-through history blocks a newer page from ever interrupting it.
"So yeah, the the the the subfolder isn't going to change anything at all."
— 40:22
Roughly weekly, it runs a test auction between two competing pages, which is why the situation never settles on its own. A page with a longer click history blocks a newer one from ever getting a turn.
41:31 - Cannibalisation does not need an identical slug
Cannibalisation needs no identical slug: if the engine treats words as synonymous (data and storage, network and storage) each page enters the other's index, the results parser hides one, its click-through suffers, one drops, the other rises, and a new test restarts the loop.
"It's where the two pages where the second page can come in. And"
— 41:31
Synonymy rather than spelling is what triggers the cycle, so two pages can compete without sharing a word in their URLs. The data-centre naming example is where it shows.
42:55 - A narrow niche written many ways forces a decision about the slug
For a narrow niche written many ways, decide whether the keyword is required or implied in the slug — high topical authority lets it be implied — and check whether you already rank before adding a specialist page; hyper-specialisation is where cannibalisation comes from.
"What would you what would you say to websites that are like, 'Oh my"
— 42:55
Decide whether the keyword needs to be in the slug or can be implied — high topical authority permits implication — and check whether you already rank before building a specialist page. Hyper-specialisation is where the whole problem originates.
43:55 - The assumption of not already ranking is usually wrong
Do not assume you do not already rank; a specialist page for a term you rank for on page three adds a new cannibalisation chance — truly differentiate slug keywords (add a specific product name) and test by searching, and if unrelated pages appear the terms are synonymous.
"Don't assume that you don't already rank for it. Um, and if you're ranking"
— 43:55
A specialist page for a term you already rank for on the third page adds a new competitor rather than a new asset. Differentiate the slug properly, test by searching, and if unrelated pages appear the terms are being read as synonyms.
44:52 - Question pages can cannibalise a heading that already ranks
"People also ask" pages can cannibalise when the same question already ranks as an H2 elsewhere on the site.
"So do you ever find that the uh actually I I I had"
— 44:52
A question already ranking as a heading elsewhere on the site will compete with a dedicated page for the same question. The heading was doing the job, and the new page arrives to take it away.
45:42 - The arithmetic shows why two very different pages tie exactly
The maths: an old page with low relevance (10 of 100) but high authority (1000) yields about 100, while a new fully-relevant page (100%) with low authority (100) also yields 100 — equal, so they cannibalise perfectly.
"So if the H2 on that page ranks and then you put and"
— 45:42
An old page at ten per cent relevance with a thousand in authority scores the same as a new, perfectly relevant page with a hundred. Equal scores is exactly the condition for a permanent stalemate, which is why it never resolves.
47:15 - A specialist page is for developing authority on a term not yet held
Decision rule: a question page or specialist page is for DEVELOPING authority on a term you do not yet rank for; if a rival page is not ranking, remove it before adding yours; if you already rank, do not add it.
"They will absolutely cannibalize each other all day, every day. So the PAAA is"
— 47:15
It is for developing authority on something not yet held. If a competing page is not ranking, remove it before adding yours; if you already rank, do not add it at all. The rule is short and it settles most of the cases.
47:54 - The check before publishing is whether the site already ranks
Before publishing, check whether you already rank; on a large site use a rank-tracking report, because when longtail keywords appear the urge to split into specialist pages is exactly the danger.
"And so, how much how much do you think about when you're putting"
— 47:54
On a large site that means a rank-tracking report rather than intuition. The moment long-tail terms start appearing is exactly when the urge to split into specialist pages arrives, and exactly when it is most dangerous.
48:44 - The fastest diagnostic is a manual removal request
Diagnostic kit: rank-tracking cannibalisation reports are about 50% reliable; when a rank drops, check cannibalisation; the fastest test is a manual removal request — the other page should bounce back within 12 to 24 hours — then re-home the removed page in another index.
"So one, have a SER report. Make sure you don't, you know, sort of"
— 48:44
Cannibalisation reports run around half reliable, so a rank drop is a prompt to check rather than a diagnosis. The fastest test is a manual removal request: the other page should recover within a day, after which the removed one gets re-homed.
49:34 - Positions five to seven are the safe zone to specialise
Striking-distance rule: positions five to seven are a safe zone to specialise, but three to seven is too dangerous — if you rank three to seven, build topical authority on related, non-cannibalising keywords instead of doubling down.
"how about, you know, we were talking about striking distance keywords, keywords where"
— 49:34
Positions five to seven are safe ground for a specialist page. Three to seven is too dangerous, and if you sit in that band the right move is building topical authority on related terms that will not compete.
51:18 - Publishing a dedicated page means removing the old section
When you publish a dedicated page for a term a page ranks for via an H2, remove that section from the old page or turn it into a word and link out — but if the word stays and the old page has click-through history, its mere presence can still cannibalise.
"So you have like uh you have a page about crawl budget and"
— 51:18
Remove the section from the old page, or reduce it to a word and a link. Even then, if the word remains and the old page has click history, its presence alone can keep the competition running.
52:01 - One page can rank for ten thousand keywords until the authority falls
With lots of authority one page can rank for 10,000 keywords, but as authority falls that breadth collapses fast, which is why after a core or spam update or lost backlinks pages built on a now-unsupported targeting method start losing traffic.
"this a very difficult area and and and I think again looking at the"
— 52:01
With enough authority a single page ranks for ten thousand terms, and as authority falls that breadth collapses quickly. It explains why, after an update or a loss of links, pages built on a now-unsupported approach start bleeding traffic.
55:55 - The commonest small-business case is the homepage already ranking for everything
The commonest small-business case: a service homepage already ranking for 125 variations of "expert / consultant / agency in a city" cannot add those pages without ejecting the homepage from the index; a new page is forced by its slug into the same index and struggles unless huge backlinks overwhelm it.
"new new websites who want to do SEO worry about this? Who who are"
— 55:55
A service homepage already ranking for a hundred and twenty-five variations of expert, consultant and agency in a city cannot add those pages without ejecting itself. The new page is forced by its slug into the same index and loses.
57:38 - Fifty town pages can all cannibalise one another
Fifty town pages for one state can all cannibalise if the engine does not treat the town as a differentiator, so doing "the right thing" and adding pages can lock you out; check by searching a town name to see if you already rank under the state-level page.
"think that's where most of the cannibalization happens. like um you are ranking really"
— 57:38
If the town is not being read as a differentiator, fifty town pages compete with each other and doing the apparently right thing locks the site out. Searching one town name and seeing whether the state page already ranks answers it.
58:36 - Ejecting versus keeping is diagnosed at the keyword level
Eject versus keep-and-interlink: if a page that consistently ranked goes intermittent, the other page is also intermittent, and the average position drops, they are blocking each other — diagnose at the keyword level.
"How do you decide when when it's okay to just remove uh >>"
— 58:36
A page that ranked consistently and turns intermittent, while another does the same and the average falls, is two pages blocking each other. The diagnosis lives at the keyword level, not the page level.
59:17 - A real audit cannot be done in an hour
A real audit cannot be done in an hour: you must understand the company's strategy, intent, competition and user mindset, which takes weeks of osmosis, and an AI assistant can hallucinate cannibalisation and cause over-pruning that is reversible.
"That's why I think site audits are so problematic because a lot of"
— 59:17
It takes weeks of absorbing the company's strategy, intent, competition and the mindset of its users. An assistant asked to do it will hallucinate cannibalisation and cause over-pruning — recoverable, but only because the pages can be restored. The engine's own language model cannot disambiguate an acronym with two meanings either, which is why the work targets phrases and patterns rather than meanings.
1:01:00 - Watch the money and lead-generation phrases
Watch your money and lead-generation key phrases; if a vital lead-gen keyword falls after heavy publishing, investigate that page — and note that search console hides pages once they stop ranking, whereas rank-tracking tools keep a historical look-back.
"once websites uh and SEOs and marketers actually go in and do the"
— 1:01:00
If a vital lead-generation term falls after a burst of publishing, that page is the place to look. The console hides pages once they stop ranking, so a tracking tool with history is what preserves the evidence.
1:02:37 - Diagnosis happens at keyword level, over days rather than months
Diagnose at keyword level: filter the keyword in search console and look at pages over the last 24 hours or week; if pages alternate up and down on different days, "taking turns", that is cannibalisation, because a page can rank for its headline keyword yet be hidden for one term.
"look and see uh filter in on the keyword in search console and then"
— 1:02:37
Filter to the keyword in the console and look across the last day or week. Pages alternating up and down on different days — taking turns — is the signature, because a page can rank for its headline term while being hidden for another.
1:03:30 - The failed rocket launch and the oscillating line are different problems
The "failed rocket launch" pattern: a new post that spikes then flatlines lacked authority to cannibalise, while one that oscillates on a broken line is cannibalising; if a duplicate-keyword post has only three clicks, remove it, and a normalising or rising average confirms the fix.
"it where you've let's say you've got multiple teams writing about the same lettuce"
— 1:03:30
A post that spikes and then flatlines never had the authority to compete with anything. One that oscillates on a broken line is competing. If a duplicate-term post has three clicks, remove it, and a rising average confirms the fix.
1:05:15 - The myth recap runs four for four
True/false myth recap: you can optimise crawling — false; fewer pages equals more crawl budget — false; crawling is a single event — false (multi-stage, up to five crawlers per page); more crawling equals better indexing or rankings — false.
"Yeah. Um, I want to do a uh a quick true or false"
— 1:05:15
Crawling cannot be optimised: false, with authority. Fewer pages means more budget: false. Crawling is a single event: false — it is multi-stage, with up to five crawlers touching a page. More crawling means better indexing: false.
1:05:58 - "Changing dates means more crawling" is largely false
"Changing dates equals more crawling" is largely false; sitemaps only work when you are highly authoritative and effectively get a listener bot polling every second, the way a major news publisher's pages get near-instant indexing; no authority means the engine will not check your sitemap for about three weeks.
"Changing dates equals more crawling. >> It it can do. It's largely false,"
— 1:05:58
It can happen, and mostly it does not. Sitemaps work when authority is high enough that a listener effectively polls constantly, the way a major news publisher gets near-instant indexing. Without it, the file may go unread for three weeks.
1:07:02 - At middling authority, editing dates too often costs the trust
At middling authority (the 35 to 75 band), editing last-modified dates too often makes the engine store a checksum, notice only a one-byte change across edits, stop trusting your last-modified, and stretch your index cycles further apart — with no going back.
"you have a mediocre, if you're middle of the road authority where most people"
— 1:07:02
In the middle band, frequent edits to the last-modified date cause the engine to store a checksum, notice that almost nothing changed, stop trusting the field, and lengthen the cycle. There is no route back from that.
1:08:01 - Crawl-to-index is a checklist, and passing any point can trigger it
Crawl-to-index is a checklist — does it have authority, authority for this index, when was it last indexed, what is the file size — and passing any point can send a page straight to indexing; adding 100 lines to a thin one-liner can trigger a reindex, which is the source of the "freshness" and "I added content" myths.
"I was going to say there's a there's a process to understanding like"
— 1:08:01
Does it have authority, authority for this index, when was it last indexed, what size is it — and passing any point can send it straight through. Adding a hundred lines to a one-line page can trigger a reindex, which is where the freshness myth comes from.
1:08:56 - Low-authority sites have to pass more of those checks
Low-authority sites must pass more of these checks; a title-only change on a low-authority page will not reindex, which is one of the few legitimate reasons for a manual crawl, while authoritative sites pass fewer checks.
"you just changed your page title, you have low authority, you're not going to"
— 1:08:56
They have to satisfy more of the list, which is why a title-only change on a low-authority page will not trigger anything. That is one of the few genuinely legitimate reasons to request a crawl by hand.
1:09:17 - Sitemaps help only where there is authority
XML sitemaps help only if you have authority; the engine's own dev guide says a small, fully internally-linked site does not need one, and "discovered, not indexed" means they can already find the page, so a sitemap will not help.
"Last one, which we just said, site maps. Do they solve crawling? >>"
— 1:09:17
The engine's own developer guidance says a small, fully internally linked site does not need one. And discovered-but-not-indexed means the page was already found, so the file cannot help with the thing it is being deployed against.
1:09:44 - The footer sitemap passes authority and solves more than the XML one
HTML sitemaps — unofficially "just a page" — actually pass authority and, placed in the footer, solve more than XML sitemaps despite only incremental authority; he always recommends them.
"HTML sitemaps, which aren't an official thing, right? Because a HTML sitemap is just"
— 1:09:44
It is unofficially just a page, which is precisely why it works: it passes authority. Placed in the footer it solves more than the XML file does, on incremental authority alone, and it is the recommendation he makes every time.
1:10:56 - Nobody should be forced from the homepage to a services page
Do not force people from the homepage to a "services" or "products" page; list services on the homepage and link directly to the most important ones, then fan out — very large software and hardware firms do not use products/services grouping.
"Um, so a lot of a lot of website design still come out"
— 1:10:56
List the services on the homepage and link directly to the most important ones, then fan out from there. The largest software and hardware firms do not group everything behind a products-and-services menu, and the convention is thirty years stale.
Personal Branding Lessons
An episode that replaces a popular explanation with a less comfortable one. The moves below are what survives the demolition.
Verify page-level crawling yourself rather than taking the advice
Log into the console, look at the index dates on the pages that earn clicks, then look at the pages that earn none. The difference between them is visible in a few minutes, and it settles the page-level-versus-domain-level question without anyone's opinion. Checking beats accepting, and here it is cheap. 08:42
Stop inventing signals that nothing in the system produces
Crawlers fetch pages and report status. That is the entire job. Pillar pages do not send signals, structures do not communicate intent, and the courier does not read the parcel. A large part of the field's vocabulary describes mechanisms nobody can point to. 20:26
Check whether you already rank before adding a specialist page
The moment long-tail terms start appearing in the reports is exactly when the urge to split them into dedicated pages arrives, and exactly when it is most dangerous. A specialist page for a term you already rank for on the third page adds a competitor rather than an asset. 47:54
Treat positions five to seven as the only safe zone to specialise
Between three and seven is too dangerous to double down on. If you sit in that band, the productive move is building topical authority on related terms that will not compete — which is slower and does not carry the risk of ejecting the page that was already working. 49:34
Diagnose at the keyword level over days
Filter to the keyword in the console and look across the last day or week rather than the last quarter. Pages alternating up and down on different days — taking turns — is the signature, because a page can rank perfectly well for its headline term while being hidden for another one. 1:02:37
Use a manual removal request as the fastest cannibalisation test
Automated reports run around half reliable, so a rank drop is a prompt to investigate rather than an answer. Removing one page by hand and watching whether the other recovers within a day settles it quickly, after which the removed page gets re-homed somewhere it will not compete. 48:44
Put an HTML sitemap in the footer
It is unofficially just a page, which is exactly why it works: it passes authority, which the XML file does not. Placed in the footer it solves more than the official version does, on incremental authority alone. It is the recommendation made every time, and it costs nothing. 1:09:44
List the services on the homepage and link straight to them
Do not force people through a generic services or products page on the way to the thing they came for. List them on the homepage, link directly to the most important ones, and fan out from there. The largest software and hardware firms abandoned the grouping years ago. 1:10:56
Questions
Each answer ends at the moment in the recording where it is given.
Why is my page "crawled, not indexed", and will a new sitemap fix it?
No, a new sitemap will not fix it. The status itself tells you the page was reached and retrieved, so there is no technical impediment and nothing for a sitemap to solve. Crawled-not-indexed means the page was found and lacks the one thing that would get it kept: authority — relevant inbound links and topical standing. Every reply pointing at the sitemap is chasing a problem the status has already ruled out. 01:53
Can I optimise crawling at all, and if so how?
Not as a standalone lever — there is no dial to turn, no signal to invent, and the last crawl-throttle control is gone. The only thing that genuinely raises crawl frequency is authority, because the most-clicked pages get crawled the most and your refresh rate is a property of the pool your authority puts you in. So you optimise crawling by earning standing and internal links, not by editing sitemaps or pruning pages. 19:25
Does reducing the number of pages on my site increase my crawl budget?
No. Pages that earn clicks sit permanently in a high-priority queue and are always crawled, so deleting other pages frees nothing for them and concentrates no authority. Crawl budget is not the constraint people imagine; a small site is not being rationed. What pulls a new page forward is an internal link from a page that already ranks, which lifts it out of the low-priority queue — discovery, not the deletion of neighbours. 10:06
Why won't my brand-new website rank no matter what I fix?
Because the sites that present with this problem share a profile: brand new, no inbound links, no topical authority, pages not linked from anything authoritative. Ranking first for a competitive head term from a fresh domain will not happen, and no crawl trick or pruning changes it. The missing pieces are relevant links and standing, which behave like leverage a new site simply has not built yet — a gap explored in leverage theory for individuals. 28:08
What is keyword cannibalisation and how do I know if I have it?
It is when two of your pages target the same keyword — or synonyms and spelling variants — and the engine files them into one index where they block each other, so neither reaches the results. It is not a penalty, just a semantic misfile. The tell is alternation: filter one keyword and watch the pages take turns up and down on different days while the average position sinks. Diagnose at the keyword level, never the whole-page level. 35:18
Why do both of my pages disappear from the results when they target the same keyword?
Because of a handoff between two parsers. One prepares the document for the engine's requirements and the other builds the user's results page, and each picks a different one of the two colliding pages — so both appear to rank in your report, yet neither survives to the final results the searcher sees. The report showing both pages ranking is exactly the illusion that hides the problem; the user gets neither. 37:08
Should I create a separate page for a keyword I already rank for on page three?
Be careful — probably not. A specialist page is for developing authority on a term you do not yet rank for. If you already rank, even on page three, a new page only adds another entry to the same index and a fresh chance of cannibalisation. Between positions three and seven especially, the move is to build authority on related, non-cannibalising terms rather than doubling down — a case of not over-polishing the spot you already stand on. 45:42
How do I fix cannibalisation once I find it?
Isolate it, then eject. The fastest confirming test is a manual removal request: pull one page and see whether the other recovers within roughly twelve to twenty-four hours, which proves the block was real. Then re-home the removed page in a different index rather than deleting the work. Removing only a section can be enough, but if the old page keeps the word and has click-through history, its mere presence can keep cannibalising — so ejecting the page is often the only clean remedy. 48:44
Do sitemaps actually help, and should I use XML or HTML?
XML sitemaps help only if you already have authority; a small, fully internally-linked site does not need one, because the engine can already find every page. An HTML sitemap — technically just a page — actually passes authority, and placed in the footer it solves more than an XML file despite only incremental gains. So the practical answer is an HTML sitemap in the footer, and no reliance on XML to rescue a page that is found but not kept. 1:09:17
Sources
SEO Crawling Myths: Why Crawl Budget Isn’t Your Problem
The sections above follow the episode's own order. Each timestamp is the point in the recording where the idea is discussed, so the post reads as a map of the conversation rather than a rearrangement of it.
Quotations come from the episode transcript with spoken filler ("okay", "right", "you know", "um", "actually", "kind of", "sort of") and false starts removed. No words are added: square brackets mark anything inserted for sense, and an ellipsis marks any cut. British spellings are restored, and speech-to-text errors were repaired against context only where the intended word is unambiguous. Where a line appears both in a cold-open teaser and in its place in the conversation, the timestamp cites the second — the point at which it is actually said.
Quotes are stamped with the moment they are said rather than with a speaker's name.