01About Me 02Services 03Expertise 04Pricing 05FAQ 06Contact Us Book a Call Privacy Policy · Terms · Affiliate Disclosure

Keyword Clustering With Embeddings: One Page or Five?

Most keyword cannibalisation is not a mistake anyone made on purpose. It is the result of a keyword list being treated as a content plan.

You export 3,000 keywords. You sort by volume. You start writing. Eighteen months later you have four articles competing for the same intent, none of them ranking, and a client asking why traffic has flattened despite publishing every week.

I have untangled this on more sites than I would like to count, including this one. The fix is not writing more. It is deciding, before you write anything, which keywords are the same page.

The Question a Keyword Tool Cannot Answer

Keyword tools group by string similarity. “best running shoes” and “best running shoe” land together because the characters match. That part is easy and every tool does it.

The hard cases look nothing alike:

  • “why is my site not on google” and “how to get indexed”
  • “crm pricing” and “how much does a crm cost”
  • “seo for tilda” and “tilda site not ranking”

Zero string overlap in some of those. Identical intent in all of them. Write separate pages for each and you have built your own competition.

The reverse trap is just as common. “seo consultant” and “seo consultant salary” share almost every character and are completely different pages: one is a buyer, one is a jobseeker. Merge them and you serve neither.

What Clustering by Embeddings Actually Does

An embedding turns a phrase into a list of numbers representing its meaning. Phrases that mean similar things end up numerically close together, regardless of the words used.

So keyword clustering with embeddings groups by meaning rather than spelling. “why is my site not on google” and “how to get indexed” sit near each other because the model understands they are asking the same thing.

This is the same underlying technology behind how modern search systems match queries to content, which is the useful part: you are grouping your keywords roughly the way the search engine already groups them.

The Method, Start to Finish

  1. Clean the export. Strip branded terms, obvious junk, and anything with no impressions. A dirty list produces dirty clusters and you will not trust the output.
  2. Embed every keyword. Any current embedding model will do. This is a cheap, fast operation even on tens of thousands of rows.
  3. Cluster the vectors. Use a method that does not force you to declare the number of clusters up front, and that is allowed to leave outliers unassigned. You do not know how many topics you have, that is the thing you are trying to find out.
  4. Label each cluster by its highest-volume member. That term is usually the page you should build.
  5. Check the SERP before you commit. This step is not optional, see below.

The Step That Makes It Trustworthy

Embeddings tell you what is semantically similar. They do not tell you what Google treats as the same page. Those are close, but not identical, and the gap is where errors live.

So validate the clusters against reality: take the top two or three keywords in a cluster and compare who actually ranks for each. If the top ten results overlap heavily, Google considers them one page and your cluster is correct. If the results are largely different sites, Google sees two different intents, and you should split the cluster no matter how similar the embeddings claim they are.

Overlap of roughly half the results or more is a reliable signal for merging. Little to no overlap means split. The ambiguous middle is where you use judgement, and where a human who knows the market beats any model.

Signal
Decision
High embedding similarity + heavy SERP overlap
One page, confidently
High embedding similarity + different SERPs
Split, different intent despite similar wording
Low embedding similarity + heavy SERP overlap
One page, the model missed the connection
Low similarity + different SERPs
Separate pages, obviously

Reading the Output as a Site Structure

Once clustered, the shape of your site tends to become obvious rather than debatable.

Large clusters with several distinct sub-groups inside them are hub pages with supporting articles underneath. Medium clusters are single strong pages. Tiny clusters of one or two low-volume terms usually belong as a section inside a bigger page, not as a page of their own.

The cluster sizes also tell you where you are thin. A topic you consider core to the business that produces one small cluster is a topic you have never actually covered, whatever the content calendar says.

Where this connects to the wider structure question, my write-up on topical authority covers how those hubs should link together.

Auditing a Site You Have Already Built

The same technique works backwards, and this is where it earns its keep fastest.

Embed your existing page titles and primary keywords instead of a fresh keyword list, then cluster those. Any cluster containing two or more of your own published URLs is a cannibalisation candidate. You now have a ranked list of internal conflicts, generated in an afternoon, rather than a vague sense that something is competing somewhere.

From there the decision on each pair is the usual one: merge into whichever URL already holds the inbound links and history, redirect the loser, and keep the better body copy. Consolidating is almost always stronger than trying to differentiate two pages that were always about the same thing.

Where This Goes Wrong

Three failures account for most bad clustering.

Trusting the clusters without checking a single SERP. The output looks organised and authoritative, which is exactly what makes an unchecked mistake expensive.

Clustering too aggressively. Push the similarity threshold too far and you end up with one enormous cluster called “SEO” that suggests you write one page about everything.

Ignoring commercial intent. Two keywords can be semantically identical and worth wildly different amounts. “free invoice template” and “invoicing software” are close in meaning and nowhere near each other in value. Cluster by meaning, then prioritise by money.

Is It Worth the Setup?

Under a few hundred keywords, no. Sort them by hand, you will do a better job than any model.

Past a thousand, manual grouping stops being reliable, because you cannot hold that many relationships in your head and you will quietly create duplicates. That is the point where this pays for itself, and it keeps paying every time you plan a new section of the site.

The output is not a content calendar. It is a map of how many pages your market actually supports: which is usually fewer than the keyword export implies, and that is the useful finding.

Not sure whether your pages are competing with each other?

Send me your URL list. I will cluster it and show you exactly which of your own pages are fighting for the same query.

📞 Book a free 20-minute review, or see the keyword research service.

Serving all 50 US states, remote. ✉ info@shazzseo.com

Written by Shahzaib Ul Hassan, senior AI SEO consultant and founder of ShazzSEO. Ranking sites since 2009. 500+ websites optimized, 3,000+ students trained.

Leave a Comment