🏆 US-Registered Digital Marketing Agency Trusted by 200+ brands · USA · UK · Canada · AUS
Advertisement
Advertisement
KEYWORD RESEARCH

Keyword Cluster Generator — group keywords into pages, not lists

Paste a keyword list and get topic clusters with a suggested pillar page for each one.

Volume is optional but improves the pillar suggestion — the highest-volume term in a cluster becomes the pillar.
Raise the threshold for tighter, more literal clusters. Lower it to pull loosely related terms together.
Clusters found
0
 
0
Keywords processed
0
Avg keywords per cluster
0
Single-keyword clusters
0
Total volume clustered
Tip: a cluster is a page, not a folder. If two keywords land in the same cluster but need completely different answers, split them by hand — the tool reads words, you read intent.
Advertisement

A keyword export is a list. A content plan is a set of pages. The work in between — deciding which of those 600 phrases belong on the same page and which deserve their own — is where most content strategies stall, and it is the job this keyword cluster generator does in a few seconds. Paste your keywords, and it groups them by shared vocabulary, ranks the groups by search volume, and nominates a pillar keyword for each.

Arb Digital's content team clusters before writing a single brief, because publishing one page per keyword is the fastest known route to a site that competes with itself. Nine articles about nearly the same thing will always lose to one page that covers the topic properly.

What This Keyword Cluster Generator Does

Every keyword is broken into tokens, common stop words are removed, and simple plurals are folded together so that "shoe" and "shoes" count as the same token. The tool then measures how much vocabulary each pair of keywords shares and groups anything above your similarity threshold into a cluster.

Clusters are seeded from the highest-volume keyword downward, so the biggest term in each group becomes its pillar — the page the cluster should be built around. Everything else in the cluster is a supporting term for that page, or a candidate for a sub-page linked from it. The output panel lists each cluster with its pillar, its total combined volume, and every member keyword, ready to paste into a content plan.

Processing happens entirely in your browser. Nothing is uploaded, so you can safely paste a client's full keyword export or an unreleased product roadmap.

How to Use It

  1. Paste your keyword list. One keyword per line. Add a comma and the monthly volume if you have it — the pillar suggestion and the cluster ranking both improve considerably with volume data.
  2. Set the similarity threshold. Start at 40%. If you get one giant cluster, raise it. If you get dozens of clusters of one, lower it.
  3. Set a minimum cluster size. Setting this to 2 hides orphan keywords so you can focus on the topics with real depth behind them.
  4. Read the pillar for each cluster. That is your page. The remaining keywords in the cluster are the headings, sections, and questions that page needs to cover.
  5. Copy the output into your content plan, then apply judgement — merge clusters that clearly answer the same question and split any that mix incompatible intents.

The Formula: How Clusters Are Formed

Similarity between two keywords is measured with the Jaccard index — the size of the shared token set divided by the size of the combined token set:

Similarity = (Tokens in both) ÷ (Tokens in either)

Take "trail running shoes" and "best trail running shoes". After stop-word removal, the token sets are {trail, running, shoe} and {best, trail, running, shoe}. Three tokens appear in both, four appear in total, so similarity is 3 ÷ 4 = 75% — comfortably clustered at any sensible threshold. Now compare "running shoes" with "marathon shoes": {running, shoe} against {marathon, shoe} gives 1 ÷ 3 = 33%, below the 40% default, so they stay in separate clusters.

Clustering is greedy rather than exhaustive. The highest-volume unassigned keyword becomes a seed, everything above threshold joins it, and the process repeats with the next unassigned keyword. This is fast, deterministic, and easy to reason about — run it twice with the same settings and you get the same clusters, which matters when the output goes into a plan someone else has to follow.

Advertisement

Token Overlap Is a Proxy for Intent, Not Intent Itself

The limitation is worth stating plainly. Two keywords that share words usually share meaning, but not always, and two keywords with no shared words can mean exactly the same thing. "Cheap flights" and "budget airfare" are the same query to a searcher and share zero tokens. No purely lexical method catches that, and no tool that runs offline in a browser will.

This is a well-known limit of bag-of-words matching, and the standard treatment of tokenization and stemming in the Stanford information retrieval textbook explains why aggressive normalisation helps in some cases and destroys meaning in others. The practical response is to run the clustering, then scan for synonym pairs across clusters and merge them manually. That takes minutes on a list of a few hundred keywords and produces a materially better plan.

The opposite error is more common though: keywords that share words but need different pages. "Running shoes" and "running shoes for flat feet" overlap heavily, yet the second has a specific problem behind it that deserves its own treatment. Overlap tells you where to look; only reading the query tells you what to do.

Choosing the Right Threshold

Threshold is the only setting that meaningfully changes the output, and the right value depends on how long your keywords are. Short two-word phrases share tokens easily, so a low threshold collapses everything into one blob. Long-tail question keywords contain many tokens, so a high threshold isolates almost every line into its own cluster.

A workable routine: run at 40%, look at the largest cluster, and ask whether one page could genuinely satisfy every keyword in it. If yes, keep going. If the cluster contains obviously different jobs, raise the threshold by ten and re-run. Two or three passes usually settle it. Aim for clusters of roughly five to twenty-five keywords — small enough to be one page, large enough to justify a substantial one.

What to Actually Build From a Cluster

Each cluster maps to one primary page. The pillar keyword becomes the page's target term and shapes the title and H1. The remaining keywords become sections, subheadings, and FAQ entries within that page, in descending order of volume. A cluster of eighteen keywords does not become eighteen articles; it becomes one thorough article with roughly eighteen things it definitively covers.

Where a member keyword clearly needs more depth than a section allows — a sizing guide, a comparison table, a technical explainer — promote it to a supporting page and link it from the pillar, with a link back. That hub-and-spoke shape is the whole point of clustering: concentrated topical coverage with internal links carrying context between related pages. Once you have decided which pages exist, the internal link opportunity calculator will tell you how many links each of them still needs.

Length follows from the cluster, not from a word-count target. A cluster with twenty distinct sub-questions needs a long page; a cluster with four does not. Checking the result against competing pages with the SEO content length checker is a better guide than any universal rule.

Splitting a Cluster on Intent

Some clusters must be split even when the similarity score says otherwise, and the signal is usually a modifier. "Buy", "price", and "near me" indicate someone ready to transact. "How to", "why", and "what is" indicate someone learning. "Best", "vs", and "review" indicate someone comparing. A cluster mixing all three is three pages wearing a trench coat.

Google's guidance on creating helpful, people-first content is essentially an argument for this discipline: a page should have a clear purpose and satisfy the person who arrived. A page attempting to be a buying guide, a tutorial, and a product listing simultaneously satisfies nobody fully. When in doubt, check the live results for the pillar keyword — if the top ten are all one format, that format is the answer.

Have the keywords but no plan to turn them into pages?

Arb Digital's content team builds clustered content roadmaps, writes the pillar pages, and structures the internal linking that makes a topic cluster rank as a group rather than a collection of individual posts.

Content Marketing SEO Services

Common Mistakes to Avoid

  • Writing one page per keyword — the single most reliable way to create pages that compete with each other for the same query.
  • Trusting the clusters without reading them — token overlap is a starting point, and synonym pairs will always need a manual merge.
  • Mixing transactional and informational terms in one page because they happened to share vocabulary.
  • Choosing the pillar by volume alone — sometimes the second-largest term is the one that describes the page a reader actually wants.
  • Treating cluster size as a word-count target — coverage should be driven by the questions, not by a number.

Related Free Tools From Arb Digital

Expand a thin cluster with the LSI keyword generator, judge whether the pillar term is winnable using the keyword difficulty checker, and check that no existing page already targets it with the keyword cannibalization checker. Estimate the payoff with the SEO traffic forecast calculator, then title the pillar page with the blog title generator. Everything else lives in the free online tools hub.

Frequently Asked Questions

What is keyword clustering?

It is the process of grouping keywords that a single page could rank for, so you plan one strong page per group instead of one thin page per keyword. Clusters map to pages; individual keywords map to sections within them.

How does this tool decide which keywords belong together?

It compares the words in each keyword after removing stop words and folding simple plurals, then groups any pair sharing more vocabulary than your threshold. The measure used is the shared token count divided by the combined token count.

What similarity threshold should I use?

Forty per cent is a sensible starting point. Raise it if you end up with one oversized cluster, lower it if almost every keyword sits alone. Two or three passes usually produce a workable set.

Will it group synonyms that share no words?

No. Purely lexical clustering cannot connect phrases like cheap flights and budget airfare, because they have no tokens in common. Scan the finished clusters for synonym pairs and merge those by hand.

How many keywords should one cluster contain?

Roughly five to twenty-five works well for a single substantial page. Much larger usually means the cluster is really two topics; much smaller often means the term belongs as a section inside a bigger page.

How is the pillar keyword chosen?

The highest-volume term in the cluster becomes the pillar, because it is normally the broadest phrasing of the topic. If no volume is supplied, the first keyword encountered in the cluster is used instead.

Is my keyword list sent anywhere?

No. All clustering runs in your browser with JavaScript. Nothing is uploaded, logged, or stored, so client lists and unpublished plans stay on your machine.

Clusters produced here are a planning aid based on word overlap alone — always confirm groupings against live search results before committing them to a content plan.

Advertisement
Advertisement
Arb Digital assistant

👋 Hey! Want to grow your business? Ask me anything — a free marketing proposal is on the table!