Free SEO Tool
Keyword Cannibalisation Checker
Drop in an export from Ahrefs, Semrush, Looker Studio or Search Console and this finds every query where more than one of your URLs is ranking. Then it does the part no other free checker does: it separates the conflicts that are actually costing you traffic from the ones that are working exactly as intended, names the page that should own each query, and builds you the redirect map. Nothing is uploaded. The file is read in your browser and never leaves it.
Drop your export here
CSV, TSV or TXT. Or click to choose a file.
Include a header row. Tabs, commas and semicolons all work.
Comma separated. Branded queries are exempted, because Google deliberately shows several pages from one site on a brand search.
Impressions for the query, combined across the competing URLs. Leave at 0 to see everything.
Adjust the thresholds
These are the numbers behind every verdict. They are set where I set them for client work, and you can move any of them. Nothing here is hidden from you, because a verdict you cannot inspect is a verdict you cannot trust.
Two URLs must both rank this high to count as competing.
A gap wider than this means Google has already chosen.
Percent. One page holding this much is not a real split.
Percent. A query worth less than this to every page involved is not worth a merge.
A page ranking for this many queries has a tail worth protecting.
What is keyword cannibalisation?
Keyword cannibalisation is when two or more pages on your own site rank for the same query and neither one wins properly.
The usual story goes that Google gets confused, splits your authority between the pages, and demotes both. Fix it by merging them, the advice says, and your rankings recover.
Most of that story is wrong.
I have read the four articles that own this topic in Google, and I have also gone looking for what Google itself has said about it.
The gap between those two things is the reason this page exists.
So here I will show you what cannibalisation actually is, how to tell a real conflict from healthy co-ranking in your own data, which page should win when there is a genuine fight, and why the standard fix is the wrong one more often than anyone admits.
First, the definition worth using.
The one I work from is not "two pages target the same keyword". It is this: you have a cannibalisation problem when consolidating the pages would earn you more total traffic than leaving them alone.
That is an outcome test, not a symptom test, and it is the only version of the idea that survives contact with real data.
Is keyword cannibalisation actually real?
Here is something that should give the whole industry pause.
Google has never used the term.
I had every published Google SEO Office Hours transcript searched, all fifteen of them, covering November 2022 to August 2024.
The word "cannibalisation" appears zero times, in either spelling.
It does not appear in Google's ranking systems guide, its canonicalisation documentation, its canonicalisation troubleshooting guide, or its spam policies either. For a concept the SEO industry treats as a named technical fault with a standard remedy, that absence is remarkable.
The closest thing to a Google statement is John Mueller on Bluesky in September 2025, and he was distancing himself from the framing rather than endorsing it.
"If you have 3 different pages appearing in the same search result, that doesn't seem problematic to me just because it's 'more than 1'. You need to look at the details, you need to know your site, and your potential users."
John Mueller, Bluesky, 20 September 2025
In the same thread he added a parenthetical that most of the coverage dropped entirely: "IMO it's not really 'cannibalization' if it's theoretical." In other words, two pages that could in principle compete but never actually appear together are not a problem at all.
So where did the "split authority" idea come from?
It comes from a Google blog post in 2008 called "Demystifying the duplicate content penalty", and the industry has been quoting half of it ever since.
Google wrote that when it fails to detect that several URLs are duplicates, "this may dilute the strength of that content's ranking signals by splitting them across multiple URLs."
Read that carefully.
The dilution happens when Google cannot tell that two URLs are the same page. It is a statement about undetected duplicate URLs.
However, it is not a statement about two genuinely different articles on related topics, which is what almost everyone applies it to.
The same post is blunter than anything written since:
"Let's put this to bed once and for all, folks: There's no such thing as a 'duplicate content penalty.'"
Google Search Central, 12 September 2008
And the word "duplicate" appears nowhere in Google's spam policies at all.
So is there nothing to fix?
No, there is something real here. It is just narrower and less dramatic than the folklore.
Two things Google does document are worth knowing.
- Site diversity limits what can show at once. Google's ranking systems guide says its site diversity system "generally won't show more than two web page listings from the same site in our top results", with the caveat that it may show more "in cases where our systems determine it's especially relevant to do so". So if you have five pages chasing one query, at most two of them were ever going to appear. The other three were never in the race.
- Google clusters near-identical pages and picks one. The canonicalisation documentation describes Google finding pages whose primary content is "the same or very similar", clustering them, and choosing the one that is "objectively the most complete and useful for search users". Note the scope: this is about near-duplicates, not about a buying guide and a category page.
That gives you the honest version.
Cannibalisation is real when your pages are similar enough that Google is genuinely picking between them, and when the one it picks is not the one you would have picked.
It is not real just because two URLs share a keyword, and it is not a penalty.
The good news? That distinction is measurable, and it is exactly what this tool measures.
How this checker works
You give it an export containing queries and the URLs that ranked for them. It groups every query, finds the ones with more than one URL, and then scores each conflict against five tests.
Everything happens in your browser.
That is not a marketing line, it is an architectural fact with a consequence you should care about: your Search Console or Ahrefs export contains your commercial performance data, and every other free cannibalisation checker sends it to a server.
This one has no server to send it to.
You can disconnect your internet, load the page from cache, and it will still work.
However you look at it, if you are an agency handling a client's data, or you would simply rather not hand your query data to an SEO SaaS vendor you have never heard of, that difference matters.
The tool reads six export formats and recognises three more that look usable but are not, which I will come back to because it catches a lot of people out.
How the verdict is calculated
Every conflict lands in one of six classes. The thresholds are all visible in the tool and all adjustable, because a score you cannot inspect is a score you cannot argue with.
- Real conflict. Two or more of your URLs rank inside the top 20, the impressions are genuinely split rather than dominated by one page, and the query matters to at least one of the pages involved. This is the pattern worth your time.
- Different page types, check the intent. An editorial page and a commercial page co-ranking. A buying guide and a collection page both showing for "best running shoes" is frequently deliberate and frequently profitable. Do not merge these.
-
Duplicate URL path. The same page reachable two ways, such as Shopify's
/collections/x/products/yroute or a parameter variant. This is a canonicalisation problem and merging content would not touch it. - One page is taking almost everything. There is a split on paper but 85% or more of the impressions sit with one URL. Almost nothing is leaking.
- Google has already chosen. The gap between your best and second URL is more than 30 positions. The trailing page is not competing for anything.
- Marginal to both pages. The contested query is worth under 5% of the impressions on every URL involved, and each of those URLs ranks for at least ten other queries. Merging to win this one term risks the rest of that traffic for very little.
That last test is the one I care most about, and it is the argument Ahrefs has been making in prose since 2021 without anyone turning it into a number.
If a page earns 40,000 impressions across 300 queries and one contested term accounts for 400 of them, redirecting that page away is a terrible trade.
You would win a keyword and lose a library.
The tool also builds a click-through rate curve from your own file rather than importing one.
Published CTR curves are third-party estimates that swing wildly by vertical, and your export already contains the answer for your site: the clicks and impressions at every position you hold.
So when your data includes both, the estimated recovery figure comes from your own measured CTR at the winning position, not from somebody else's benchmark.
Which export do you actually need?
This is where most people fall at the first hurdle, and it is worth being precise about.
You cannot get query-to-URL pairs out of the standard Search Console export.
The Performance report groups by one dimension at a time.
You pick the Queries tab or the Pages tab, and the export gives you whichever one you picked.
Google's own documented method for seeing which pages showed for a query is a three-step drill-down: click the Queries tab, click a query to filter, then look at the Pages tab. One query at a time.
If a combined table existed, Google would not document it that way.
So a Queries export gives you queries and no URLs, and a Pages export gives you URLs and no queries.
Neither can find cannibalisation, and the tool will tell you so rather than silently returning nothing.
That single limitation is why the search results for this topic are full of Looker Studio templates and spreadsheet workarounds.
Here is what does work.
-
Ahrefs, Site Explorer, Organic keywords. The easiest one click option. You need the Keyword, Current URL and Current position columns. Worth knowing: Ahrefs writes files with a
.csvextension that are frequently UTF-16 with tab separators, which is why they break so many tools. This one sniffs the bytes and handles it. - Semrush, Organic Research, Positions. Comma or semicolon separated, both fine. Note the export column names differ from the ones shown in the interface.
- Looker Studio, connected to Search Console. Build a table with Query and Landing Page as dimensions and export it. This gives you real Search Console data with the pairing intact, capped at 50,000 rows per day.
- The Search Analytics for Sheets add-on. Group by Query and Page, then download as CSV. Free, and it does the API call for you.
- The Search Console API or the BigQuery bulk export. The API supports grouping by page and query together, up to 50,000 rows per day per property. The BigQuery bulk export has no documented row limit and is the only complete option.
- Anything else with two useful columns. Any CSV with a query column and a URL column will work, and if the names are unfamiliar the tool will ask you to map them.
One caveat that applies to every Search Console route.
Google anonymises rare queries, defined as those not issued by more than a few dozen users over a two to three month period, and they are omitted from the tables entirely.
Your export is not the whole picture, and neither is anybody else's.
Which page should win the query?
Every other free checker stops at "these pages conflict". That is the easy half.
The hard half is deciding which one survives, and the tool commits to an answer for every real conflict. It ranks the competing URLs on four signals, in this order of weight.
- Current position. The page Google already prefers for this query starts well ahead, because you are working with the algorithm rather than against it.
- Impression share. Which URL is actually being shown, not which one you think should be.
- Overall query coverage. A page ranking for 200 queries is a stronger asset than one ranking for 3, and should rarely be the one you redirect away.
- URL and query match. A slug that matches the query is a targeting signal worth something, though it is the lightest of the four.
You get the winner named, with the reason spelled out, and every losing URL listed underneath it.
Where the fix is a merge, the tool builds the redirect map for you as a two-column list you can copy straight out.
When the wrong page is ranking
This is the situation that sends most people looking for a cannibalisation tool in the first place, and almost nothing written on the topic addresses it.
Your blog post outranks your product page. Google shows the 2019 article instead of the 2026 rewrite. The category page you actually monetise sits at position 14 while a tangential guide sits at 4.
Deleting the winner is not the answer.
Work through these in order instead.
- Check your internal anchor text first. If forty internal links point at the blog post using the query as anchor text, you have told Google which page is about that topic. This is the most common cause and the easiest to fix.
- Check where the backlinks sit. External links to the wrong page will keep it in front. You cannot always fix this, but you should know about it before you spend a week on the page itself.
- Check the title and H1 of the page you want. If the commercial page does not use the query and the blog post does, Google is reading the situation correctly and you are the one who is wrong.
- Check the intent honestly. Look at the actual results page. If the other nine results are all guides, Google has decided this query wants a guide. Your product page is not going to win it, and forcing the issue wastes the effort.
- Then swap the signals, rather than removing the page. Re-point the internal links to the page you want, move the section that matches the query onto it, link from the losing page to the winner with the query as anchor text, and sharpen both titles so the difference is obvious.
Google documents this approach explicitly.
Its canonicalisation troubleshooting guide says to "ensure clustered pages are sufficiently different" and notes that pages "will generally split out faster if the difference between the new content and the other clustered pages is clear and significant."
Differentiation is a first-class fix in Google's own documentation.
Merging is not the only option and it is frequently the worse one.
One more thing from the same doc, because it saves a lot of panic: Google says it "might hold pages in a duplicate cluster for up to two weeks" after you fix the content. Nothing you do here works overnight.
Keyword cannibalisation on Shopify
Every article ranking for this topic treats cannibalisation as a blog problem, and their fixes assume you can merge and delete freely.
On a store you frequently cannot.
I checked the four articles that own the head term.
Ahrefs mentions ecommerce zero times. Seobility mentions it zero times. Yoast gives it one sentence in the whole piece. Semrush uses a t-shirt as a canonical tag illustration.
Not one of them mentions Shopify at all, which is odd given how reliably the platform generates this exact problem.
Here is what actually causes it on a Shopify store.
-
The duplicate product path. Every product is reachable at
/products/handleand at/collections/whatever/products/handle. That is a canonical issue rather than a content one, and no amount of rewriting fixes it. -
Collection against collection. Creating a new collection takes about fifteen seconds, so most stores accumulate
/collections/mens-running-shoes,/collections/running-shoes-menand/collections/mens-trainersover a couple of years. These are the genuine merge candidates. - Collection against product. A one-product collection and its product page will fight, and the collection usually loses. Either broaden the collection or drop it.
- Blog against collection. Your "best running shoes for flat feet" article and your running shoes collection both chasing the same term. This is the one people get wrong most often. It is usually not a fault, it is two intents, and merging them costs you the guide's entire long tail.
-
Tag and filter URLs. Faceted navigation generates near-duplicate indexable pages at scale. Handle these with canonicals and your
robots.txt, not with a merge. -
Pages against collections. A
/pages/landing page built for a campaign, still live two years later, still competing with the collection it was meant to support.
The Shopify fix stack is also different from the one in every article on this topic.
Redirects go in Content, then Menus, then URL Redirects, and there is a bulk CSV import if you have more than a handful.
Canonical handling lives in your theme.liquid, and crawl control goes through robots.txt.liquid, which you can build with my robots.txt generator.
Two Shopify limits are worth knowing before you plan a big consolidation.
A standard store caps at 100,000 URL redirects and Shopify Plus at 20,000,000, which is per Shopify's own documentation.
Shopify's redirect importer also expects its own column headers, so download the sample template from your admin and match it rather than trusting any tool's guess, including mine.
The map this page gives you is a plain from and to list, deliberately, because I would rather hand you something honest than something that fails on import.
If you are consolidating as part of a replatform or a URL change, my guide to Shopify migration SEO covers the redirect side properly, and Shopify site structure covers how to stop generating these conflicts in the first place.
The fixes, ranked
The advice on this topic contradicts itself badly, so here is each remedy weighed against what Google actually documents.
- Merge, then 301 redirect. The strongest fix when it applies. Google lists redirects as the most powerful of its consolidation signals, ahead of canonical annotations and well ahead of sitemap inclusion. Use it when the losing page has little independent value. Update your internal links afterwards so they point at the surviving URL rather than through the redirect.
- Differentiate the pages. Underrated, and the right answer more often than people think. Google documents it directly, and it is the only fix that does not risk the losing page's other rankings. This is the default when both pages earn traffic of their own.
- Canonical tag. Correct for duplicates, wrong for cannibalisation. Google calls it a strong signal but explicitly "a hint, not a rule", and it may pick a different canonical anyway. Use it for the Shopify duplicate product path or a parameter variant. Do not use it to paper over two genuinely different pages, because Ahrefs is right about this one: canonicalisation is not a cannibalisation fix.
- Noindex. Almost never. Semrush lists it as a last resort, Ahrefs calls it terrible, and Ahrefs is closer to right. A noindexed page consolidates nothing. You lose the page and keep none of its value, which is the worst of both options.
- Re-target one page onto a different keyword. Sounds neat, works badly. Seobility recommends it. The problem is that you cannot de-optimise a page for one keyword without affecting the others it ranks for, so you tend to lose more than you intended.
- Move a page to a subdomain. No. Seobility suggests this to get two rankings for one term. Google's site diversity documentation says subdomains are generally treated as part of the root domain for exactly this purpose, so the trick does not work and you have split your site for nothing.
- Delete the page. Rarely, and never without checking. Look at its traffic, its backlinks and its other rankings first. A page that looks stale can still be earning.
What not to do
Three mistakes account for most of the damage I see from cannibalisation cleanups.
Do not merge pages that serve different intents. If one answers a research question and the other sells something, they are not competing, they are covering a funnel. Merging them usually loses the top of it.
Do not fix at the keyword level. Pages rank for hundreds of terms. Optimising a decision around one of them is how people lose a page's entire long tail to win a keyword worth forty clicks a month.
Do not act without a baseline. Record the combined clicks, impressions and total ranking keywords for every affected URL before you touch anything. If you skip this you will never know whether the fix worked, and you will have no way back if it did not.
What this tool will not do
Being straight about the limits, since nobody else in these results is.
- It does not fetch or crawl your site. It reads the file you give it and nothing else. It cannot see your content, your canonical tags, your internal links or your titles, all of which matter to the diagnosis.
- It cannot see queries that are not in your export. Search Console omits anonymised queries. Ahrefs and Semrush show estimated rankings, not your actual impressions. Neither is complete and this tool inherits whatever gaps your file has.
- It cannot detect ranking instability over time. The strongest proof of a real conflict is Google flipping between your URLs for one query week after week. A single export is one moment. Run it monthly and compare if you want that signal.
- It cannot judge search intent. It reads page types from your URL structure, which is a decent proxy on Shopify and a rough one elsewhere. Looking at the actual results page is still your job.
- The click estimate is an estimate. It uses your own measured CTR by position, which is better than an imported curve, but it still assumes a consolidated page holds the better position. Sometimes it will not.
- It will not tell you whether the merged page is any good. Consolidating two thin pages produces one thin page.
Frequently asked questions
Is it cannibalisation or cannibalization?
Both, and Google treats them as the same query. I use the British spelling because I am British. The American spelling carries roughly five times the search volume, so you will see it more often, and one of the tools currently ranking third in Google for the American spelling has the British spelling in its own URL.
Does keyword cannibalisation actually hurt rankings?
Sometimes, and less often than the advice suggests. Google has never documented a cannibalisation penalty and John Mueller has publicly said that several pages appearing for one query does not seem problematic by itself. What is real is that Google clusters near-identical pages and picks one, and its site diversity system generally shows no more than two pages from one site. If your pages are genuinely similar and Google is picking the wrong one, you have something to fix. If they are different pages that happen to share a term, you probably do not.
Why will my Search Console export not work?
Because the Performance report groups by one dimension at a time. The Queries export gives you queries without URLs, the Pages export gives you URLs without queries, and you need both in the same row. Use Looker Studio with Query and Landing Page as dimensions, the Search Analytics for Sheets add-on, the Search Console API, or the BigQuery bulk export. An Ahrefs or Semrush export is the quickest route if you have one.
How many URLs per query is too many?
There is no threshold, which is why this tool does not use one. Two URLs splitting impressions evenly on a commercially important term is worth fixing. Five URLs where one takes 95% of the impressions is not. The number of URLs tells you almost nothing on its own, and any tool that scores on it is guessing.
Should I merge my blog post into my product page?
Usually not. They answer different questions, and the blog post is almost certainly ranking for a long tail of research queries the product page will never touch. Keep both, sharpen the difference between them, and link the post to the product page using the commercial query as anchor text.
How long does a fix take to show up?
Google documents that it might hold pages in a duplicate cluster for up to two weeks after you change the content. In practice I would give a merge and redirect four to eight weeks before judging it, and longer on a large site. Do not panic in week one.
Is my data uploaded anywhere?
No. The file is read by JavaScript in your browser and never transmitted. There is no account, no server and no logging. You can open this page, disconnect from the internet and it will still work, which is the simplest way to prove the point to yourself.
What is the difference between cannibalisation and duplicate content?
Duplicate content is the same content at more than one URL, which is a canonicalisation problem you fix with redirects or canonical tags. Cannibalisation is different content competing for the same query, which is a targeting problem you fix by merging or differentiating. Treating one as the other is the single most common mistake on this topic, and it is why this tool separates duplicate URL paths into their own class.
Can two pages from my site rank for the same keyword on purpose?
Yes, and on brand queries it is normal. Google's site diversity system generally caps a single site at two listings but explicitly allows more where it judges that useful, which is why a brand search often returns your homepage plus sitelinks plus a second page. The tool exempts branded queries for this reason, so add your brand terms before you run it.
Does cannibalisation affect AI search and AI Overviews?
Probably, though nobody has proven it properly yet and I am not going to pretend otherwise. The reasoning is that answer engines cite at passage level and tend to pull a small number of pages per domain, so several near-identical pages give a model no clear best source. I would treat that as a sensible hypothesis rather than an established fact. If AI visibility is the actual concern, my AI visibility audit measures it directly.
How often should I run this?
Quarterly for most sites, monthly if you publish a lot. The useful trick is to keep the exports: comparing two runs three months apart shows you whether Google is flipping between your URLs for a query, and that oscillation is the strongest evidence of a real conflict there is.
What size site does this handle?
It processes whatever your browser can hold, and tens of thousands of rows are fine. There is no keyword cap, no daily limit and no paywall, which is the main practical difference between this and the free checkers that stop at ten keywords.
Found conflicts you are not sure how to unpick?
Cannibalisation is one of those problems where the diagnosis is the easy part and the decision is the hard one. Merging the wrong two pages can cost a site more traffic than the conflict ever did, and I have been called in to reverse exactly that more than once.
I run this check on every technical SEO audit, alongside the crawl and indexing work that usually explains why the wrong page is winning in the first place. If you want someone to look at what your export is actually telling you, get in touch or email me directly at graeme@gwcontent.com.
More free tools are on the SEO tools hub, including the hreflang generator if your conflicts turn out to be international rather than internal.