Content Gap Finder
We crawl your competitor's pages and yours, read what each page says it is about, and show you the topics they publish that you never cover. Based entirely on a live crawl — no invented keywords, and no claim about what anyone ranks for.
🔒 No signup · Read-only crawl · Usually done in 20–40 seconds
What This Tool Honestly Shows You
Most tools that use the phrase "content gap" mean something specific and expensive: the keywords a competitor ranks for that you do not, pulled from a licensed SERP database costing thousands a year. This tool does not have that data and will never pretend it does. Instead it measures something adjacent, free, and entirely verifiable — what your competitor has actually chosen to publish about.
It works by crawling a bounded sample of their pages and reading each one's title, H1, subheadings and meta description. Those are the fields where a page states its own subject, and they are written by the competitor's own team. Do the same on your domain, subtract your topic vocabulary from theirs, and what remains is the set of subjects they cover and you do not. That is the entire method, and every phrase in your report can be traced back to a specific page on their site — the report links the evidence.
The distinction matters commercially. Publishing intent is a leading indicator; rankings are a lagging one. A competitor who has just built out fifteen pages on a topic you ignore is telling you where they intend to compete, whether or not those pages have earned positions yet. What this cannot tell you is whether that bet is working for them — for that you would need ranking data, and you should be sceptical of any free tool that claims to have it.
How the Gap Is Calculated
Discover their pages
We read robots.txt, follow any Sitemap: directive it declares, and fall back to /sitemap.xml and then homepage links. URLs are sampled across the sitemap, not taken from the top.
Crawl politely, within limits
Up to 30 pages per domain, fetched in small concurrent batches with short timeouts and an overall time budget. Disallowed paths are dropped before any request is made.
Extract what each page is about
Title, H1, H2 and H3 headings and meta description are normalised into one- to three-word phrases, with function words and brand names stripped out.
Diff, then cluster
A phrase counts as a gap only if it appears on several of their pages and appears nowhere in the pages we crawled on yours. AI then groups the survivors into themes — using only phrases the crawl found.
Because the sample is bounded, treat the result as a strong signal rather than a complete inventory. Thirty pages is enough to reveal what a site is built around; it is not their whole archive. If a competitor has two thousand posts, run the report, act on the clearest themes, and re-run it in a few months to see what they have added.
The natural next steps sit in the other two tools: check the real search demand behind each gap topic with the keyword research tool, and see how the two sites compare technically with the competitor comparison tool. If you would rather hand the whole programme over, that is what our content marketing service is for.
What It Does and Doesn't Claim
Topics they publish
Taken directly from their own titles and headings, with the source page linked so you can check any term yourself.
Topics you don't
Absent from every title, heading and meta description across the pages crawled on your domain, checked as a phrase and as a substring.
Shared ground
The themes you both cover, which tells you how genuinely comparable the two sites are before you act on any gap.
Not rankings
No SERP position data exists here. Nothing in the report says or implies that a competitor ranks for anything.
Not traffic or volume
No search volume, traffic estimate or difficulty score is shown, because none was measured. The keyword tool handles live demand data.
Not a full archive
A bounded sample per domain, stated on every report along with exactly how many pages were analysed and how they were found.
Content Gap Questions
What exactly does this content gap finder measure?
It crawls a bounded set of pages on the competitor domain you enter — found through their sitemap, or through links on their homepage when no sitemap exists — and does the same on your domain. From each page it reads the title, the H1, the H2 and H3 headings and the meta description, because those are where a page declares what it is about. It then builds a vocabulary of the topics each site publishes and subtracts yours from theirs. The output is the phrases that appear across the competitor's pages and nowhere in the pages we crawled on yours.
Does this show me the keywords my competitor ranks for?
No, and it deliberately does not claim to. Knowing what a domain ranks for requires SERP-tracking or clickstream data from a paid provider, and nothing here has that. What this tool observes is what a competitor chose to publish, which is a different and often more actionable thing — publishing intent is visible in their headings, whether or not those pages currently rank. Treat the output as "topics they cover and you do not", because that is precisely what was measured.
How many pages does it crawl?
Up to 30 per domain by default, with a hard ceiling of 40, and the whole job runs against a wall-clock budget. Pages are discovered from the sitemap where one is available, sampled across the file rather than taken from the top, so a news-ordered sitemap does not give you thirty pages from the same week. The crawl is deliberately small: it is polite to the site being crawled and it keeps the report fast.
Does it respect robots.txt?
Yes. robots.txt is the first thing requested on each domain, and any URL its Disallow rules cover is dropped before it is ever fetched. The crawler identifies itself with a descriptive user agent that names the tool and links back to this site, so any administrator reading their logs can see exactly what visited and why.
Why do some of the results look like odd fragments?
The vocabulary is built from real phrases in real headings, so it inherits whatever the competitor actually wrote — including navigation labels and product names. Nothing is cleaned up by inventing better-looking terms, because that would mean showing you something the crawl did not find. Read the list as raw evidence and let the clustered view do the interpreting.
Is the AI adding keywords of its own?
It cannot. The clustering step is only allowed to return phrases copied character for character from the crawled list, and anything it returns that is not in that list is discarded before you see it. It groups and explains; it never contributes a topic.
What should I do with the gaps it finds?
Not all of them deserve a page. Work down the list and keep the topics that a customer of yours would plausibly search for and that you could genuinely write something useful about. Discard the ones that are competitor-specific product names or industry jargon nobody searches. Then take the survivors into the keyword research tool to see the real autocomplete demand around each one before committing to a brief.
Can I run this against my own two sites, or a site I do not own?
You can crawl any publicly accessible site — the requests are the same read-only GETs a browser makes. Private, local and internal addresses are refused. Comparing two of your own properties is a legitimate use and is a quick way to find topics one site covers that a sister site has never touched.
Find Out What They're Publishing and You're Not
Two domains, one crawl, an evidence-linked list of the topics your competitor covers that never appear on your site.
Run a Free Content Gap ReportWant a content plan built from it? Talk to our team about content marketing.