CMS Canvas
CMS CanvasSEO Crawler

The Free SEO + AI SEO Crawler for Agencies

Find every SEO problem on any website — and make sure AI can find you too.

CMS Canvas scans a site end to end, scores its SEO health, and runs a full AI SEO audit — so you rank on Google and get cited by ChatGPT, Perplexity, and Google AI Overviews.

Scan, understand, and fix SEO issues with AI.

AI SEO / AI Readiness Audit

Checks whether ChatGPT, Claude, Perplexity, and Google AI Overviews can actually read, understand, and cite your site — AI crawler access, llms.txt, answer-ready content, and trust signals.

Deep Site Crawl

Scans every page of a website the way Google does — broken links, redirect chains, missing titles, duplicate content, sitemap gaps, and 30+ other checks.

SEO Health Score & Core Web Vitals

One clear 0–100 score plus Google's own page-speed grades, so clients instantly understand where they stand.

AI Plain-English Overview

AI turns raw findings into a prioritized action plan written in business language — what to fix first and why it matters.

One-Click Pitch Emails

Generate a ready-to-send outreach email to the site owner, citing their real issues, with a branded PDF report to attach.

Validate Fixes

Re-scan after the work is done and show clients a before/after comparison proving exactly what was fixed.

Reports & Quotes

Downloadable PDF reports and fix-it quotes, ready to send to clients.

AI SEO — the new search battleground

Is your site invisible to ChatGPT, Perplexity, and AI Overviews?

More and more buyers ask AI instead of Google. If AI assistants can't crawl, understand, or trust your site, you don't exist in those answers. Every CMS Canvas scan includes a full AI Readiness audit:

AI Crawler Access

Detects when robots.txt accidentally blocks GPTBot, ClaudeBot, PerplexityBot, Google-Extended, and other AI crawlers — the #1 reason sites never get mentioned by AI assistants.

llms.txt Check

Checks for the emerging llms.txt standard — a plain-text map that tells LLMs exactly which of your pages matter most, so AI tools cite the right content.

Answer-Ready Content

Flags pages AI can't quote — missing question headings, lists, tables, and FAQ schema — and shows exactly how to restructure them to win AI Overview citations.

AI Trust Signals

Verifies the Organization and Person schema that tells AI systems who is behind the site — the credibility markup AI weighs when choosing which sources to recommend.

Plus JavaScript-rendering warnings (most AI crawlers can't run JavaScript at all) and structured-data validation on every page.

Why choose CMS Canvas over Screaming Frog, Ahrefs, Semrush, or Sitebulb?

Those are all respected tools — but they're built for full-time SEO specialists. CMS Canvas is built for everyone who isn't one. It finds the problems, explains what each one means in plain language, and then walks you through fixing every issue. So whether you're an agency proving value to a client or a business owner improving your own site, you can get real results — no SEO experience, and no specialist on payroll.

100% cloud-based

No desktop install, no crawls hogging your machine.

Free to start

Full crawls on the free tier — no all-in-one suite subscription.

AI plain-English reports

Findings become a prioritized action plan clients understand.

Built to win clients

One-click pitch emails, branded PDFs, quotes, and fix validation.

CMS Canvas vs. Screaming Frog

A powerful desktop crawler loved by technical SEOs — but it runs on your own machine, output is raw spreadsheet-style data, and turning a crawl into something a client understands is on you.

CMS Canvas runs entirely in the cloud — nothing to install, no laptop tied up mid-crawl — and every scan comes with an AI plain-English overview and a branded PDF report you can hand straight to a client.

CMS Canvas vs. Ahrefs Site Audit

A solid site audit, but it comes bundled inside a large all-in-one SEO platform built around backlinks and keyword research — you pay for the whole suite to get the audit.

If site audits are what you need, CMS Canvas gives you free crawls with an SEO health score, Core Web Vitals, and client-ready reports — no all-in-one subscription required.

CMS Canvas vs. Semrush Site Audit

Another strong audit inside a sprawling marketing platform. Reports are built for in-house marketers reviewing dashboards, not for agencies pitching new clients.

CMS Canvas is built for the agency workflow: scan a prospect's site, get a prioritized action plan in business language, then send a one-click pitch email citing their real issues — with a fix-it quote attached.

CMS Canvas vs. Sitebulb

A well-regarded desktop (and cloud) auditing tool with rich visualizations, aimed at consultants doing deep technical analysis.

CMS Canvas skips the install and the learning curve: free cloud crawls, a single 0–100 health score clients instantly grasp, and before/after re-scans that prove the value of your work.

Transparent pricing — no hidden fees

Every plan includes every feature. You only pay for how many pages you crawl — nothing else.

Free

€0forever

Try it on any site — no credit card, no trial clock.

  • Crawl up to 75 pages per scan
  • Full SEO health score & Core Web Vitals
  • AI plain-English action plan
  • All 25+ technical checks included
Most popular

Pro

€115per month

For agencies auditing full client sites every month.

  • 1,000 crawled pages per month
  • Full-site crawls, no per-scan cap
  • PDF reports, quotes & pitch emails
  • Re-scan to validate fixes
  • Cancel anytime — no contracts

Top-Up

€115one-time

Ran out mid-month? Add pages without changing your plan.

  • 1,000 extra crawled pages
  • Pages never expire mid-project
  • No subscription change required
  • Pay only when you need it

Prices include everything — no setup fees, no per-user charges, no feature paywalls. Upgrade, downgrade, or cancel anytime.

SEO Crawling FAQs

Everything you need to know about how search engines crawl, index, and rank websites.

Definitions & basics

Crawling is the process where search engines send automated programs (crawlers) to discover and download pages on the web by following links and sitemaps. It's the first step in getting a page to appear in search results — before a page can be indexed or ranked, it has to be crawled.

A web crawler — also called a spider or bot — is software that automatically browses the web, following links from page to page and downloading what it finds. "Crawler," "spider," and "bot" are interchangeable terms. Search engines run crawlers to build their indexes, and SEO tools like CMS Canvas run crawlers to audit sites the same way search engines see them.

Googlebot is Google's own web crawler. It discovers pages via links and sitemaps, fetches them, and passes the content to Google's indexing systems. Googlebot has a smartphone version and a desktop version, with the smartphone crawler doing the vast majority of crawling since Google moved to mobile-first indexing.

Search engines start from a list of known URLs (from previous crawls, sitemaps, and links on other sites), fetch those pages, extract every link they contain, and add new URLs to a queue to crawl next. They repeat this continuously, revisiting pages at a frequency based on how important the page seems and how often it changes. Pages blocked by robots.txt or unreachable by links may never be discovered.

Crawlability is how easily a search engine can access and move through your site. A crawlable site has working internal links, a clean robots.txt, an up-to-date XML sitemap, fast server responses, and no broken links or redirect chains blocking the path. Poor crawlability means pages get missed — and unmissed pages can't rank.

Crawling vs. indexing

Crawling is discovering and downloading a page; indexing is analyzing that page and storing it in the search engine's database so it can appear in results. Crawling comes first and is purely about access; indexing is a quality decision about whether the page is worth storing and showing.

Indexing is when a search engine processes a crawled page — reading its content, evaluating its quality and uniqueness — and adds it to the index, the giant database that search results are pulled from. If a page isn't in the index, it cannot rank for anything.

No. Google crawls far more pages than it indexes. After crawling, a page can be excluded for thin or duplicate content, a noindex tag, a canonical pointing elsewhere, low perceived quality, or simply because Google didn't consider it valuable enough. Crawled-but-not-indexed is one of the most common issues in Search Console.

The usual causes: content is thin, duplicated, or too similar to other pages; a noindex meta tag or wrong canonical tag; the page loads too slowly or errors intermittently; or Google judged the page low value. Fix it by strengthening the content, checking your meta tags and canonicals, and making the page genuinely distinct. An SEO crawler like CMS Canvas flags the technical causes — noindex tags, canonicals, duplicates — in one scan.

After crawling the raw HTML, Google queues pages that rely on JavaScript for a second phase called rendering, where it executes the JavaScript to see the final page — much like a browser does. Only after rendering does indexing happen. Heavy JavaScript sites can see delays or gaps if important content only appears after rendering.

Crawl budget

Crawl budget is the number of URLs a search engine will crawl on your site within a given timeframe. It's determined by two things: crawl capacity (how much load your server can handle without slowing down) and crawl demand (how much Google wants your content, based on popularity and how often it changes).

Google balances crawl capacity against crawl demand. Fast, healthy servers earn more crawling; slow responses and errors reduce it. Sites with popular, frequently-updated content generate more demand, so their pages get revisited more often. There's no fixed number — it adjusts continuously.

Speed up your server, fix broken links and long redirect chains, remove or consolidate duplicate and low-value pages, keep your XML sitemap clean and current, and block genuinely useless URLs (like faceted-filter combinations) in robots.txt. In short: make every crawl request count toward pages that matter.

Mostly no — this is a common misconception. Google says crawl budget is only a real concern for very large sites (roughly a million+ pages) or sites with rapidly changing content in the tens of thousands of pages. If your site has a few hundred or thousand pages, focus on content quality and crawlability basics instead.

The big offenders: infinite URL spaces from faceted navigation and filter parameters, soft-404s, long redirect chains, duplicate pages, session-ID URLs, hacked or spammy pages, and huge numbers of low-value pages. Every wasted fetch is a fetch not spent on the pages you want indexed.

Crawl depth

Crawl depth is how many clicks a page sits from the homepage (or another starting point). The homepage is depth 0, pages it links to are depth 1, and so on. Deeper pages are discovered later, crawled less often, and typically receive less link authority.

A common rule of thumb is to keep important pages within 3 clicks of the homepage. It's not a hard limit, but pages buried 5+ clicks deep get crawled less frequently and signal to search engines that they're less important.

Depth isn't a direct ranking factor, but it correlates strongly with outcomes: shallow pages get crawled more often, indexed faster, and inherit more internal link equity — all of which help rankings. Deeply buried pages often underperform simply because search engines treat them as low priority.

Strengthen internal linking: link to important deep pages from the homepage or top-level category pages, add breadcrumbs, build hub/category pages that group related content, use sensible pagination, and link between related articles or products. A crawl report that shows each page's depth — like the one CMS Canvas produces — tells you exactly which pages to surface.

Controlling & troubleshooting crawling

Submit an XML sitemap in Google Search Console and use URL Inspection's "Request Indexing" for individual priority pages. Beyond that, faster crawling comes from earning it: a quick server, fresh content published regularly, good internal linking, and links from other sites.

It varies per page, not per site. News homepages can be crawled every few minutes; a static page on a small site might be revisited every few weeks or months. Frequency follows how often a page changes and how important Google considers it. Search Console's Crawl Stats report shows your actual numbers.

Use the URL Inspection tool in Google Search Console — it shows the last crawl date, how the page was discovered, and whether it was indexed. Your server access logs also record every Googlebot visit, and the Crawl Stats report gives a site-wide view.

Use robots.txt to block crawling — but know the difference: robots.txt stops crawling, not indexing (a blocked URL can still appear in results if other sites link to it). To keep a page out of search results, use a noindex meta tag and let Google crawl it to see the tag. Never combine noindex with a robots.txt block, because Google can't read a tag on a page it isn't allowed to fetch.

robots.txt is a plain text file at the root of your domain that tells crawlers which paths they may or may not fetch. It's the first thing a crawler checks. Mistakes in it are dangerous — one wrong Disallow line can hide your entire site from search engines.

An XML sitemap is a list of the URLs you want crawled, with optional hints like when each was last modified. It helps search engines discover pages that are new, deep, or poorly linked, and it speeds up re-crawling of updated content. It's a strong hint, not a command — listing a URL doesn't guarantee crawling or indexing.

Crawl errors happen when a bot requests a URL and fails: 404s from deleted or mistyped pages, 5xx server errors, DNS failures, redirect loops, and timeouts. Fix them by redirecting removed pages to relevant replacements, repairing broken internal links, and resolving server issues. A site crawl surfaces every one of them — CMS Canvas checks all of this on every scan.

Internal links are the paths crawlers follow — a page with no internal links pointing to it (an orphan page) may never be found at all. Pages with more internal links get discovered sooner and crawled more often, and anchor text tells search engines what the linked page is about.

Impact on rankings

Crawling doesn't rank pages directly, but it's the gateway: a page that isn't crawled can't be indexed, and a page that isn't indexed can't rank. Crawl problems — blocked resources, errors, buried pages — quietly remove pages from the competition before ranking even begins.

Effectively no. In rare cases Google shows a URL it has never fetched (known only from links pointing to it) with no description, but it won't rank meaningfully for anything. Real rankings require crawling, then indexing.

Not directly — being crawled more often doesn't boost rankings by itself. But frequent crawling means changes and improvements get picked up faster, and it usually reflects that Google already considers the site important. Slow crawling delays how quickly your SEO work shows results.

Run a free SEO + AI SEO audit on your first site

Create a free account and run a full SEO and AI-readiness crawl in minutes — free accounts include 300 pages of scanning (up to 75 pages per scan), and Pro unlocks full-site crawls of 1,000 pages a month.