Chapter 05 · Knowledge Base

Serpio's site index is not Google's index

The Website crawler builds Serpio's private index of your pages for writing and internal links. What the count means, what gets crawled, and how to fill gaps.

Updated Sep 20, 2026 · 5 min read

On this page

The Website crawler panel at the top of the Knowledge base page reports how many of your pages Serpio has indexed. That number is often lower than the count in Google Search Console, and people ask why. The short answer: it is a different index, built for a different job.

Where to find it

Expand Website crawler. The panel is titled Serpio site index and shows a line such as 12 pages indexed · last refreshed 3d ago, a Refresh button, and a larger card repeating the page count with the note Available to Serpio for writing and internal-link suggestions. While a crawl runs, the title changes to Crawling your site…, the button reads Crawling…, and a Crawl in progress badge appears. The panel says a crawl usually takes 30 to 90 seconds.

The pages themselves are listed with your other sources. Choose the Website filter on the Knowledge base page to see them, open any page to read its preview, and set its usage like any other source.

Not Google's index

Google's index records which of your URLs Google has crawled and may show in search results. Serpio's site index records which of your pages Serpio has read, summarised, and stored so the writer can use them. A page can be in one and not the other. Refreshing Serpio's index does nothing to Google, and being in Serpio's index says nothing about rankings.

What the index is used for

  • Brand context. Crawled pages default to Brand knowledge · Facts allowed, so the writer learns what you sell, how you describe it, and what you charge from your own words.
  • Article briefs. The most common topics across your indexed pages are added to each article brief, so the writer frames the piece around what you actually offer. Idea collections are prepared from public research and do not read the index.
  • Internal links. When an article is written, Serpio picks link targets from indexed pages on your domain, including in-page sections it found such as a features anchor, and never invents a URL. See Internal links.
  • Evidence. Passages from crawled pages appear in an article's Knowledge used panel labelled Website knowledge.

What gets crawled

Serpio discovers pages from your sitemap and by following links from your homepage, keeps only pages on your own domain, and ranks them before applying an automatic cap of about 100 pages per crawl. The cap keeps processing predictable on large sites. Within it, pages are prioritised roughly as follows.

PriorityPages
HighestThe homepage.
HighProduct, feature, solution, pricing, plans, demo, trial, get-started, docs, help, support, guides, and resources pages.
MediumAbout and contact pages, then other pages by depth, shallower first.
LowThe blog index, then individual blog posts and news articles.
LowestPrivacy, terms, cookie, and legal pages.

For each page Serpio keeps the title, description, headings, the main text, and a short excerpt, classifies it as product, blog, category, or page, and extracts topics and entities. Pages that return no readable text, or that fail to load, are recorded as failed; the panel shows the last error line under the count.

Weekly recrawl and Refresh

Serpio recrawls every onboarded site once a week, on Sunday. A recrawl repeats discovery and skips pages whose content hash has not changed, so an unchanged site produces no new versions. Pages you have paused, or set to Manually, are left alone. Press Refresh when you have shipped something the writer should know about now: a new pricing page, a renamed product, a new docs section. Refreshing does not spend article credits and does not affect Google.

Why a page might be missing, and how to add it

  • It fell outside the automatic cap. Blog posts are ranked lowest and are the first to be cut on large sites.
  • It is not in your sitemap and not linked from pages that were crawled.
  • It lives on a different origin, such as a docs. subdomain or a separate app domain. Only pages on the brand's own origin are kept.
  • It failed to load, returned something other than HTML, or had too little readable text.
  • It was crawled but then paused, or set to Manually with an existing version, so the recrawl skipped it.

If the panel shows We couldn't fetch pages from this site automatically. You can add pages manually under Knowledge base., discovery found nothing at all, usually because the site blocks crawlers or has no sitemap and no crawlable links. The fix in every case is the same.

  1. 1

    Add it as a URL source

    Press Add knowledge, choose Add URL, and paste the page address. URL sources are not subject to the cap and can live on any public domain, including subdomains.

  2. 2

    Set usage

    For your own pages choose Brand knowledge so the writer treats it like a crawled page. Pick Weekly under Check URL for changes to keep it in step with the recrawl.

  3. 3

    Confirm it is Ready

    The row appears under the Added URLs filter. Once it reads Ready, the writer can use it. It becomes an internal-link candidate only if it is on the brand's own domain; a subdomain page is used as knowledge but not linked.

Common questions

  • Why does Serpio show 40 pages when Google shows 400?

    Different indexes. Serpio caps automatic discovery at about 100 pages and prioritises product and support pages over blog archives. Google indexes whatever it decides to. Neither number affects the other.

  • Can I raise the cap?

    Not from the dashboard. Add the specific pages you need as URL sources; they are not counted against the cap.

  • Does refreshing the site index re-run the site report?

    No. The site report is a separate analysis with its own re-run. See the guide on when to re-analyze.

Your next article starts here.

Turn a topic that is trending right now into your first article.

Write my first article

3 free articles a month. No credit card.