SilentChat

How does the web crawler work exactly?

Last updated: September 9, 2026

The KB Web Crawler

Instead of manually feeding your knowledge base with Markdown, the crawler can index your existing website or help documentation.

Setup

  1. Verify Domain (Settings → Knowledge Base → Add Domain → Set DNS-TXT-Record)
  2. Enter Start URL (e.g., https://help.your-company.de)
  3. Path Filter (optional): only under /docs, or /help, or custom patterns
  4. Depth: maximum 3 hops from the seed (default), up to 6 on Pro

What the Crawler Does

  • Follows internal links starting from the seed URLs (no external crawling)
  • Respects robots.txt + Meta-Robots
  • Extracts main content (Mozilla Readability + own pruner)
  • Strips navigation, footer, ads, modal wrappers
  • Per page: creates 1 article + vector embedding
  • Follows User-Agent: SilentChatBot/1.0 (+https://silentchat.de/crawler)

Refresh Intervals

  • Manual: Settings → "Re-crawl now" starts immediately
  • Automatic: weekly (default), daily on Enterprise
  • Only changed pages are re-embedded (ETag + Content-Hash-Check)

Limits

PlanPages per DomainDomains
Starter1001
Growth5003
Pro2,00010
Enterprise10,000Unlimited

Common Issues

  • JS-rendered Pages: the crawler uses a headless Chrome — works. But login-protected content is not crawlable.
  • PDFs: PDFs are indexed separately via document upload, not via crawler.
  • Privacy: nothing containing PII (form confirmation pages, etc.) should be crawled — crawler has no cookie authentication.
How does the web crawler work exactly? — Help Center — SilentChat | SilentChat