How does the web crawler work exactly?
Last updated: September 9, 2026
The KB Web Crawler
Instead of manually feeding your knowledge base with Markdown, the crawler can index your existing website or help documentation.
Setup
- Verify Domain (Settings → Knowledge Base → Add Domain → Set DNS-TXT-Record)
- Enter Start URL (e.g., https://help.your-company.de)
- Path Filter (optional): only under /docs, or /help, or custom patterns
- Depth: maximum 3 hops from the seed (default), up to 6 on Pro
What the Crawler Does
- Follows internal links starting from the seed URLs (no external crawling)
- Respects
robots.txt+ Meta-Robots - Extracts main content (Mozilla Readability + own pruner)
- Strips navigation, footer, ads, modal wrappers
- Per page: creates 1 article + vector embedding
- Follows User-Agent:
SilentChatBot/1.0 (+https://silentchat.de/crawler)
Refresh Intervals
- Manual: Settings → "Re-crawl now" starts immediately
- Automatic: weekly (default), daily on Enterprise
- Only changed pages are re-embedded (ETag + Content-Hash-Check)
Limits
| Plan | Pages per Domain | Domains |
|---|---|---|
| Starter | 100 | 1 |
| Growth | 500 | 3 |
| Pro | 2,000 | 10 |
| Enterprise | 10,000 | Unlimited |
Common Issues
- JS-rendered Pages: the crawler uses a headless Chrome — works. But login-protected content is not crawlable.
- PDFs: PDFs are indexed separately via document upload, not via crawler.
- Privacy: nothing containing PII (form confirmation pages, etc.) should be crawled — crawler has no cookie authentication.