How exactly does the web crawler work?
Zuletzt aktualisiert: 15. Mai 2026
The KB web crawler
Instead of feeding your knowledge base by hand with Markdown, the crawler can index your existing website or help docs.
Setup
- Verify domain (Settings → Knowledge Base → Add domain → set DNS TXT record)
- Enter start URL (e.g. https://help.your-company.com)
- Path filter (optional): only under /docs, or /help, or custom patterns
- Depth: max 3 hops from the seed (default), up to 6 on Pro
What the crawler does
- Follows internal links from the seed URLs (no external crawl)
- Respects
robots.txtand meta robots - Extracts main content (Mozilla Readability + custom pruner)
- Strips navigation, footer, ads, modal wrappers
- Per page: creates 1 article + a vector embedding
- User-Agent:
SilentChatBot/1.0 (+https://silentchat.de/crawler)
Refresh schedules
- Manual: Settings → "Re-crawl now" starts immediately
- Automatic: weekly (default), daily on Enterprise
- Only changed pages are re-embedded (ETag + content-hash check)
Limits
| Plan | Pages per domain | Domains |
|---|---|---|
| Starter | 100 | 1 |
| Growth | 500 | 3 |
| Pro | 2,000 | 10 |
| Enterprise | 10,000 | Unlimited |
Common issues
- JS-rendered pages: the crawler uses headless Chrome — works. Login-protected content is not crawlable.
- PDFs: indexed separately via document upload, not the crawler.
- Privacy: anything with PII (form confirmation pages etc.) should not be crawled — the crawler has no cookie auth.