Web Crawler
The Web Crawler automatically imports content from websites into your Knowledge Base. Ideal if you already have existing documentation or a FAQ page.
Prerequisites
- KB Crawler Feature must be activated in your plan
- The target website must be publicly accessible
Crawl URLs
- Knowledge Base → Crawler → "Crawl URLs"
- Enter one or more URLs (one per line)
- Click "Start Crawling"
- The crawler loads the pages, extracts the text, and creates articles
Import Sitemap
For larger websites:
- Knowledge Base → Crawler → "Import Sitemap"
- Enter the sitemap URL (e.g.,
https://example.com/sitemap.xml) - The crawler automatically finds all pages and imports them
Manage Crawl Sources
Under Crawler → Sources, you can see all configured URLs:
| Column | Description |
|---|---|
| URL | The crawled URL |
| Status | Active / Disabled |
| Last Crawl | When the URL was last crawled |
| Content Hash | Change detection — only changed content is updated |
Auto-Recrawl
If Auto-Recrawl is activated in your plan, sources are automatically recrawled at a configurable interval (e.g., every 7 days). This keeps your KB up to date.
Configuration: The Auto-Recrawl interval is set in the plan under kb_auto_recrawl_days.
Robots.txt
The crawler respects the robots.txt file of the target website. Pages that are blocked there will not be crawled. The robots cache is automatically updated.
Limits
| Plan | Max. crawled pages |
|---|---|
| Growth | 100 |
| Professional | 500 |
| Enterprise | Unlimited |
Crawl Jobs
Under Crawler → Jobs, you can see the status of ongoing and completed crawl jobs. Jobs can be aborted.