SilentChat

Auto-Refresh — keep KB up to date

Last updated: September 16, 2026

Your website changes — new blog posts, pricing updates, FAQ extensions. To ensure the AI bot continues to provide accurate answers, SilentChat can automatically re-index the Knowledge Base per domain. The feature is opt-in per domain and is charged against the normal AI token budget (crawling itself does not cost LLM calls, but the embedding generation of new articles does).

Activation

Settings → Domains → verified domain → "Auto-Refresh" switch

Each verified domain gets its own toggle:

  • Toggle On: the next crawl run is scheduled for the next daily tick
  • Interval Picker: 3 / 7 / 14 / 30 days (Default: 7)
  • "Re-index now": triggers a refresh immediately — useful after a major website change, without waiting for the cron tick

Unverified domains do not show the strip — the crawler requires DNS verification (see Domain Verification) before it does anything.

What Happens in the Background?

Three cron jobs orchestrate the feature:

  1. crawl_refresh_scheduler (daily) — lists all domains with next_crawl_at <= NOW(), queues a crawl job per domain + creates a crawl_runs entry, bumps next_crawl_at by the configured interval. Cap: 50 domains per tick (prevents a newly activated enterprise configuration from flooding the worker pool).

  2. crawl_runs_finalize (every 5 minutes) — finds runs without finished_at whose jobs are all completed or failed (or > 30 min old). Calculates the diff via CrawlDiffService and persists the stats:

    {
      "new": 3,
      "updated": 7,
      "removed": 2,
      "total_after": 47
    }
    
  3. crawl_runs_purge (daily) — drops crawl_runs entries older than 90 days. Consistent with the general activity log retention.

Diff Display in the Dashboard

Each domain gets a small history strip (in preparation). Currently, you see:

  • Last crawl timestamp under the toggle
  • "Last indexing: 4 days ago" as a subtle chip

In the next iteration, the drill-down modal with the article list per run (new / updated / removed) will be added — until then, the counts are available in the backend log via crawl_refresh tenant=... run=... diff=new:3 updated:7 removed:0 total:47.

What is Crawled?

Currently, the scheduled refresh crawls only the root URL of the verified domain. Multi-page discovery via sitemap.xml is a separate plan in the roadmap (see crawler-refresh-scheduler.md).

This works well for:

  • Pricing pages
  • Changelog pages
  • FAQ entry pages
  • Single-page pricing comparisons

If you want to index multiple sub-pages today, use the manual KB → Auto-Import path (Plan #164) instead, which processes a list of URLs.

Token Consumption

  • Crawling itself: 0 LLM tokens (HTTP request + HTML parse)
  • Embedding calculation for new / changed article chunks: against embed_index budget
  • Scale: A 5 KB webpage produces 5-10 chunks ≈ 5,000 tokens for embedding indexing. With a 7-day interval ≈ 22,000 tokens/month per domain.

GDPR

Auto-Refresh does not send PII to external services. The crawler only reads the publicly accessible website — same data protection implications as during the initial indexing in the onboarding wizard (Plan #163/#164). Embeddings are computed by IONOS in Germany — the same provider as for the AI chatbot.

Frequently Asked Questions

My Auto-Refresh toggle is on but nothing happens. Check domains.next_crawl_at via SQL — if in the future, the scheduler is waiting. Filter the backend log for crawl_refresh_scheduler. If the day has passed but the tick did not run, the lock is likely still held by another cron worker — it will run on the next day tick.

Can I globally disable Auto-Refresh? Per domain is sufficient — toggle off, next_crawl_at is set to NULL, the scheduler skips the domain. Globally would only work by stopping the cron container (or skip logic via RuntimeConfig — will be added if desired).

What happens if the crawl fails mid-run? The crawl_runs_finalize job also picks up stuck runs after 30 minutes and stamps them with the diff counts up to that point. The next scheduler tick re-enqueues — no drop effect.

How do I see the history? Via API: GET /api/v1/domains/<id>/crawl-runs lists the last 20 runs with timestamps + stats. The drill-down UI for this is in preparation.

Does Auto-Refresh consume my plan limits? Crawling itself does not (no LLM call). But the embedding generation per new chunk does — this runs against embed_index and is subject to the monthly embedding quota (see embedding-token-tracking.md for configuration).

Auto-Refresh — keep KB up to date — Help Center — SilentChat | SilentChat