Websites
Set up Websites
Ingest public web pages — a handbook site, a help centre, published policies — and keep them in step as the site changes.
Before you start
- No credential to create — this reads whatever a signed-out visitor could reach. Pages behind a login can't be fetched.
- Decide how to point at your pages: an explicit URL list, a sitemap, or both together.
- Works for an internal address too, not just the public internet.
- Documents → Website / Help centre: paste your URLs and/or a Sitemap URL.
- Choose access groups, Test, then Sync.
What comes in
- Each URL you list, and every page a sitemap lists, becomes one document — titled from the words in the page's own title tag, when it has one.
- A PDF, Word, Excel or PowerPoint file is ingested the same way when a URL points straight at it; linked images, audio and video are skipped.
- Only the addresses you configure are ever fetched — links inside a page are not followed, so nothing outside your list can come in.
| URLs | One page URL per line, fetched exactly as given. |
|---|---|
| Sitemap URL | A sitemap.xml — every page it lists is ingested. Sitemap index files (the WordPress/Shopify layout) and gzipped sitemaps both work; combine with the URL list or use either alone. |
| Exclude patterns | Under Advanced — leave out pages whose URL or filename matches a pattern, for example *.pdf or /blog/*. |
Access
Everyone in the access groups you choose can ask about every page this source brings in. Public pages carry no reader list of their own to mirror, so the groups you pick when setting it up are the only control.
Every sync re-checks each page and updates the ones that changed; robots.txt is not consulted, so point it at exactly what should be searchable. A sitemap listing more than 5,000 pages only syncs the first 5,000. Test only confirms the first URL responds — a page that goes down later is skipped and noted, not treated as a failure.
Was this page helpful?
Last updated 20 Sep 2026