Choose Which Pages Train Your AI Knowledge Base
If you're using an AI knowledge base to help answer customer questions automatically, it needs to learn from your content first. Rather than dumping your entire website in and hoping for the best, you can preview exactly which pages get crawled and pick only the ones that are actually useful.
Why curating matters
Not every page on your site should train your AI — think about your privacy policy, old blog posts, or duplicate pages that would just add noise. Being selective up front means faster setup and a more accurate assistant, since it's only learning from content you've actually chosen.
Where to find this
Go to AI Agents > Knowledge Base, open the knowledge base you want to add content to, and click into the Web Crawler tab. Click + Add Website to get started.
Choose how broadly to search
You'll be asked to pick one of three scopes:
- Exact URL — just one specific page
- Path-based — a whole section of your site, like everything under yourbusiness.com/help
- Domain-wide — your entire website
Review what gets found
Once you submit a URL or domain, Vendi Today checks for a sitemap and shows you a list of every page it found — for example, "Sitemap found — 426 pages." You can page through the results, and if the sitemap doesn't cover everything you need, turn on crawling so it follows internal links to discover additional pages beyond what the sitemap listed.
Pick exactly what you want
Use the search box to filter the list down, then check or uncheck individual pages. A running counter shows how many you've selected against your knowledge base's page limit, so you'll know if you're approaching capacity before you commit.
Train it
Once you're happy with your selections, click Train Selected. Training runs in the background, so you're free to keep working while it processes — no need to sit and wait for it to finish.
Was this article helpful?
