How to Add Your Business Knowledge
Supply knowledge by pasting text, uploading PDF/DOCX/TXT/Markdown/CSV, or importing one public web page. See processing status, reindex and edit sources.
Three ways in
The Grounded AI Agent & Knowledge Base answers from content you supply, so the first real task is supplying it. There are three routes, and each suits a different kind of material.
Paste text. The fastest option and the most underused. If the answer lives in an email you have sent forty times, paste it. No formatting to clean up, no file to find.
Upload files. PDF, DOCX, TXT, Markdown or CSV. Right for material that already exists as a document: service specifications, terms, price lists, product data exported as CSV.
Import a page by URL. Pulls the readable text of one public web page. Good for content that lives on your own site already, like an FAQ page or a service description.

What to add first
Start with what your team repeats. Open your sent folder, or think about the last twenty enquiries, and write down the questions that came up more than twice. That list is your first knowledge base, and it is almost always more useful than uploading your entire document library.
A practical starting set for most businesses:
- Opening hours, including anything seasonal
- Service or product descriptions, in customer language
- Coverage area or delivery terms
- Pricing structure, or how a quote is arrived at if you do not publish prices
- Lead times and what affects them
- Warranty, guarantee or returns terms
- The three things people always misunderstand about what you do
That last one is worth an evening on its own. Whatever your team spends most time correcting is exactly what the AI needs written down.
Processing, status and reindexing
After you add a source, it processes. The list shows status, so you can see when a source is ready rather than guessing why an answer has not changed yet.
You can edit source text directly in the workspace, which is the right move for a small correction, a changed price, a retired service. After an edit, reindex so retrieval picks up the new text. Sources can also be deleted when they stop being true, and deleting outdated material is as valuable as adding new material. Stale content that still retrieves well is worse than no content at all.

Write for retrieval, not for a brochure
Two habits make a noticeable difference to answer quality.
Use the customer’s words. If people say “callout fee” and your document says “attendance charge”, retrieval has a harder job. Include both.
Keep each source about one thing. A single document covering pricing, hours, coverage and terms retrieves worse than four short ones. Retrieval works on passages, so clean topical boundaries help it find the right passage.
Short question-and-answer formatting works particularly well, because it mirrors how the question will arrive.
The limits, stated plainly
Three boundaries to plan around rather than discover:
One page per URL import. There is no full-site crawl and no continuous website sync. If your FAQ page changes, reimport it or edit the source. Nothing updates itself.
Scanned PDFs need OCR first. Built-in OCR is not part of the product, so a scanned document with no text layer will not produce usable knowledge.
Quotas count sources and calculated pages. A knowledge page is calculated from 3,000 characters of extracted text, which will not always match one original PDF page. Internal search chunks are not what gets counted.
Test what you added
The fastest way to check your knowledge base is to interrogate it. Ask the chat the ten questions from your list. Anything it answers well is done. Anything it refuses is a missing source, and you now know exactly which one. Anything it answers wrongly is usually two sources disagreeing with each other, which is worth finding early.
Run that test again after any significant edit. It takes ten minutes and catches most of what would otherwise reach a customer.
Organising more than a handful of sources
Once you pass a dozen sources, structure starts to matter. A knowledge base is not a filing cabinet, nothing needs to be tidy for its own sake, but a few decisions make retrieval noticeably better.
Split by topic rather than by document. If your terms and conditions cover returns, delivery and warranty, three focused sources will retrieve better than one long one, because each passage sits closer to the question it answers. Conversely, do not shred a coherent explanation into fragments; a paragraph that only makes sense with its neighbours should stay with them.
Name sources so a colleague can audit them later. “Delivery and lead times, updated Sept” is a name that tells you when to check it. “Doc 4” is not.
Keep one source per fact. The most common cause of a confidently wrong answer is two sources that disagree: an old price list and a new one, both indexed and both retrievable. Whichever wins is essentially arbitrary. When you update something, update or delete the old version rather than adding alongside it.
A maintenance rhythm that survives a busy month
Knowledge bases decay quietly. Nothing breaks; the answers just drift further from what you would say today.
A workable rhythm is quarterly, plus an event trigger. Quarterly, reread the sources with a date, a price or a lead time in them. On an event such as a price change, a retired service or a revised policy, update the source the same week, while someone still remembers the detail.
The other half of maintenance is your transcripts. Refusals tell you what is missing; wrong answers usually tell you what is contradictory. Ten minutes reading conversations once a month will surface more genuine problems than any amount of speculative tidying.
One habit worth keeping
Whenever you answer a question you have answered before, ask whether it is in the knowledge base. If not, paste it in while the wording is fresh. Two minutes at the moment you notice beats an hour of retrospective auditing later, and it is how a knowledge base stays current without ever becoming a project.
Read next: Strict, Balanced and Open grounding modes compared, or how to reduce AI chatbot hallucinations.
Learn more about Grounded AI Agent & Knowledge Base
An AI agent that answers from your own knowledge base, with grounding modes, configurable refusals, and handoff to a person.
Questions people ask about this
What file types can I upload?
PDF, DOCX, TXT, Markdown and CSV. You can also paste text directly, which is often the fastest route for a short FAQ, and import the readable text of one public web page by URL.
Can it crawl my whole website?
No. A URL import handles one public page. Full-site crawling and continuous website synchronisation are not part of the product, so add the specific pages that carry answers rather than pointing us at a domain.
What about scanned PDFs?
They need OCR before import, built-in OCR is not part of the product. Run the file through OCR first so it contains real text, then upload it.
Related guides
How to Reduce AI Chatbot Hallucinations with Grounding Controls
Why chatbots invent facts, and how strict retrieval gating, refusal thresholds, handoff and source attribution help reduce it, controls that guide behaviour.
Read guideStrict, Balanced and Open Grounding Modes Compared
Compare onmsg's grounding modes: Strict (refuse when no relevant content), Balanced (prefer your knowledge), and Open (general conversation), plus when each fits.
Read guideWebsite Data Privacy: What Visitors' Chat Data Means for GDPR and CCPA
What chat and enquiry data is collected, how retention settings differ by plan, and general GDPR/CCPA considerations for chat intake, guidance, not legal advice.
Read guideWhat Does "Grounded AI" Mean for Website Chat?
Grounded AI answers from your own content instead of general knowledge. What grounding is, why it matters, and how it differs from a general chatbot.
Read guideWant to try this on your own site?
No credit card required.