codexier.

AI & Automation

Chatbots That Answer From Your Own Documents

By CodexierPublished 7 min read

A chatbot that answers from your own documents is not 'trained' on them in the way the phrase suggests. It looks the relevant passage up at the moment of the question and writes an answer from it. That mechanism decides everything: which documents work, why answers go stale, and how access control has to be built. This guide explains it plainly.

How retrieval-based chatbots work

The consequence worth understanding is that the chatbot is only as good as the passage it finds. If the right answer is spread across three documents, or buried in a table the indexer could not read, the model gets the wrong context and answers confidently from it. Most 'the bot made something up' complaints are really 'the bot was handed the wrong passage'. That failure mode is a sourcing problem before it is a model problem, so this guide focuses on the sources.

Which documents make good sources

Good sources share three traits: each answers a question on its own, they are written in the language your customers ask in, and they carry a date. Swedish companies typically have some of both kinds.

SourceWorks well becauseWatch out for
FAQ pages and help articlesOne question, one answer, already in the customer's wordsOld answers nobody removed
Terms, delivery and return policiesPrecise wording the bot can quoteSeveral versions in circulation
Product sheets and price listsStructured, factualTables in PDFs that index badly; prices excl. and incl. VAT mixed
Internal manuals and routinesHigh value for staff-facing botsWritten for insiders, full of unexplained abbreviations
Email threads and chat logsContain real answersPersonal data, contradictions, no clear final version

As a rule: if a new employee could answer the question from the document, so can the bot. If they would need to ask a colleague, the document is not a source yet.

Format matters less than structure. Clean HTML, Markdown, Word files and text-based PDFs all index well. Scanned PDFs, screenshots of tables and slide decks with the meaning in the layout do not: the indexer sees either nothing or a jumble.

Keeping answers current

Because nothing is memorised, freshness is a process problem, not a technical one. The index has to be rebuilt when documents change, and someone has to own the documents. The failure pattern is predictable: the bot is launched with a clean set of sources, prices change in the spring, and by summer it is quoting last year's delivery terms.

  • Point the index at the live source (the website, the shared drive folder) rather than at copies made on launch day.
  • Schedule re-indexing, daily for a webshop, weekly for a service company.
  • Give every source document an owner and a review date, and show the date in the bot's answer.
  • Keep a log of questions the bot could not answer; it is the best list of missing documents you will ever get.
  • Remove sources rather than letting them rot. A missing answer is better than a wrong one.

Access control for internal content

A customer-facing bot should only ever see public documents; that is the simplest rule and it removes most risk. The moment the same system also serves staff with internal manuals, salary routines or customer records, retrieval must respect permissions: the search step should only return passages the person asking is allowed to read. Filtering the answer afterwards is not enough, because the model has already seen the text.

Public bot

Index only what is already on the website or in published PDFs. No customer data, no internal routines.

Staff bot

Log in with the company identity (Microsoft 365, Google Workspace) and filter retrieval by the same groups that govern the documents.

Both

Two separate indexes. Sharing one and hoping the prompt keeps secrets is not a control.

Personal data in sources also triggers GDPR: a processor agreement with the model provider, a documented lawful basis, and a plan for deletion requests, which means being able to remove a person from the index, not only from the source.

Signs your content is not ready

Sometimes the honest recommendation is to write first and build later. If the answers to your most common questions live in people's heads or in email threads, the project's first phase is documentation, and a chatbot on top of nothing answers nothing. If two documents contradict each other on delivery times, the bot will pick one at random. If the website is in Swedish but customers ask in English, the retrieval quality drops unless sources exist in both languages, so plan for bilingual sources from the start.

When you should not buy this from us: fewer than a handful of recurring questions, or questions that need judgement (a legal assessment, a medical answer) rather than lookup. For that, a well-written FAQ page and a fast reply routine beat any bot. When the sources are ready, our chatbot setup is fixed-price at 14,990 kr and includes the indexing, the handover rules and a test round on your real questions.

Frequently asked questions

Is the chatbot trained on my documents?

Not in the sense of changing the model. The documents are indexed and searched at question time, and the model writes an answer from what it finds. This is why updates take effect immediately after re-indexing and why the bot can cite its source. Fine-tuning a model on company text is a different, rarely necessary approach.

Do my documents get used to train the provider's model?

That depends on the provider and the contract. Business and API plans from the major providers generally state that customer data is not used for training, but you should have that in the processor agreement rather than assume it. Ask where the index and the logs are stored, and for how long.

How many documents does it need?

Fewer than most people think. A well-maintained FAQ, your terms and a price list often cover most customer questions. Volume is not the goal; coverage of the questions actually asked is. Start with the top twenty questions from your inbox and add sources until each has a clear answer.

Can it answer from Fortnox or my CRM as well?

Yes, but that is a different mechanism: a live lookup through an API rather than a document index, and it needs authentication so the bot only fetches data the asking customer is entitled to. It is common for order status and invoice questions, and it is scoped and priced separately from the document part.

What happens when the bot does not know?

It should say so and hand over: collect the question and contact details, or route to a person. A bot that always answers is a bot that sometimes invents. The handover rules are part of any serious setup, and the unanswered-question log becomes your list of documents to write next.

Are your documents ready to be a source?

Send us a few of them and your most common questions before the call. In fifteen minutes we tell you what would work now, what needs rewriting first, and what a setup would cost.

Book a free 15-minute call