Blog

Chatbot Trained on Your Own Data? A Readiness Checklist

A chatbot trained on your own data is only as good as your documents. Plain-English RAG explainer, data-readiness checklist and an honest FAQ-widget test.

Chatbot trained on your own data? A readiness checklist: RAG explained, readiness checklist and FAQ test, Banxal cover.

A customer asks the chatbot on your website if you’re open on the bank holiday Monday. It answers straight away and politely: yes, 9am to 5pm.

You’re closed. The bot found its answer in an old opening-hours PDF that nobody remembered was still in the shared folder.

The AI did its job. The data was the problem. Most guides to building a chatbot trained on your own data stop at the upload button. This post covers what comes before it: how these chatbots work in plain English, a data-readiness checklist, a hypothetical worked example, and an honest test of whether you need one at all.

TL;DR

Do you actually need a RAG chatbot — or will an FAQ widget do?

Often you don’t — there are three options, each more work to run well than the last.

1. An FAQ widget or help-centre search. You write the answers; the widget shows them. Nothing is generated, so nothing is made up. If most questions are repeats (“Do you deliver to…?”, “What’s your returns window?”), this is predictable and cheap.

2. An AI assistant that reads everything, every time. If your material is small, skip retrieval and give the AI all of it with each question. Anthropic: “If your knowledge base is smaller than 200,000 tokens (about 500 pages of material), you can just include the entire knowledge base in the prompt” (Anthropic). The trade-off is that you pay for reading it on every message. Our guide to AI agent cost per month shows how to work that out.

3. A RAG chatbot. Worth it when your material is too big for one prompt, changes often (stock, prices, policies), needs to be shown to some people and hidden from others, or answers must cite their source.

One more check: if answering means looking something up in a live system (“Where’s my order?”, “Is Tuesday at 3pm free?”), you’re not dealing with a document problem. You need an integration or an agent. See Zapier, n8n or an AI agent.

Decision gate flowchart “RAG chatbot or FAQ widget?” with four yes/no questions in order — mostly repeat questions leads to an FAQ widget, needing live data leads to an integration or AI agent, under ~500 pages and rarely changing leads to an AI assistant with the whole knowledge base in the prompt, otherwise a RAG chatbot

How a chatbot trained on your own data actually works

A correction first: for most of these tools, “trained on your own data” isn’t quite accurate. AWS describes RAG as having a model reference “an authoritative knowledge base outside of its training data sources before generating a response”, “all without the need to retrain the model” (AWS). IBM draws the same line: “RAG lets an LLM query an external data source while fine-tuning trains an LLM on domain-specific data” (IBM).

Think of a capable new hire with a filing cabinet. They don’t memorise it; when a question comes in, they pull the right folder and answer from that page. Updating is cheap (swap the page), and the cabinet is the weak point (a bad page means a bad answer).

Here’s the pipeline behind that picture.

Left-to-right RAG pipeline diagram: your documents, chunks, embeddings, retrieval, grounded answer — with a dashed loop showing embeddings update when documents change

  1. Documents. Gather your source material and clean it up. Pinecone’s walkthrough describes cleaning the data before anything else happens (Pinecone).
  2. Chunks. Long documents are split into short passages. OpenAI’s hosted version, for example, splits files into 800-token pieces (a token is roughly a word or part of a word) with a 400-token overlap by default (OpenAI).
  3. Embeddings. Each chunk is turned into a list of numbers, “a numerical representation of the data’s meaning” (Pinecone), and stored in a vector database (AWS).
  4. Retrieval. The customer’s question is turned into numbers the same way and matched against the stored chunks. This is how “Can I bring my dog?” can find a passage titled “Pet policy” even though the words barely overlap. OpenAI notes that semantic search surfaces similar results “even when they match few or no keywords” (OpenAI).
  5. Grounded answer. The question and the matching passages go to the AI model, which writes the answer from them. Both AWS and IBM point out that RAG systems can cite their sources, so you and your customers can check them.

Two things that trip up real chatbots

Chunks lose their context. Anthropic’s example is a passage that reads, on its own, “The company’s revenue grew by 3% over the previous quarter.” Which company? Which quarter? (Anthropic). The small-business version: a price-list line saying “£60, includes follow-up” whose heading got split off. Anthropic’s fix was to add context to each chunk before storing it and pair meaning-based search with keyword search. In their tests that cut failed retrievals by 49% (67% with an extra re-ranking step). You can do a version of this yourself in how you write your documents.

Exact codes need exact matching. Search by meaning can miss precise identifiers. Anthropic’s example is “Error code TS-999”, which keyword matching finds reliably. If customers ask about SKUs, part numbers or plan names, add keyword search too.

The data-readiness checklist

This decides whether your chatbot gives good answers, whichever tool you choose. Go through it before you upload anything.

One-page AI chatbot data-readiness checklist with five grouped sections and tick boxes: source-of-truth content, freshness and ownership, sensitive data and permissions, FAQ gap-check, and format cleanup

1. Source-of-truth content

2. Freshness and ownership

3. Sensitive data and permissions

4. FAQ gap-check

5. Format cleanup

Worked example (hypothetical): a small dental practice

An illustrative scenario, not a client. Imagine a small dental practice with two clinics. The front desk spends much of the day answering the same calls about prices, opening hours, accepted insurers and how to prepare for a procedure. The owner wants “a chatbot trained on our website and leaflets.”

The gate. Questions repeat, but answers depend on the clinic and the treatment, so a fixed FAQ feels stiff. Booking questions need the live diary, so they route to the existing booking link instead. Everything else (website pages, a price list, a dozen leaflets) is well under 500 pages, so no RAG yet — an assistant with the whole knowledge base in the prompt will do for version one.

The checklist turns up the real work:

None of this needed an engineer, though skipping it would have produced wrong answers. If the practice later adds clinics or a staff-only assistant, the cleaned-up content moves straight into a RAG setup.

Platform vs. custom build: what changes once your data is ready

Once your data is ready, picking the tool is the easier decision — the work above carries over whichever you choose.

No-code chatbot platform Hosted retrieval API (e.g. OpenAI file search) Custom RAG build
Setup Upload files, paste a snippet on your site A developer connects files to a model Designed around your systems
Chunking & search Fixed by the vendor Adjustable. OpenAI allows 100–4,096-token chunks (OpenAI) Fully controllable, incl. keyword + meaning-based search
Permissions Usually one bot = one set of files Tag-based filtering Per-user access tied to your logins
Live data Rarely With extra development Yes, via integrations
Best for Public FAQ-style content Small teams with a developer Multiple sources, staff + customer bots, citations

Running costs: storing the documents is usually the small part. OpenAI’s vector storage, for instance, is free up to 1 GB and $0.10 per GB per day after that (OpenAI). The model’s cost per conversation usually matters more, which is what our cost-per-month guide works through.

A platform is probably enough for public content and simple questions. You’ve outgrown it when you need per-user answers, live system data, reliable exact-code lookups, or a way to test answer quality before customers see it.

Want a second opinion before you build?

Worked through the checklist and still unsure whether you need an FAQ widget, a simple assistant or a full RAG chatbot? Our AI integration & agents team builds chatbots and assistants on your own data, and we’ll tell you plainly if a simpler option will do. Tell us what your customers keep asking.

FAQ

Can I train ChatGPT on my own business data?

You usually don’t need to. Most “own data” chatbots use RAG: they look up your documents at question time and leave the model unchanged. Fine-tuning, which does change the model, is “computationally expensive and resource-intensive” (IBM) and rarely a small business’s first step.

How much data do I need for a chatbot trained on my own data?

Less than you’d think, but it must be current, consistent and cover the questions people actually ask. Under about 500 pages, Anthropic suggests skipping retrieval and putting it all in the prompt (Anthropic).

Is it safe to put customer data into an AI chatbot?

Keep customer and staff data out of a public bot’s knowledge base entirely. Anything in it can end up shown to anyone. Give staff assistants a separate, access-controlled collection.

How do I stop the chatbot making things up?

Fill the gaps, remove outdated and conflicting documents, have the bot cite sources, and hand over to a person when it can’t find an answer. RAG gives the model real material to work from. It can’t fix missing or wrong material.


What’s the one question your customers ask that you’d most like a chatbot to answer, and is it actually written down anywhere yet? Tell us in the comments.


If you’ve tried a chatbot on your own documents, what tripped it up first: outdated files, missing answers, or something else? Tell us in the comments.

Building something like this?

Banxal designs and ships AI-first software in weeks. Tell us what you’re working on and we’ll give you an honest take.