# AI Agent Knowledge Base: How to Build One Your Agent Actually Uses

> Build an AI agent knowledge base that stays accurate. What to include, files vs RAG, fixing stale and contradictory info, and a 20-question test.

Last updated: 2026-09-28T10:00:00

Canonical URL: https://crevio.co/blog/ai-agent-knowledge-base

*Last updated: September 2026*

**An AI agent knowledge base rarely fails because the agent can't find your refund policy. It fails because the agent finds two of them.** In 2024, a Canadian tribunal ordered Air Canada to compensate a customer after the airline's website chatbot described a bereavement refund that its own policy page ruled out. The model did its job. The knowledge behind it disagreed with itself.

This guide is for business owners who want an AI agent to know the business well enough to act on it, whether the agent runs operations for you or answers your customers. It covers what to put in the knowledge base, what to leave out, how agents actually read it, and how to test it before a customer does.

- **An AI agent knowledge base is the written record your agent works from**: policies, product facts, FAQs, procedures, and what the agent has learned about you
- **Sort by how often things change.** Stable rules go in files. Anything that changes more than monthly (prices, stock, orders) should be looked up live, never pasted in
- **Most small businesses do not need vector search.** Below about 500 pages, an agent can open the right files directly, which is simpler and much easier to fix
- **Contradictions do more damage than gaps.** An agent that finds no answer can say so. An agent that finds two answers picks one
- **Test with 20 real customer questions** in a fresh chat, and fix the source file, never the chat

## What Is an AI Agent Knowledge Base?

An AI agent knowledge base is the collection of written business knowledge an agent reads before it acts: your policies, product details, answers to common questions, step-by-step procedures, and the facts it has picked up while working with you. It is what turns a general-purpose model into something that knows *your* refund window, *your* best customer, and *your* way of saying no.

It holds four kinds of knowledge, and they behave differently:

| Kind | Example | How it goes wrong |
|---|---|---|
| **Facts** | "The Pro course has 12 modules and lifetime access" | Goes stale when the product changes |
| **Rules** | "Refunds within 30 days of purchase, no questions asked" | Contradicted by an older copy somewhere else |
| **Procedures** | "How we onboard a new coaching client" | Too vague to follow, so the agent improvises |
| **Memory** | "Axel prefers short replies and hates exclamation marks" | Learned from one bad day and never corrected |

The term covers two jobs. One is an agent that works *for* you: it drafts emails, updates products, and runs reports, and it needs to know how your business works. The other is an agent that answers *your customers* from a support knowledge base. The same knowledge base can serve both. The stakes differ, because a customer acts on what the agent tells them. We cover the support side in its own section below.

## What to Put in Your AI Agent Knowledge Base

Start with what you already explain repeatedly. For most small businesses, that is a short list:

1. **One paragraph on the business.** What you sell, who buys it, and what you are not
2. **Customer-facing policies.** Refunds, cancellations, shipping or delivery, privacy, and support hours, in their current wording
3. **Product facts.** What each product includes, who it is for, and who it is *not* for
4. **FAQs in your customers' words.** Copy real questions from your inbox, not the phrasing on your sales page
5. **Voice and tone.** Three example replies you were happy with beat a paragraph of adjectives
6. **Escalation rules.** What the agent must hand to you: legal threats, refunds over a set amount, anything medical or financial
7. **Procedures you repeat.** Onboarding a client, launching a product, answering a chargeback

The harder decision is where each piece lives. Sort by how often it changes, not by topic:

![Diagram showing three layers of business knowledge for an AI agent: standing instructions read on every task, reference files opened when needed, and live connections looked up fresh](https://crevio.co/vite/assets/where-business-knowledge-belongs-i0z68ico.svg)

### What to leave out

- **Anything that changes more than monthly.** Prices, discount codes, stock, bookings, and order status belong in the system that owns them. The agent should look them up through an integration. A price typed into a document is a future wrong answer
- **Passwords and API keys.** Knowledge files get read, quoted, and copied into drafts. Credentials belong in a proper secrets store
- **Old versions.** Delete the 2024 refund policy. Keeping it "for reference" in the same folder is how the Air Canada problem starts
- **Marketing copy that overstates.** If your sales page says "unlimited support," the agent will promise unlimited support. It cannot tell enthusiasm from policy

## How Agents Read a Knowledge Base: Files, Search, or RAG

There are three ways an agent can get at your knowledge, and the right one depends mostly on size.

| Approach | How it works | Best for | Main weakness |
|---|---|---|---|
| **Load everything** | The whole knowledge base goes into every request | A few pages of core instructions | Cost, and quality drops as the prompt grows |
| **Files plus search** | A short index is always loaded; the agent searches and opens files as needed | Most small businesses | The agent must choose the right file |
| **Vector search (RAG)** | Documents are cut into chunks; the most similar chunks are fetched per question | Thousands of pages, large support libraries | Chunks lose context, and misses are hard to debug |

For a small business, the numbers favor the simplest option. In its write-up on retrieval, [Anthropic notes](https://www.anthropic.com/engineering/contextual-retrieval) that if a knowledge base is smaller than 200,000 tokens, "about 500 pages of material," you can give the model the whole thing "with no need for RAG or similar methods." The policies, product facts, and FAQs of most small businesses fit in a few dozen pages.

That does not mean you should stuff everything into every request. [Chroma's context rot research](https://www.trychroma.com/research/context-rot) tested 18 leading models and found that "model performance consistently degrades with increasing input length," even on simple tasks. So the practical middle ground for a small business is files plus search: a short index the agent always sees, and topic files it opens when a task calls for them.

Vector search has a real place. With thousands of help articles, it earns its complexity. It is also where knowledge bases get hard to reason about. In Anthropic's own tests, a standard retrieval setup left the right passage out of its top 20 results 5.7% of the time, and it took several extra techniques to bring that down to 1.9%. With files, when the agent gets something wrong, you can see which file it opened and fix that file.

## How to Structure an AI Agent Knowledge Base So the Agent Uses It

A knowledge base that is technically complete can still be ignored. These habits make the difference:

- **Keep an index.** One line per file: a link plus a hook that says when to open it, such as "Refund policy: 30-day window, exceptions for bundles." The agent reads the index first. A file the index does not mention is effectively invisible
- **One topic per file, one fact in one place.** If the refund window appears in three files, it will eventually say three different things
- **Write rules with numbers.** "Refunds within 30 days of purchase" beats "about a month." "Escalate refunds over $200" beats "escalate large refunds"
- **Include the why and the edge case.** "No refunds on coaching sessions already held, because the time is spent. Exception: if we cancelled." Agents handle edge cases far better when they know the reason for the rule
- **Say what to do when it is not covered.** "If the answer is not here, say you'll check and flag it to me." Without this line, the agent fills the gap with something plausible
- **Show, don't describe.** Two real replies you liked teach tone better than "friendly but professional"
- **Date anything time-bound.** "Effective September 1, 2026" lets the agent and you spot what is out of date

Repeatable procedures deserve their own format. We covered that in [how to write AI agent skills](/blog/ai-agent-skills): a skill holds *how* to do a task, while the knowledge base holds *what is true*.

## Stale Information and Contradictions: How Knowledge Bases Go Wrong

Gaps are the failure everyone plans for. They are also the least dangerous, because an agent with no answer can say it does not know.

**Contradictions are worse.** In Chroma's tests, "even a single distractor reduces performance," where a distractor is text that looks relevant but does not answer the question. In a business, the worst distractors also give a different answer: the old pricing page, last year's FAQ, or the offhand "we can make an exception" in a note from six months ago. The agent does not know which one you meant. It picks.

The Air Canada case shows the cost. The chatbot told a grieving customer he could apply for a bereavement fare after travel; the airline's actual policy did not allow that. The tribunal held the airline responsible for what its chatbot said and ordered it to pay the fare difference. The [American Bar Association called it](https://www.americanbar.org/groups/business_law/resources/business-law-today/2024-february/bc-tribunal-confirms-companies-remain-liable-information-provided-ai-chatbot/) a reminder that companies remain liable for their AI tools.

**Staleness is the slow version of the same problem.** A document written in March is correct in March. By September it has quietly become one of two answers: the old one in the file and the new one in your head.

Three rules prevent most of this:

1. **Change a fact where it is stated.** Do not add a new file that contradicts the old one. Edit or delete the original
2. **Newer wins, explicitly.** If two sources can disagree, tell the agent which one is authoritative
3. **One owner per file.** Someone has to notice when the refund policy changes. In a one-person business, that is you, so keep the list short

## How to Keep Your Knowledge Base Current

Maintenance is where knowledge bases usually decay. The fix is to tie updates to events you already notice, not to a calendar reminder you will snooze.

| When this happens | Update this |
|---|---|
| You change a price or launch a product | The product facts file (and nothing else, if prices are looked up live) |
| You change a policy | That policy file, then delete any old copy |
| You correct the agent twice on the same thing | Add or fix the file that should have covered it |
| A customer asks something new three times | Add it to the FAQ in their wording |
| You hire someone or change who handles what | The escalation rules |

Once a month, spend 15 minutes asking the agent to summarize what it believes about your refunds, your products, and your customers. Wrong beliefs surface quickly when you read them as a list.

If your agent keeps its own memory, treat what it saves like a new hire's notes: useful, usually right, and worth skimming. Memory that writes itself still needs a reader.

## How to Test Your AI Agent Knowledge Base

Testing is the step almost everyone skips, and it takes under an hour.

![Four-step loop for testing an AI agent knowledge base: collect 20 real questions, ask in a fresh chat, grade each answer, and fix the source file rather than the answer](https://crevio.co/vite/assets/knowledge-base-test-loop-nlvijmc1.svg)

1. **Collect 20 real questions** from your inbox, DMs, and sales calls. Add three it should refuse or escalate, such as a refund far outside policy
2. **Ask each one in a fresh chat.** Earlier conversation leaks hints. Ask the agent where the answer came from
3. **Grade each answer**: right, wrong, made up, or "I don't know." Count "I don't know" as a pass when the knowledge base genuinely does not cover the question
4. **Fix the source, not the answer.** Correcting the agent in the chat fixes one conversation. Correcting the file fixes every conversation after it

Then run two targeted checks. **The contradiction test:** change a policy, then ask about it. If the agent quotes the old version, a copy is hiding somewhere. **The gap test:** ask something you know is not covered. The right answer is a clear "I'll need to check," not a confident guess.

Save the 20 questions. Re-run them after every significant change. It is the closest thing a small business has to a regression test.

## AI Knowledge Bases for Customer Support Agents

Much of what is written about AI knowledge bases is about support: an AI agent that answers customers from your help center. The principles above apply, with higher stakes.

The economics are why businesses try it. [Gartner's benchmarking](https://www.gartner.com/en/documents/5164231) puts the median cost of a self-service contact at $1.84, against $13.50 for an assisted one. That saving only holds if the answers are right, because customers act on them without checking.

What changes for a customer-facing agent:

- **Write for the question, not the product page.** Customers ask "can I get my money back if I haven't started the course?", not "refund policy"
- **Escalation rules become the most important file.** Decide in advance what the agent must never promise: refunds outside policy, delivery dates, legal or medical advice
- **Read the unresolved conversations weekly.** Every question the agent could not answer is a missing file
- **Draft before send while you build trust.** For anything involving money, have the agent draft and a person approve. We laid out which actions deserve that in [human in the loop AI agents](/blog/human-in-the-loop-ai-agents)

If you run a high volume of customer questions, dedicated support tools such as [Intercom's Fin](https://fin.ai/) and [Zendesk](https://www.zendesk.com/) are built for exactly this job, with help-center syncing and resolution reporting.

## How Crevio's Agent Stores Business Knowledge

[Crevio](https://crevio.co/) is an [AI business builder](/blog/ai-business-builder): you describe what you want to sell, and its agent builds the store, runs payments, writes the marketing, and keeps working on the business with you. Everything above shaped how that agent keeps what it knows about your business.

![Crevio homepage: AI that builds your store, processes your payments, writes your marketing, and grows your sales](https://crevio.co/vite/assets/crevio-homepage-k3m1mvi1.png)

Here is how it works today:

- **Memory is plain files, in two scopes.** Business memory (how the business works) is shared with everyone on your team. Personal memory (your own preferences) stays private to you. Each has a short index that goes into every chat, capped at 6,000 characters for the business and 2,000 for personal notes. Topic files are opened only when a task needs them
- **It saves when it should.** The agent writes a memory when you ask it to remember something, when you correct it, or when you state a lasting fact such as a policy, number, or rule. It does not save passwords, and it does not copy things Crevio already records
- **Live data stays live.** Products, orders, customers, and your site are looked up directly, not remembered, so a price change never leaves a stale copy behind
- **It tidies itself.** Periodically, at most once a day, a cleanup pass merges duplicate memories, resolves contradictions in favor of the newer fact, and removes what is stale. It takes a snapshot first
- **No vector search.** The agent searches files on its own computer and your past chats as plain text, then reads what it finds. You can attach files to a chat to give it documents to work from

You can also shape knowledge directly. Each bot you create has a Persona field that is added to its instructions, which is the right home for tone and standing rules:

![Crevio bot settings with Name, Handle, Role, and a Persona field that is added to the bot's instructions](https://crevio.co/vite/assets/crevio-bot-persona-field-d2jyuuyk.png)

Procedures go in skills, which you can install, upload, or let the agent write from experience. And if the agent has learned something wrong, Settings → Advanced has a reset:

![Crevio's AI Workspace Knowledge setting with a Reset AI knowledge button that wipes accumulated notes without touching the storefront](https://crevio.co/vite/assets/crevio-reset-ai-knowledge-nkms8zwb.png)

The honest caveats: there is no dedicated page for browsing and editing memory yet, so you review it by asking the agent in chat what it knows. Reset is all-or-nothing. And Crevio is not a customer-facing support bot; its agent works for you, not in a help widget on your site. Crevio has a free plan with 20 AI credits a month, so you can tell it your refund policy today and check whether it still knows it next week.

## What Nobody Tells You

**Writing it exposes your own contradictions.** The first draft of a refund policy for an agent usually reveals that you have been handling refunds three different ways. That is a business problem the agent surfaced, not an AI problem.

**Agents follow written rules more literally than staff do.** A human reads "no refunds after 30 days" and still makes an exception for a customer whose card was double-charged. An agent needs the exception written down.

**Deleting is the highest-value edit.** Most knowledge base improvements after the first month are removals: the old FAQ, the duplicated policy, the note that is no longer true. Less text with no contradictions beats more text with a few.

**The knowledge base is the difference between an agent and an assistant.** A correction that still holds a month later, after other work has happened, is the test we used in [AI employee vs AI agent](/blog/ai-employee-vs-ai-agent). Without durable knowledge, every session starts from zero.

## AI Agent Knowledge Base FAQ

### What is an AI agent knowledge base?

An AI agent knowledge base is the written business knowledge an agent reads before it acts: policies, product facts, FAQs, procedures, and remembered preferences. It lets a general AI model answer and act the way your business would, instead of guessing from general knowledge.

### Do I need RAG or a vector database for my AI agent?

Probably not, if you are a small business. Anthropic suggests that knowledge under roughly 500 pages can be given to a model directly. A short index plus topic files the agent opens on demand is simpler, cheaper, and easier to fix. RAG makes sense at thousands of documents.

### What is the difference between agent memory and a knowledge base?

A knowledge base is what you write deliberately: policies, product facts, procedures. Memory is what the agent picks up while working with you, like preferences and corrections. Both need the same care: one fact in one place, and old versions removed.

### How often should I update my AI agent knowledge base?

Update it when something changes, not on a schedule: a new price, a changed policy, or a correction you have made twice. Add a monthly 15-minute review where you ask the agent to summarize what it believes about your products and policies.

## The Short Version

An AI agent knowledge base is only as good as its least accurate file. Keep it small, put each fact in one place, let live systems answer live questions, and test it with the questions your customers actually ask. The agent will follow what you wrote down, so make sure it is still true.

## Related Blog Posts

- [AI Agent Skills: What They Are and How to Write One That Works](/blog/ai-agent-skills)
- [Human in the Loop AI Agents: What to Approve and What to Let Run](/blog/human-in-the-loop-ai-agents)
- [AI Employee vs AI Agent: The Difference That Actually Matters](/blog/ai-employee-vs-ai-agent)
- [AI Agents for Small Business: Which Lanes to Hand Over First](/blog/ai-agents-for-small-business)
