Skip to content

Knowledge bases

A knowledge base is a collection of documents an agent can search while it answers. Manage them under Build → Knowledge Base.

  1. Open Knowledge Base and create a new one.
  2. Enter a Name and a Description, then select Create Knowledge Base.
  3. Open the knowledge base again to upload documents. You can only upload after the knowledge base has been saved.
  4. Attach it to an agent in the visual canvas.
  5. Test the agent in chat with questions the documents answer.
Field What it does
Name Required. Letters, numbers, spaces, underscores and hyphens, up to 100 characters.
Description What the knowledge base contains. The agent reads it, together with the name, to decide when to search this knowledge base, so be specific.
Format What is indexed
PDF (.pdf) The text of each page.
Word (.docx) The document text, including tables.
Excel (.xlsx, .xls) Every sheet, one row per line.
CSV (.csv) Every row.
PowerPoint (.pptx) The text and tables of each slide.

Drop files onto the upload area or select Select Files. You can pick several at once; they are uploaded one after another, and any file with another extension is skipped.

Each document shows its size, its page count, and a status that updates on its own while the page is open:

Status Meaning
uploading The file is being received.
processing The text is being extracted, split into chunks and indexed.
completed The document is searchable.
failed Processing did not finish. The reason is shown next to the file.

An agent only searches documents that have completed. A knowledge base with no completed documents is not offered to the agent at all.

When a knowledge base is attached, the agent gets a search it can run against it. On each turn the model decides whether to search, using the knowledge base’s name and description. The search returns the passages closest in meaning to the query, and the agent writes its answer from them, citing each passage inline as [1], [2] and so on.

Because the model decides when to search, a vague description leads to missed searches. Say what is inside and when it applies, for example “Refund, exchange and warranty policy for retail customers”.

Expand Advanced RAG Configuration on the knowledge base to tune how documents are split and searched. The defaults suit most documents.

Setting Default Range What it does
Chunk Size 2000 500–8000 Characters per chunk when a document is indexed. Larger chunks carry more context each.
Chunk Overlap 200 0–2000 Characters shared by consecutive chunks, so a sentence split at a boundary keeps its context.
Results per Query 5 1–20 How many chunks each search returns to the agent.
Score Threshold 0.00 0–1 Minimum similarity for a chunk to be returned. 0 means no filter.
Search Strategy Similarity Similarity returns the most relevant chunks first. MMR (Diverse) favours results that are relevant but not repetitive.
MMR Fetch K 20 5–100 MMR only. How many candidates are considered before re-ranking.
Diversity (λ) 0.50 0–1 MMR only. 0 favours diversity, 1 favours relevance.

Retrieval settings apply to the next search. Chunk Size and Chunk Overlap apply only to documents uploaded after you change them; documents already indexed keep the chunks they were created with. To re-index an existing document with new chunk settings, delete it and upload it again.

The embedding model that turns text into searchable vectors is a workspace setting (EMBEDDING_PROVIDER and EMBEDDING_MODEL in Workspace settings). Changing it requires every document to be processed again.

  • Update a document: delete the old file and upload the new version. Uploading a second copy without deleting the first leaves both searchable, and the agent may quote the outdated one.
  • Delete a document: select the × next to it. Its passages stop being searchable.
  • Delete a knowledge base: from the Knowledge Bases list. All of its documents are removed.
  • Keep content clean, specific and current. An outdated document produces confident, outdated answers.
  • Prefer several small, topic-specific knowledge bases over one large mixed one; the agent chooses between them by name and description.
  • Split very large files into topic-sized documents. They upload more reliably and retrieve more precisely.
  • When an answer is wrong, check first whether the source document is wrong or missing.
  • Protect the important answers with an evaluation. The groundedness scorer checks that answers are supported by what the agent actually retrieved.