Upload Files
Uploading files is the fastest way to give your agent structured knowledge. Manuals, policy documents, product catalogues, help articles, and data exports can all be uploaded and indexed in minutes.
Supported file formats
| Format | Extensions | Notes |
|---|---|---|
| Text-layer PDFs only. Scanned image PDFs (no text layer) will produce poor results — use a text-based PDF or export from the source application. | ||
| Word Document | .docx, .doc | Headings, paragraphs, and tables are all extracted. Embedded images are skipped. |
| Plain Text | .txt | Raw text files. Simplest format — no formatting overhead. |
| CSV | .csv | Tabular data. Each row is treated as a searchable unit. Best for product lists, FAQs in spreadsheet form, or data tables. |
| Excel Spreadsheet | .xlsx, .xls | All sheets are extracted and indexed. Cell values are read; charts and images are skipped. |
How file processing works
When you upload a file, it goes through a three-stage pipeline automatically:
Upload and extraction
Chunking and embedding
Indexed and live
Status indicators
Each file in your knowledge base shows a status badge:
- •Pending — file is queued, processing has not yet started.
- •Processing — chunks are being created and embedded. Large files may take a minute or two.
- •Completed — file is fully indexed and live. You can see how many chunks and tokens were generated.
- •Failed — processing encountered an error. The most common cause is a scanned PDF with no text layer, a password-protected file, or a corrupted document.
Uploading files
Open your agent and go to Knowledge Base
Upload your files
Monitor status
Updating and deleting files
To update a file, delete the existing version and re-upload the new one. When a file is deleted, all its vector embeddings are immediately removed from the index — the agent stops drawing on that content at once.
Tips for the best results
- •Use text-based PDFs, not scanned image PDFs. Export from the original application (Word, Google Docs) rather than scanning a printout.
- •Add clear headings and section titles. The chunking algorithm respects headings, so well-structured documents produce more accurate retrieval.
- •Keep files focused on a single topic where possible. A 10-page product FAQ retrieves better than a 200-page company handbook where every topic is mixed together.
- •For very large documents (100+ pages), consider splitting them into topic-specific files — the agent will retrieve only the relevant chunk regardless, but smaller files make status tracking easier.
- •CSV files work well for structured data like pricing tables, product attributes, or FAQ lists — one question/answer per row.
Was this page helpful?
