RAG
Let your agent search an existing knowledge base with Pinecone, Weaviate, or Qdrant Cloud.
The RAG tool lets your agent answer questions using a knowledge base you already maintain. It sends a text query to your provider and returns matching passages and sources. Your provider handles query vectorization; your existing pipeline continues to manage documents and ingestion.
Choose a provider
Select the provider in packages/agent/tools/rag/provider/index.ts, just like the payment provider:
export * from "./pinecone";Replace "./pinecone" with "./weaviate" or "./qdrant" to switch. Keep one export active. Each provider implements the same SearchKnowledgeBase function type from rag/types.ts and reads its own environment variables when called.
Match the stored passage field first
Setting textField correctly is required for retrieval to work. Credentials alone are not enough. Before enabling the tool, inspect a document in your provider's console or a returned search record and find the field containing the actual passage text. Set textField in packages/agent/tools/rag/config.ts to that exact, case-sensitive name.
There is no universal passage field name. content in the example below is only an example, not a provider standard or an automatically detected value:
| Stored passage field | Required setting |
|---|---|
content | textField: "content" |
text | textField: "text" |
chunk_text | textField: "chunk_text" |
For Pinecone, use the field name inside a search hit's fields; for Weaviate, use the collection property name; for Qdrant, use the payload field name. The configured field must contain a non-empty string in every returned match. An ID, embedding vector, or source URL does not replace the passage text.
The adapter uses this mapping to read the text passed to the agent. A wrong mapping can make the tool fail with missing or invalid text field even when authentication succeeds and the provider returns matching records. One match with missing, blank, or non-string passage text fails the entire search. Recheck the mapping whenever you switch providers, indexes, collections, or ingestion formats.
Enable the tool
Add your provider's connection settings to the root .env.local, then open packages/agent/tools/rag/config.ts:
export const config: RagConfig = {
// Enable after choosing a provider, setting credentials, and matching stored fields.
enabled: true,
limit: 5,
// Required: replace with the exact field containing your stored passage text.
textField: "content",
sourceField: "source",
metadataFields: [],
};The tool ships with enabled: false. Set it to true after configuring your provider and matching the stored passage field, then restart the API. The tool is already wired into the agent in apps/api/src/index.ts.
Set sourceField to a source string field, such as a URL, or remove it if your documents have no source. metadataFields lists any additional stored fields you want returned. limit defaults to 5 and accepts values from 1 to 10.
Use a reference corpus that all users admitted to this agent are allowed to search. Selecting an organization in chat does not automatically filter the external knowledge base. Private per-user or per-organization retrieval needs your own trusted server-side authorization and scoping.
Pinecone
Use an existing index with integrated embedding. Copy its index host and API key into .env.local:
PINECONE_INDEX_HOST="https://your-index-host"
PINECONE_API_KEY="your-api-key"
PINECONE_NAMESPACE="__default__"Use the namespace containing your documents; it defaults to __default__ when omitted. If your passages are stored in chunk_text, set textField: "chunk_text" in rag/config.ts.
Pinecone's integrated embedding field mapping identifies the stored field used for embedding. For example, field_map: { "text": "chunk_text" } embeds the document's chunk_text field; set textField: "chunk_text" to retrieve that passage. If your records use text, set textField: "text" instead.
The provider sends your question as query.inputs.text to Pinecone's record-search API. That query input stays named text, independently of your stored passage field. The adapter explicitly requests configured fields in the response, so requesting content will not retrieve a passage stored as text. Pinecone's search documentation describes this fields selection; omitting it in a diagnostic request returns all fields so you can inspect their names.
Weaviate
Use an existing collection with a compatible server-side vectorizer:
WEAVIATE_URL="https://your-weaviate-instance"
WEAVIATE_API_KEY="your-api-key"
WEAVIATE_COLLECTION="Document"
# Optional, for named vectors and multi-tenant collections:
WEAVIATE_TARGET_VECTOR=""
WEAVIATE_TENANT=""The provider uses nearText to search that collection. Match textField, sourceField, and metadataFields to your collection's property names.
If your existing OpenAI vectorizer needs credentials at query time, set WEAVIATE_OPENAI_API_KEY. The provider forwards it as X-OpenAI-Api-Key. For another vectorizer, add its required X-... header in rag/provider/weaviate/index.ts, reading its value from a server environment variable. No additional header is needed when credentials are already configured on the Weaviate server.
Qdrant Cloud
Use an existing collection on Qdrant Cloud with Cloud Inference:
QDRANT_URL="https://your-qdrant-cluster:6333"
QDRANT_API_KEY="your-api-key"
QDRANT_COLLECTION="documents"
QDRANT_MODEL="sentence-transformers/all-minilm-l6-v2"
# Optional, if your collection uses named vectors:
QDRANT_VECTOR=""The model above is an example. Use the compatible model identifier from your existing setup and check that it is available in your cluster's Inference tab. The provider sends text and this model identifier to Qdrant; Cloud Inference handles vectorization and search in one request. This integration targets Cloud Inference, not self-hosted vector-only collections.
Try it in chat
Enable the AI Agent, restart the API, and ask a question covered by your documents:
Search our knowledge base for the refund policy and include the source in your answer.
The agent calls search_knowledge_base with a text query. Its result looks like this:
{
"results": [
{
"id": "refund-policy",
"text": "Refunds are processed within five business days.",
"source": "https://your-product.com/help/refunds"
}
]
}The agent uses those passages to write its answer. Results preserve provider order and include selected metadata; scores and distances remain provider-specific. Tool activity and saved results use the existing chat experience. This native agent tool is not exposed through the MCP Server.
Troubleshooting
- Tool unavailable: check
enabled: trueand restart the API. - Configuration error: check the named setting and your selected provider's environment variables. Other providers need no credentials.
- HTTP 401 or 403: check the API key and its access to the configured target.
- HTTP 400 or GraphQL error: confirm text search works in the existing service. Check the Weaviate vectorizer, collection and property names, or the Qdrant model and named vector.
- No matches: an empty
resultslist is a successful search. Check your configured namespace, collection or tenant, then try a question you know the documents answer. - Missing or invalid text field: set
textFieldto the stored string field containing the passages. Missing text is an error, not silently discarded. - Response or result too large: reduce
limitor return fewer metadata fields. Responses and serialized results are limited to 512 KiB. - Request timed out: check provider availability and latency. Requests have a 20-second deadline and are not automatically retried.
Questions are limited to 4,000 characters. Cancelling a run cancels its retrieval request. To disable retrieval, set enabled to false and restart the API; your remote knowledge base and saved conversations remain available.