Skip to content

Vector

The Vector module stores and searches embeddings for retrieval-augmented generation (RAG). It has two independent kinds of provider: stores (Pinecone, Amazon S3 Vectors) hold and search vectors, and embedding providers (OpenAI, Google Gemini, Ollama) turn text into vectors. Agents in the AI module use its operations as tools, so the agent decides when to retrieve; there is no fixed retrieval pipeline.

Install: Pinecone and every embedding provider need nothing extra. Amazon S3 Vectors needs boto3:

Terminal window
pip install "kirak[vector-s3vectors]"

kirak[all] includes it.

Enable:

{
"modules": ["vector"],
"vector": {
"store": { "default_provider": "pine", "providers": { "pine": { "type": "pinecone" } } },
"embedding": { "default_provider": "openai", "providers": { "openai": { "type": "openai", "model": "text-embedding-3-small" } } }
}
}
KIRAK_VECTOR_STORE_PINE_API_KEY=...
KIRAK_VECTOR_EMBEDDING_OPENAI_API_KEY=sk-...

Either section may be left out: with pre-computed vectors you need no embedding providers, and with embed() alone you need no store. An enabled module with neither fails at startup. Every setting is listed in the configuration reference.


Each section lists named instances, like the other modules’ providers. A call uses the section’s default_provider unless it passes "provider" (the store instance) or "embedding_provider" (the embedding instance for upsert/search; embed/embed_batch take "provider"). Only instances listed in kirak.json can be used; anything else fails with PROVIDER_NOT_CONFIGURED (400).

Store type Notes
pinecone Serverless indexes. Metrics cosine, euclidean, dotproduct. Namespaces supported.
s3_vectors One instance = one existing vector bucket (create it in AWS first). Metrics cosine and euclidean only. No namespaces and no delete by filter (both return 400). Available in a limited set of AWS regions.
Embedding type Notes
openai /v1/embeddings; base_url for OpenAI-compatible servers; optional dimensions.
google Gemini batchEmbedContents; queries and documents get the retrieval task types. Optional dimensions.
ollama A local or self-hosted Ollama server; no secrets.

Embedding providers split large inputs into batches and retry rate limits, server errors and connection failures with exponential backoff (max_retries, default 3).


Every operation takes a params dict and returns the standard envelope. Index management and writes (create_index, delete_index, upsert, delete) need an admin, system or superadmin caller; the others need any authenticated caller. Guests are refused (401).

await kirak.vector.create_index({"name": "docs", "dimensions": 1536, "metric": "cosine"})
# -> data: {"name": "docs", "dimensions": 1536, "metric": "cosine", ...}

metric defaults to cosine. Pinecone waits until the new index is ready (ready_timeout). An existing name fails with VECTOR_INDEX_EXISTS (409).

delete_index, list_indexes, describe_index

Section titled “delete_index, list_indexes, describe_index”
await kirak.vector.delete_index({"name": "docs"})
await kirak.vector.list_indexes({}) # data: {"indexes": [...]}
await kirak.vector.describe_index({"name": "docs"}) # data: {"name", "dimensions", "metric", ...}

An unknown index fails with VECTOR_INDEX_NOT_FOUND (404).

await kirak.vector.upsert({
"index": "docs",
"vectors": [
{"id": "faq-1", "text": "Refunds are issued within 5 days.", "metadata": {"source": "faq"}},
{"id": "faq-2", "values": [0.12, -0.03, ...]},
],
"namespace": "tenant-42", # optional (Pinecone only)
})
# -> data: {"upserted": 2}

Each item needs an id and exactly one of values (a vector) or text (embedded first, all texts in one batch, with the embedding provider). metadata is optional.

The operation agents use to retrieve.

result = await kirak.vector.search({
"index": "docs",
"text": "How long do refunds take?", # or "vector": [...]
"top_k": 5, # default 10
"filter": {"source": {"$eq": "faq"}}, # optional metadata filter, passed to the store
})
# -> data: {"matches": [{"id": "faq-1", "score": 0.91, "metadata": {"source": "faq"}}, ...]}

Pass exactly one of text or vector. Matches come most similar first. For cosine, score is the cosine similarity on both stores; S3 Vectors matches also carry the raw distance.

await kirak.vector.fetch({"index": "docs", "ids": ["faq-1", "faq-2"]}) # data: {"vectors": [...]}
await kirak.vector.delete({"index": "docs", "ids": ["faq-2"]})
await kirak.vector.delete({"index": "docs", "filter": {"source": {"$eq": "old"}}}) # Pinecone only

delete takes exactly one of ids or filter, so a call can never clear an index by accident.

Embeddings without storing them.

await kirak.vector.embed({"text": "hello"}) # data: {"embedding": [...], "dimensions": 1536}
await kirak.vector.embed_batch({"texts": ["a", "b"]}) # data: {"embeddings": [[...], [...]], "dimensions": 1536}

embed embeds a search query; embed_batch embeds documents. Some models treat the two differently.

Before calling the store, upsert and search compare every vector with the index’s declared dimension and fail with VECTOR_DIMENSION_MISMATCH (400) on a mismatch, instead of an opaque provider error. The dimension is read once per index and cached.

Errors: MISSING_PARAMS / INVALID_PARAMS (400), PROVIDER_NOT_CONFIGURED (400), VECTOR_DIMENSION_MISMATCH (400), VECTOR_INDEX_NOT_FOUND (404), VECTOR_INDEX_EXISTS (409), VECTOR_STORE_ERROR (500, or 400 for a request the store rejects as invalid), EMBEDDING_ERROR (500).


There are two ways to give an agent retrieval. Both run as the user who started the agent run, under the access rules above, including when the agent runs from a script or scheduled job.

A tool in the agent file is the shortest. The key vector.search names the operation:

{
"name": "support",
"model": "openai:gpt-4o-mini",
"instructions": "Answer from the knowledge base. Search the 'docs' index before answering.",
"tools": {
"vector.search": {
"parameters": {
"type": "object",
"properties": {
"index": {"type": "string"},
"text": {"type": "string"},
"top_k": {"type": "integer"}
},
"required": ["index", "text"]
}
}
}
}

Optional parameters the model leaves out (here top_k) fall back to the operation’s defaults. The tool returns the full search result, metadata included. Without parameters an agent’s model gets an open schema and has to infer the arguments, so always declare them.

An agent-scoped Python tool gives more control: it can fix arguments the model should not choose, such as the index, and trim the results to what the model needs:

from pydantic_ai import RunContext
from kirak.ai import KirakDeps
async def on_kirak_ready(kirak):
@kirak.agent("support").tool
async def search_knowledge_base(ctx: RunContext[KirakDeps], question: str, top_k: int = 5) -> list:
"""Search the knowledge base. Returns the most relevant passages with their source."""
result = await kirak.vector.search({"index": "docs", "text": question, "top_k": top_k})
return [m["metadata"] for m in result["data"]["matches"]]

pydantic-ai builds this tool’s name and schema from the function signature. examples/17_vector_rag.py is a complete app using this form.

Scoping data per user or tenant is the application’s job: the access rules above let any signed-in caller search every vector in an index. Store the tenant in each vector’s metadata (or use one namespace per tenant) and restrict every search to it. With a tool in the agent file the model chooses the arguments, so force the filter in a before_search hook, which replaces whatever the model sent:

from kirak.core.context import get_user_context
from kirak.core.exceptions import PermissionDenied
@kirak.vector.hook("before_search")
def only_this_tenant(params):
user = get_user_context()
tenant = tenant_for(user["user_id"]) if user else None # your app's lookup
if tenant is None:
raise PermissionDenied("No tenant for this search")
return {**params, "filter": {"tenant": {"$eq": tenant}}}

tenant_for stands for the app’s own mapping from user to tenant: the caller Kirak passes carries only user_id, email and role. Keep the hook sync. A sync hook that raises stops the search; an async hook that raises is skipped and the search runs with the original, unfiltered params. Outside a request (a script or scheduled job) get_user_context() returns None in hooks, so this hook refuses those searches. When the lookup needs the database, scope in an agent-scoped Python tool instead: it can await the lookup, reads the caller from ctx.deps.current_user, and passes the filter itself.


Every operation runs before_<operation> and after_<operation> hooks, e.g. before_upsert, after_search, before_embed. A before_ hook receives the params and may return changed params; an after_ hook receives the result.

@kirak.vector.hook("after_upsert")
async def log_upsert(result):
logger.info("Upserted %s vectors", result["data"]["upserted"])

Unlike most modules, Vector mounts no routes; its operations are Python methods only. That is deliberate:

  • Index management (create_index, delete_index) is infrastructure, like creating or dropping a table: run it from a deploy script, on_kirak_ready or an admin tool, not through a public API.
  • embed over HTTP would let any logged-in user spend the app’s embedding quota.
  • Raw search and fetch would expose every stored document to every logged-in user unless the app scoped them. Retrieval normally happens inside an agent tool, in-process.
  • list_indexes / describe_index return infrastructure details.

When an app does want an endpoint, it writes a route with its own authorization and scoping and calls the module from it. The access rules above still apply, with the caller taken from the request:

from fastapi import Body, Request
from kirak.core.context import set_request_context
@app.post("/api/kb/documents")
async def add_documents(request: Request, documents: list = Body(...)):
set_request_context(request) # lets the Vector module identify the caller
return await app.get_kirak().vector.upsert({
"index": "docs",
"vectors": [{"id": d["id"], "text": d["text"], "metadata": {"text": d["text"]}} for d in documents],
})

examples/17_vector_rag.py and examples/19_document_ingestion_recipe.py have complete routes.


Kirak’s job starts at text: upsert embeds and stores it. Getting text out of PDFs, Word files or web pages, and splitting it into chunks, is left to the application, because the right choices depend on the documents and the libraries are heavy. Popular options: Unstructured, Docling, LlamaIndex readers and LangChain document loaders. Extract and chunk with one of them, then call upsert with text items; examples/19_document_ingestion_recipe.py does this with Unstructured.


The shared contract, the credential check (check()), entry points and testing are covered in Adding a Provider.

Subclass VectorStoreProvider or EmbeddingProvider (kirak/vector/providers/base.py) and register it for its category, "store" or "embedding":

from kirak.vector.providers.base import EmbeddingProvider
class CohereEmbedding(EmbeddingProvider):
TYPE_NAME = "cohere"
SECRET_FIELDS = ("api_key",) # read from KIRAK_VECTOR_EMBEDDING_<INSTANCE>_API_KEY
async def embed_query(self, text): ...
async def embed_documents(self, texts): ...
async def on_kirak_ready(kirak):
kirak.vector.register_provider("embedding", "cohere", CohereEmbedding)

A store provider implements create_index, delete_index, list_indexes, describe_index, upsert, delete, fetch and query. Providers can also be published as pip packages with an entry point in the kirak.vector_store_providers or kirak.vector_embedding_providers group. examples/18_custom_vector_provider.py is a complete in-memory store.