Vector
The Vector module stores and searches embeddings for retrieval-augmented generation (RAG). It has two independent kinds of provider: stores (Pinecone, Amazon S3 Vectors) hold and search vectors, and embedding providers (OpenAI, Google Gemini, Ollama) turn text into vectors. Agents in the AI module use its operations as tools, so the agent decides when to retrieve; there is no fixed retrieval pipeline.
Install: Pinecone and every embedding provider need nothing extra. Amazon S3 Vectors needs boto3:
pip install "kirak[vector-s3vectors]"kirak[all] includes it.
Enable:
{ "modules": ["vector"], "vector": { "store": { "default_provider": "pine", "providers": { "pine": { "type": "pinecone" } } }, "embedding": { "default_provider": "openai", "providers": { "openai": { "type": "openai", "model": "text-embedding-3-small" } } } }}KIRAK_VECTOR_STORE_PINE_API_KEY=...KIRAK_VECTOR_EMBEDDING_OPENAI_API_KEY=sk-...Either section may be left out: with pre-computed vectors you need no embedding providers, and with
embed() alone you need no store. An enabled module with neither fails at startup. Every setting is
listed in the configuration reference.
Providers
Section titled “Providers”Each section lists named instances, like the other modules’ providers. A call uses the section’s
default_provider unless it passes "provider" (the store instance) or "embedding_provider" (the
embedding instance for upsert/search; embed/embed_batch take "provider"). Only instances
listed in kirak.json can be used; anything else fails with PROVIDER_NOT_CONFIGURED (400).
Store type |
Notes |
|---|---|
pinecone |
Serverless indexes. Metrics cosine, euclidean, dotproduct. Namespaces supported. |
s3_vectors |
One instance = one existing vector bucket (create it in AWS first). Metrics cosine and euclidean only. No namespaces and no delete by filter (both return 400). Available in a limited set of AWS regions. |
Embedding type |
Notes |
|---|---|
openai |
/v1/embeddings; base_url for OpenAI-compatible servers; optional dimensions. |
google |
Gemini batchEmbedContents; queries and documents get the retrieval task types. Optional dimensions. |
ollama |
A local or self-hosted Ollama server; no secrets. |
Embedding providers split large inputs into batches and retry rate limits, server errors and
connection failures with exponential backoff (max_retries, default 3).
Operations
Section titled “Operations”Every operation takes a params dict and returns the standard envelope. Index management and writes
(create_index, delete_index, upsert, delete) need an admin, system or superadmin caller; the others
need any authenticated caller. Guests are refused (401).
create_index
Section titled “create_index”await kirak.vector.create_index({"name": "docs", "dimensions": 1536, "metric": "cosine"})# -> data: {"name": "docs", "dimensions": 1536, "metric": "cosine", ...}metric defaults to cosine. Pinecone waits until the new index is ready (ready_timeout).
An existing name fails with VECTOR_INDEX_EXISTS (409).
delete_index, list_indexes, describe_index
Section titled “delete_index, list_indexes, describe_index”await kirak.vector.delete_index({"name": "docs"})await kirak.vector.list_indexes({}) # data: {"indexes": [...]}await kirak.vector.describe_index({"name": "docs"}) # data: {"name", "dimensions", "metric", ...}An unknown index fails with VECTOR_INDEX_NOT_FOUND (404).
upsert
Section titled “upsert”await kirak.vector.upsert({ "index": "docs", "vectors": [ {"id": "faq-1", "text": "Refunds are issued within 5 days.", "metadata": {"source": "faq"}}, {"id": "faq-2", "values": [0.12, -0.03, ...]}, ], "namespace": "tenant-42", # optional (Pinecone only)})# -> data: {"upserted": 2}Each item needs an id and exactly one of values (a vector) or text (embedded first, all texts in
one batch, with the embedding provider). metadata is optional.
search
Section titled “search”The operation agents use to retrieve.
result = await kirak.vector.search({ "index": "docs", "text": "How long do refunds take?", # or "vector": [...] "top_k": 5, # default 10 "filter": {"source": {"$eq": "faq"}}, # optional metadata filter, passed to the store})# -> data: {"matches": [{"id": "faq-1", "score": 0.91, "metadata": {"source": "faq"}}, ...]}Pass exactly one of text or vector. Matches come most similar first. For cosine, score is the
cosine similarity on both stores; S3 Vectors matches also carry the raw distance.
fetch and delete
Section titled “fetch and delete”await kirak.vector.fetch({"index": "docs", "ids": ["faq-1", "faq-2"]}) # data: {"vectors": [...]}await kirak.vector.delete({"index": "docs", "ids": ["faq-2"]})await kirak.vector.delete({"index": "docs", "filter": {"source": {"$eq": "old"}}}) # Pinecone onlydelete takes exactly one of ids or filter, so a call can never clear an index by accident.
embed and embed_batch
Section titled “embed and embed_batch”Embeddings without storing them.
await kirak.vector.embed({"text": "hello"}) # data: {"embedding": [...], "dimensions": 1536}await kirak.vector.embed_batch({"texts": ["a", "b"]}) # data: {"embeddings": [[...], [...]], "dimensions": 1536}embed embeds a search query; embed_batch embeds documents. Some models treat the two differently.
Dimension check
Section titled “Dimension check”Before calling the store, upsert and search compare every vector with the index’s declared
dimension and fail with VECTOR_DIMENSION_MISMATCH (400) on a mismatch, instead of an opaque provider
error. The dimension is read once per index and cached.
Errors: MISSING_PARAMS / INVALID_PARAMS (400), PROVIDER_NOT_CONFIGURED (400),
VECTOR_DIMENSION_MISMATCH (400), VECTOR_INDEX_NOT_FOUND (404), VECTOR_INDEX_EXISTS (409),
VECTOR_STORE_ERROR (500, or 400 for a request the store rejects as invalid), EMBEDDING_ERROR (500).
Using it from an agent
Section titled “Using it from an agent”There are two ways to give an agent retrieval. Both run as the user who started the agent run, under the access rules above, including when the agent runs from a script or scheduled job.
A tool in the agent file is the shortest. The key vector.search names the operation:
{ "name": "support", "model": "openai:gpt-4o-mini", "instructions": "Answer from the knowledge base. Search the 'docs' index before answering.", "tools": { "vector.search": { "parameters": { "type": "object", "properties": { "index": {"type": "string"}, "text": {"type": "string"}, "top_k": {"type": "integer"} }, "required": ["index", "text"] } } }}Optional parameters the model leaves out (here top_k) fall back to the operation’s defaults. The
tool returns the full search result, metadata included. Without parameters an agent’s model gets an
open schema and has to infer the arguments, so always declare them.
An agent-scoped Python tool gives more control: it can fix arguments the model should not choose, such as the index, and trim the results to what the model needs:
from pydantic_ai import RunContextfrom kirak.ai import KirakDeps
async def on_kirak_ready(kirak): @kirak.agent("support").tool async def search_knowledge_base(ctx: RunContext[KirakDeps], question: str, top_k: int = 5) -> list: """Search the knowledge base. Returns the most relevant passages with their source.""" result = await kirak.vector.search({"index": "docs", "text": question, "top_k": top_k}) return [m["metadata"] for m in result["data"]["matches"]]pydantic-ai builds this tool’s name and schema from the function signature. examples/17_vector_rag.py
is a complete app using this form.
Scoping data per user or tenant is the application’s job: the access rules above let any
signed-in caller search every vector in an index. Store the tenant in each vector’s metadata (or
use one namespace per tenant) and restrict every search to it. With a tool in the agent file the
model chooses the arguments, so force the filter in a before_search hook, which replaces whatever
the model sent:
from kirak.core.context import get_user_contextfrom kirak.core.exceptions import PermissionDenied
@kirak.vector.hook("before_search")def only_this_tenant(params): user = get_user_context() tenant = tenant_for(user["user_id"]) if user else None # your app's lookup if tenant is None: raise PermissionDenied("No tenant for this search") return {**params, "filter": {"tenant": {"$eq": tenant}}}tenant_for stands for the app’s own mapping from user to tenant: the caller Kirak passes carries
only user_id, email and role. Keep the hook sync. A sync hook that raises stops the search; an
async hook that raises is skipped and the search runs with the original, unfiltered params. Outside
a request (a script or scheduled job) get_user_context() returns None in hooks, so this hook
refuses those searches. When the lookup needs the database, scope in an agent-scoped Python tool
instead: it can await the lookup, reads the caller from ctx.deps.current_user, and passes the
filter itself.
Every operation runs before_<operation> and after_<operation> hooks, e.g. before_upsert,
after_search, before_embed. A before_ hook receives the params and may return changed params; an
after_ hook receives the result.
@kirak.vector.hook("after_upsert")async def log_upsert(result): logger.info("Upserted %s vectors", result["data"]["upserted"])No HTTP endpoints
Section titled “No HTTP endpoints”Unlike most modules, Vector mounts no routes; its operations are Python methods only. That is deliberate:
- Index management (
create_index,delete_index) is infrastructure, like creating or dropping a table: run it from a deploy script,on_kirak_readyor an admin tool, not through a public API. embedover HTTP would let any logged-in user spend the app’s embedding quota.- Raw
searchandfetchwould expose every stored document to every logged-in user unless the app scoped them. Retrieval normally happens inside an agent tool, in-process. list_indexes/describe_indexreturn infrastructure details.
When an app does want an endpoint, it writes a route with its own authorization and scoping and calls the module from it. The access rules above still apply, with the caller taken from the request:
from fastapi import Body, Requestfrom kirak.core.context import set_request_context
@app.post("/api/kb/documents")async def add_documents(request: Request, documents: list = Body(...)): set_request_context(request) # lets the Vector module identify the caller return await app.get_kirak().vector.upsert({ "index": "docs", "vectors": [{"id": d["id"], "text": d["text"], "metadata": {"text": d["text"]}} for d in documents], })examples/17_vector_rag.py and examples/19_document_ingestion_recipe.py have complete routes.
Bring your own ingestion
Section titled “Bring your own ingestion”Kirak’s job starts at text: upsert embeds and stores it. Getting text out of PDFs, Word files or
web pages, and splitting it into chunks, is left to the application, because the right choices depend
on the documents and the libraries are heavy. Popular options: Unstructured,
Docling, LlamaIndex readers and LangChain document loaders.
Extract and chunk with one of them, then call upsert with text items; examples/19_document_ingestion_recipe.py
does this with Unstructured.
Adding a Custom Provider
Section titled “Adding a Custom Provider”The shared contract, the credential check (
check()), entry points and testing are covered in Adding a Provider.
Subclass VectorStoreProvider or EmbeddingProvider (kirak/vector/providers/base.py) and register
it for its category, "store" or "embedding":
from kirak.vector.providers.base import EmbeddingProvider
class CohereEmbedding(EmbeddingProvider): TYPE_NAME = "cohere" SECRET_FIELDS = ("api_key",) # read from KIRAK_VECTOR_EMBEDDING_<INSTANCE>_API_KEY
async def embed_query(self, text): ... async def embed_documents(self, texts): ...
async def on_kirak_ready(kirak): kirak.vector.register_provider("embedding", "cohere", CohereEmbedding)A store provider implements create_index, delete_index, list_indexes, describe_index,
upsert, delete, fetch and query. Providers can also be published as pip packages with an entry
point in the kirak.vector_store_providers or kirak.vector_embedding_providers group.
examples/18_custom_vector_provider.py is a complete in-memory store.