Skip to content

AI Integration

The AI module provides text generation and autonomous agent runs. Embeddings are in the Vector module (kirak.vector.embed()), which also stores and searches them. It is built on pydantic-ai. Kirak ships first-class extras for OpenAI, Anthropic, and Google Gemini; any other provider pydantic-ai supports works by pairing kirak[ai] with that provider’s pydantic-ai-slim extra.

Install:

Terminal window
pip install "kirak[ai-openai]" # OpenAI GPT
pip install "kirak[ai-anthropic]" # Anthropic Claude
pip install "kirak[ai-google]" # Google Gemini
pip install "kirak[all-ai]" # all providers
pip install "kirak[ai]" # pydantic-ai only; install your provider SDK yourself

pydantic-ai is not part of the base kirak install; each extra above includes it.

Kirak does not ship its own LLM clients. It passes the "provider:model" string to pydantic-ai, and pydantic-ai calls the provider through that provider’s SDK. The ai-* extras install pydantic-ai-slim with the matching extra (ai-openai installs pydantic-ai-slim[openai], which brings in the openai package), so the SDK version is always one pydantic-ai supports.

Other providers. Any provider pydantic-ai supports works. Install kirak[ai], then add the provider through pydantic-ai-slim’s own extra rather than pinning its SDK yourself:

Terminal window
pip install "kirak[ai]" "pydantic-ai-slim[groq]" # groq:llama-3.3-70b-versatile, GROQ_API_KEY
pip install "kirak[ai]" "pydantic-ai-slim[mistral]" # mistral:mistral-large-latest, MISTRAL_API_KEY
pip install "kirak[ai]" "pydantic-ai-slim[openrouter]" # openrouter:anthropic/claude-sonnet-4-5, OPENROUTER_API_KEY

Some providers reuse another SDK (OpenRouter and Ollama use openai). See pydantic-ai’s installation page for the extra names and its models overview for each provider’s prefix, model names and API key variable. kirak validate warns about a prefix Kirak has no extra for; the agent still runs if pydantic-ai can load the provider.

Enable:

{
"modules": ["ai"],
"ai": {
"default_model": "openai:gpt-4o-mini"
}
}

Provider selection is a single "provider:model-name" string. Pass it as "model" per-call or set default_model in kirak.json.

Provider Format Example
OpenAI openai:<model> openai:gpt-4o
Anthropic anthropic:<model> anthropic:claude-haiku-4-5-20251001
Google Gemini google:<model> google:gemini-2.0-flash

API keys are read by the provider SDK from its own standard environment variable – Kirak does not read or rename them:

OPENAI_API_KEY=sk-...
ANTHROPIC_API_KEY=sk-ant-...
GOOGLE_API_KEY=AIza... # GEMINI_API_KEY is also accepted

Simple one-shot LLM call. No tools, no history, no agent definition required.

result = await kirak.ai.prompt({
"prompt": "Explain dependency injection in two sentences.",
"model": "openai:gpt-4o-mini", # optional -- overrides default_model
"max_tokens": 200,
"temperature": 0.7,
})
output = result["data"]["output"]

Response on success:

{
"statusCode": 200,
"status": "success",
"data": {
"output": "Dependency injection is a pattern where ...",
"model": "openai:gpt-4o-mini",
"usage": {
"input_tokens": 12,
"output_tokens": 48,
"total_tokens": 60,
"requests": 1,
"cost_usd": 0.000004
}
}
}
Parameter Default Description
prompt – Required. The text prompt.
model default_model pydantic-ai model string "provider:name".
instructions none System prompt prepended to the request.
max_tokens provider default Maximum output tokens.
temperature provider default Sampling temperature (0 = deterministic).
output_schema none Pydantic class or JSON Schema dict for structured output. When set, output is a validated dict instead of a plain string.

Errors: MISSING_MODEL (400), USER_ERROR (400), USAGE_LIMIT_EXCEEDED (429), UNEXPECTED_MODEL_BEHAVIOR (500), PROMPT_ERROR (500).


Run an agent with tools, optional structured output, and a full execution trace.

result = await kirak.ai.run_agent({
"agent": "customer-support",
"prompt": "List our most recent 5 customers.",
"max_budget_usd": 0.05,
})

Response on success:

{
"statusCode": 200,
"status": "success",
"data": {
"output": "Here are your 5 most recent customers: ...",
"model": "anthropic:claude-haiku-4-5-20251001",
"usage": {"input_tokens": 120, "output_tokens": 85, "total_tokens": 205, "requests": 2, "cost_usd": 0.0004},
"trace": [
{"type": "llm_call", "model": "anthropic:claude-haiku-4-5-20251001", "duration_ms": 312},
{"type": "tool_call", "name": "customer_fetch", "call_id": "c1", "args": {"limit": 5}},
{"type": "tool_result", "name": "customer_fetch", "call_id": "c1", "result": "[{\"id\": 1, ...}]", "duration_ms": 18, "status": "completed"},
{"type": "llm_call", "model": "anthropic:claude-haiku-4-5-20251001", "duration_ms": 280},
{"type": "text", "content": "Here are your 5 most recent customers: ..."}
],
"conversation_id": "uuid"
}
}

conversation_id is only present when a conversation is active (a conversation_id was passed in the request or new_conversation: true generated one).

Parameter Default Description
prompt – Required (or message alias).
agent – Named agent from agents/<name>.json. Overrides other params when set.
instructions none Additional context injected as a supplemental system prompt for this run; combined with, not replacing, the agent’s static instructions. extra_context is an alias.
context {} App-specific per-request state dict; accessible in tools via ctx.deps.context.
tools – Declared in the agent JSON file (agents/<name>.json), not accepted per-call.
output_schema none Pydantic class or JSON Schema for structured output.
max_turns 10 Maximum LLM turns before stopping.
max_budget_usd none Cost cap; raises AI_SERVICE_ERROR (500) if exceeded.
require_tool_use false Returns TOOL_NOT_USED (400) if the agent calls no tools.
goal_condition none Natural language condition evaluated by reviewer_model; the agent runs again until the reviewer judges the condition met. Has no effect without reviewer_model.
max_iterations 1 Maximum iterations when goal_condition is set.
reviewer_model none Model to score output against quality_rubric.
quality_rubric none Rubric text for the reviewer agent.
quality_threshold 2 Minimum score (0-3); lower returns QUALITY_FAILED (422). When review passes, data.review: {"score": N} is included in the response.
conversation_id none Persist and load message history across calls (requires Redis or a database).
new_conversation false Auto-generate a new conversation_id for this call and return it in the response. Has no effect when stream: true is set; pass an explicit conversation_id for streaming conversations.

The agent model is set in the agent JSON file ("model" field) and cannot be overridden per call. Use prompt for per-call model selection.

Errors: MISSING_AGENT (400), AGENT_NOT_FOUND (404), USER_ERROR (400), USAGE_LIMIT_EXCEEDED (429), TOOL_NOT_USED (400), QUALITY_FAILED (422), UNEXPECTED_MODEL_BEHAVIOR (500), AI_SERVICE_ERROR (500), AGENT_ERROR (500).

All four event types appear in the trace array:

Type Fields
llm_call model (string), duration_ms (int)
tool_call name (string), call_id (string), args (object)
tool_result name, call_id, result (string or null), duration_ms (int), status ("completed" or "failed"), error (string, only on failure)
text content (string)

llm_call appears once per LLM API request; a multi-turn trace interleaves llm_call, tool_call, tool_result, and text entries. tool_result.result is always a string (str(return_value)[:500]), not the original object. requests in the usage dict counts LLM API calls; it may be null if the provider does not report it.

Every agent run is kept inside the model’s context window automatically. The token threshold at which compaction triggers is set by compaction_token_threshold in kirak.json (default: 150,000).

  • Provider-native compaction. For anthropic:... models, pydantic-ai’s AnthropicCompaction is used: when input tokens exceed the threshold, Anthropic summarises the earlier history server-side and the run continues from that summary. For openai:... and openai-responses:... models, OpenAICompaction applies the same threshold server-side.
  • Harness compaction (all other providers). For Gemini and any other provider, Kirak uses TieredCompaction from pydantic-ai-harness: when the conversation exceeds the threshold, it first clears old tool results (ClearToolResults), then slides the window to keep only the most recent messages (SlidingWindowCompaction).
  • Disabling compaction. Set "compaction_token_threshold": null in kirak.json to disable all compaction.

A native tool names one Kirak operation as "owner.operation" (lowercase), and runs it as the caller, so the model’s access rules and row-level security apply as for REST:

  • A model’s CRUD operation: fetch, search, count, exists, create, update, upsert, delete, destroy or restore. "customer.fetch" runs kirak.fetch("customer", params).
  • A module operation, as kirak modules NAME --json lists them: "vector.search" runs kirak.vector.search(params). Only declared operations can be named, never another method of the module.
"customer.fetch" # kirak.fetch("customer", params)
"vector.search" # kirak.vector.search(params)

A ref that names no model or module operation fails schema validation at startup (kirak validate agents). A ref that passes validation but the operation is unavailable fails when the agent calls the tool.

from pydantic_ai import RunContext
from kirak.ai import KirakDeps
@kirak.agent("billing-agent").tool
async def issue_refund(ctx: RunContext[KirakDeps], order_id: str, amount: float) -> dict:
"""Issue a refund for an order."""
...
result = await kirak.ai.run_agent({
"agent": "billing-agent",
"prompt": "Issue a refund of $49.99 for order ORD-123.",
})

Tools must accept ctx: RunContext[KirakDeps] as their first parameter to access per-request context (ctx.deps.current_user, ctx.deps.kirak, ctx.deps.context). A tool registered without RunContext is still accepted but generates a warning at startup and cannot access ctx.deps.

The decorator takes no options. To set requires_confirmation or timeout (seconds the tool may run) on a Python tool, register it with a call instead:

async def issue_refund(ctx: RunContext[KirakDeps], order_id: str, amount: float) -> dict:
"""Issue a refund for an order."""
...
kirak.agent("billing-agent").tool(issue_refund, requires_confirmation=True, timeout=30)

When requires_confirmation=True, the run pauses and returns:

{
"statusCode": 200,
"status": "success",
"data": {
"status": "approval_required",
"approval_id": "uuid"
}
}

Resume with POST /ai/agents/{name}/resume:

{"approval_id": "uuid", "approved": true}

Bind to a named agent for tool registration or approval resumption:

from pydantic_ai import RunContext
from kirak.ai import KirakDeps
# Register a tool scoped to this agent
@kirak.agent("billing-agent").tool
async def get_invoice(ctx: RunContext[KirakDeps], invoice_id: str) -> dict:
"""Look up an invoice by ID."""
...
# Run the agent
result = await kirak.ai.run_agent({"agent": "billing-agent", "prompt": "What is my balance?"})

agent(name).tool registers the function as a tool scoped to that agent name. It is only called for runs that reference agent_name; all other agents ignore it.

To run a named agent, use run_agent({"agent": name, "prompt": ...}) directly.


Place .json files in agents/ (the directory name is fixed). Kirak loads them at startup and registers per-agent HTTP endpoints automatically.

agents/customer-support.json:

{
"name": "customer-support",
"description": "Handles customer support requests",
"model": "anthropic:claude-haiku-4-5-20251001",
"instructions": "You are a helpful customer support agent...",
"tools": {
"customer.fetch": {},
"order.fetch": {}
},
"max_turns": 8,
"access": "User"
}

Each file is checked at startup against the agent schema (kirak schema --print agent): an unknown key, a wrongly typed value or a malformed tool entry stops startup with a message naming it. Add "$schema": "../.kirak/agent.schema.json" after running kirak schema to get completion and checks in your editor; Kirak ignores that key.

The tools map uses the ref as the key and an options object as the value. Use {} to take all defaults. Python function tools are registered via @kirak.agent("name").tool, not in the agent JSON file.

Option Default Description
name ref with _ for . Tool name the model sees (customer_fetch). The default replaces . because model provider APIs refuse dots in tool names.
description operation description What the tool does; written for the model.
requires_confirmation false Pause the run for the caller to confirm before the tool is called.
timeout none Seconds the tool may run.
parameters none JSON Schema of the arguments the model may pass (see below).

parameters declares the arguments the LLM may pass to the tool. When omitted, the LLM receives an open schema (additionalProperties: true) and relies entirely on the tool’s docstring for guidance. When set, pydantic-ai validates the LLM’s arguments against the schema before dispatch.

{
"vector.search": {
"parameters": {
"type": "object",
"properties": {
"text": {"type": "string", "description": "What to search for"},
"index": {"type": "string", "enum": ["kb"]},
"top_k": {"type": "integer", "default": 5}
},
"required": ["text", "index"]
}
}
}

A simpler example with a single required field:

{
"customer.fetch": {
"parameters": {
"type": "object",
"properties": {
"id": {"type": "string", "description": "Customer ID"}
},
"required": ["id"]
}
}
}

The LLM sees these as typed arguments; pydantic-ai validates the call before dispatch. Without parameters, the LLM receives additionalProperties: true and must infer arguments from the docstring.

Register a tool only for a specific agent using agent(name).tool in on_kirak_ready (or any code that has the kirak instance):

from pydantic_ai import RunContext
from kirak.ai import KirakDeps
async def on_kirak_ready(kirak):
@kirak.agent("customer-support").tool
async def get_order(ctx: RunContext[KirakDeps], order_id: str) -> dict:
"""Look up an order by ID."""
return {"id": order_id, "status": "shipped", "total": 49.99}

Tools registered this way are injected at run time for that agent name and are not visible to other agents.

An agent can use the tools of any MCP server by listing it in "toolsets". This requires kirak[ai-mcp] alongside the model extra:

pip install "kirak[ai-openai,ai-mcp]"

Kirak-own server (an mcp/<name>.json server in the same project):

{
"toolsets": {
"knowledge_base": {}
}
}

The URL is derived from base_url in kirak.json as {base_url}/mcp/knowledge_base/. The caller’s bearer token is forwarded automatically – no credentials to configure. base_url must be set in kirak.json; startup fails with a clear message if it is absent.

External MCP server (any other MCP endpoint):

{
"toolsets": {
"github": {
"url": "https://api.githubcopilot.com/mcp/"
}
}
}

The auth token is read from KIRAK_AI_TOOLSET_GITHUB_TOKEN in the environment at startup. Leave the env var unset for public (unauthenticated) servers.

Both types can appear in the same agent file alongside "tools".

kirak validate agents checks that every Kirak-own toolset key has a matching mcp/{name}.json. kirak schema --print agent shows the full schema including the toolsets property.

Connection lifecycle: one connection is opened per agent run (not per LLM turn), so a 10-turn conversation makes one TCP connection to the MCP server.

Known limitation (v1): Kirak-own toolsets connect over HTTP to the same process ({base_url}/mcp/{name}/). On a single-worker deployment this is a local loopback; on multi-worker it routes through the load balancer.

Field Default Description
name – Required. Lowercase letters, digits, and hyphens; used in routes and as the agent() scope.
model – Required. "provider:name" string.
instructions – Required (or instructions_file). System prompt / agent instructions.
instructions_file – Path relative to the agent JSON file; YAML frontmatter (---...---) is stripped if present.
description "" Shown in GET /ai/agents listings.
tools {} Map of tool refs to options – see above.
toolsets {} Map of MCP server names to options – see above. Requires kirak[ai-mcp].
input_schema {} JSON Schema validated against the incoming request body.
output_schema none Dotted path to a Pydantic class for structured output.
conversation true Persist message history across calls (requires Redis or a database).
max_turns 10 Maximum LLM turns before stopping.
max_budget_usd none Cost cap; raises AI_SERVICE_ERROR (500) if exceeded.
require_tool_use false Returns TOOL_NOT_USED (400) if the agent calls no tools.
access "User" Who may call the agent’s endpoints: "Guest" (anyone, no sign-in), "User" (any signed-in caller), "Admin" or "System".
rate_limit_per_minute 60 Per-agent request rate limit.
goal_condition none Natural language condition evaluated by reviewer_model; the agent runs again until the reviewer judges the condition met. Has no effect without reviewer_model.
max_iterations 1 Maximum runs while goal_condition is not met.
reviewer_model none "provider:name" model that scores the output against quality_rubric.
quality_rubric none Rubric text for the reviewer model.
quality_threshold 2 Minimum reviewer score (0-3); lower fails with QUALITY_FAILED.
temperature provider default Sampling temperature (0 = most deterministic).

requires_auth is not an agent field: a file that has it fails to load. Use access ("Guest" for a public agent).

input_schema validates the body of POST /ai/agents/{name}/run before the agent runs. A body that breaks it (a missing required field, or an extra field when the schema sets additionalProperties: false) is a 400 VALIDATION_ERROR with the failures in details.errors, and the LLM is not called. The keys a run takes for itself (stream, conversation_id, new_conversation, instructions, context, max_budget_usd) are not checked against it.

{
"name": "order-lookup",
"model": "openai:gpt-4o-mini",
"instructions": "Look up and summarise the customer's recent orders.",
"input_schema": {
"type": "object",
"properties": {
"message": {"type": "string"},
"customer_id": {"type": "string"}
},
"required": ["message", "customer_id"],
"additionalProperties": false
},
"tools": {
"order.fetch": {
"parameters": {
"type": "object",
"properties": {
"customer_id": {"type": "string"},
"limit": {"type": "integer", "default": 10}
},
"required": ["customer_id"]
}
}
}
}

output_schema constrains the response to a Pydantic model. Specify the dotted Python import path to the class. data.output becomes a validated dict instead of a plain string.

myapp/schemas.py
from pydantic import BaseModel
class OrderSummary(BaseModel):
total_orders: int
recent_order_id: str
total_spent: float
{
"name": "order-lookup",
"model": "openai:gpt-4o-mini",
"instructions": "Return a structured order summary.",
"output_schema": "myapp.schemas.OrderSummary"
}

Response when output_schema is set:

{
"statusCode": 200,
"status": "success",
"data": {
"output": {
"total_orders": 12,
"recent_order_id": "ORD-789",
"total_spent": 549.99
},
"model": "openai:gpt-4o-mini",
"usage": {"input_tokens": 95, "output_tokens": 30, "total_tokens": 125, "requests": 1, "cost_usd": 0.000008},
"trace": [
{"type": "llm_call", "model": "openai:gpt-4o-mini", "duration_ms": 245},
{"type": "text", "content": "..."}
]
}
}

Auto-registered endpoints:

GET /ai/agents -- list all agents
GET /ai/agents/customer-support -- agent metadata
POST /ai/agents/customer-support/run -- start a run
POST /ai/agents/customer-support/resume -- resume after approval

Run via HTTP:

Terminal window
curl -X POST /ai/agents/customer-support/run \
-H "Authorization: Bearer <token>" \
-d '{"message": "Where is my order ORD-456?"}'

Pass "stream": true in the body of a named-agent run request. The response is an SSE stream:

Terminal window
curl -X POST /ai/agents/customer-support/run \
-H "Authorization: Bearer <token>" \
-d '{"message": "Where is my order?", "stream": true}'
data: {"type": "text_delta", "content": "Hello"}
data: {"type": "text_delta", "content": " world"}
data: {"type": "done", "usage": {"input_tokens": 5, "output_tokens": 3, "cost_usd": 0.000001}, "conversation_id": "uuid"}

conversation_id is only present in the done event when conversation_id was supplied in the request body.

Error event:

data: {"type": "error", "message": "..."}

Every run_agent call is recorded by the monitoring module ("monitoring" in kirak.json modules): agent, model, tokens, cost, duration, status and the tools it called (with a failed tool’s error, emails and secrets masked) – never the prompt or the output. Read it with the monitoring bearer token (KIRAK_MONITORING_INGESTION_KEY), not a user’s JWT:

GET /monitoring/ai/runs?limit=20
GET /monitoring/ai/runs/{run_id}

See Monitoring for the summary, cost, per-agent, per-model and per-tool endpoints.


{
"ai": {
"default_model": "openai:gpt-4o-mini",
"conversation_ttl_seconds": 3600,
"compaction_token_threshold": 150000
}
}
Key Default Description
default_model none pydantic-ai model string used when model is not passed per-call.
conversation_ttl_seconds 3600 Redis TTL for conversation session history.
compaction_token_threshold 150000 Token count at which compaction triggers. Set to null to disable.

Method Path Auth Description
POST /ai/prompt required One-shot text generation
GET /ai/agents public List registered agent definitions
GET /ai/agents/{name} public Agent metadata
POST /ai/agents/{name}/run per-agent Run a named agent (supports "stream": true)
POST /ai/agents/{name}/resume per-agent Resume a named agent run