AI Integration
The AI module provides text generation and autonomous agent runs. Embeddings are in the
Vector module (kirak.vector.embed()), which also stores and searches them.
It is built on pydantic-ai. Kirak ships first-class extras for OpenAI, Anthropic, and Google Gemini; any other provider pydantic-ai supports works by pairing kirak[ai] with that provider’s pydantic-ai-slim extra.
Install:
pip install "kirak[ai-openai]" # OpenAI GPTpip install "kirak[ai-anthropic]" # Anthropic Claudepip install "kirak[ai-google]" # Google Geminipip install "kirak[all-ai]" # all providerspip install "kirak[ai]" # pydantic-ai only; install your provider SDK yourselfpydantic-ai is not part of the base kirak install; each extra above includes it.
Kirak does not ship its own LLM clients. It passes the "provider:model" string to pydantic-ai,
and pydantic-ai calls the provider through that provider’s SDK. The ai-* extras install
pydantic-ai-slim with the matching extra
(ai-openai installs pydantic-ai-slim[openai], which brings in the openai package), so the
SDK version is always one pydantic-ai supports.
Other providers. Any provider pydantic-ai supports works. Install kirak[ai], then add the
provider through pydantic-ai-slim’s own extra rather than pinning its SDK yourself:
pip install "kirak[ai]" "pydantic-ai-slim[groq]" # groq:llama-3.3-70b-versatile, GROQ_API_KEYpip install "kirak[ai]" "pydantic-ai-slim[mistral]" # mistral:mistral-large-latest, MISTRAL_API_KEYpip install "kirak[ai]" "pydantic-ai-slim[openrouter]" # openrouter:anthropic/claude-sonnet-4-5, OPENROUTER_API_KEYSome providers reuse another SDK (OpenRouter and Ollama use openai). See pydantic-ai’s
installation page for the extra names and its
models overview for each provider’s prefix, model
names and API key variable. kirak validate warns about a prefix Kirak has no extra for; the
agent still runs if pydantic-ai can load the provider.
Enable:
{ "modules": ["ai"], "ai": { "default_model": "openai:gpt-4o-mini" }}Provider and model selection
Section titled “Provider and model selection”Provider selection is a single "provider:model-name" string. Pass it as
"model" per-call or set default_model in kirak.json.
| Provider | Format | Example |
|---|---|---|
| OpenAI | openai:<model> |
openai:gpt-4o |
| Anthropic | anthropic:<model> |
anthropic:claude-haiku-4-5-20251001 |
| Google Gemini | google:<model> |
google:gemini-2.0-flash |
API keys are read by the provider SDK from its own standard environment variable – Kirak does not read or rename them:
OPENAI_API_KEY=sk-...ANTHROPIC_API_KEY=sk-ant-...GOOGLE_API_KEY=AIza... # GEMINI_API_KEY is also acceptedprompt
Section titled “prompt”Simple one-shot LLM call. No tools, no history, no agent definition required.
result = await kirak.ai.prompt({ "prompt": "Explain dependency injection in two sentences.", "model": "openai:gpt-4o-mini", # optional -- overrides default_model "max_tokens": 200, "temperature": 0.7,})output = result["data"]["output"]Response on success:
{ "statusCode": 200, "status": "success", "data": { "output": "Dependency injection is a pattern where ...", "model": "openai:gpt-4o-mini", "usage": { "input_tokens": 12, "output_tokens": 48, "total_tokens": 60, "requests": 1, "cost_usd": 0.000004 } }}Parameters
Section titled “Parameters”| Parameter | Default | Description |
|---|---|---|
prompt |
– | Required. The text prompt. |
model |
default_model |
pydantic-ai model string "provider:name". |
instructions |
none | System prompt prepended to the request. |
max_tokens |
provider default | Maximum output tokens. |
temperature |
provider default | Sampling temperature (0 = deterministic). |
output_schema |
none | Pydantic class or JSON Schema dict for structured output. When set, output is a validated dict instead of a plain string. |
Errors: MISSING_MODEL (400), USER_ERROR (400), USAGE_LIMIT_EXCEEDED (429),
UNEXPECTED_MODEL_BEHAVIOR (500), PROMPT_ERROR (500).
run_agent
Section titled “run_agent”Run an agent with tools, optional structured output, and a full execution trace.
result = await kirak.ai.run_agent({ "agent": "customer-support", "prompt": "List our most recent 5 customers.", "max_budget_usd": 0.05,})Response on success:
{ "statusCode": 200, "status": "success", "data": { "output": "Here are your 5 most recent customers: ...", "model": "anthropic:claude-haiku-4-5-20251001", "usage": {"input_tokens": 120, "output_tokens": 85, "total_tokens": 205, "requests": 2, "cost_usd": 0.0004}, "trace": [ {"type": "llm_call", "model": "anthropic:claude-haiku-4-5-20251001", "duration_ms": 312}, {"type": "tool_call", "name": "customer_fetch", "call_id": "c1", "args": {"limit": 5}}, {"type": "tool_result", "name": "customer_fetch", "call_id": "c1", "result": "[{\"id\": 1, ...}]", "duration_ms": 18, "status": "completed"}, {"type": "llm_call", "model": "anthropic:claude-haiku-4-5-20251001", "duration_ms": 280}, {"type": "text", "content": "Here are your 5 most recent customers: ..."} ], "conversation_id": "uuid" }}conversation_id is only present when a conversation is active (a conversation_id was passed in the request or new_conversation: true generated one).
Parameters
Section titled “Parameters”| Parameter | Default | Description |
|---|---|---|
prompt |
– | Required (or message alias). |
agent |
– | Named agent from agents/<name>.json. Overrides other params when set. |
instructions |
none | Additional context injected as a supplemental system prompt for this run; combined with, not replacing, the agent’s static instructions. extra_context is an alias. |
context |
{} |
App-specific per-request state dict; accessible in tools via ctx.deps.context. |
tools |
– | Declared in the agent JSON file (agents/<name>.json), not accepted per-call. |
output_schema |
none | Pydantic class or JSON Schema for structured output. |
max_turns |
10 |
Maximum LLM turns before stopping. |
max_budget_usd |
none | Cost cap; raises AI_SERVICE_ERROR (500) if exceeded. |
require_tool_use |
false |
Returns TOOL_NOT_USED (400) if the agent calls no tools. |
goal_condition |
none | Natural language condition evaluated by reviewer_model; the agent runs again until the reviewer judges the condition met. Has no effect without reviewer_model. |
max_iterations |
1 |
Maximum iterations when goal_condition is set. |
reviewer_model |
none | Model to score output against quality_rubric. |
quality_rubric |
none | Rubric text for the reviewer agent. |
quality_threshold |
2 |
Minimum score (0-3); lower returns QUALITY_FAILED (422). When review passes, data.review: {"score": N} is included in the response. |
conversation_id |
none | Persist and load message history across calls (requires Redis or a database). |
new_conversation |
false |
Auto-generate a new conversation_id for this call and return it in the response. Has no effect when stream: true is set; pass an explicit conversation_id for streaming conversations. |
The agent model is set in the agent JSON file ("model" field) and cannot be overridden per call. Use prompt for per-call model selection.
Errors: MISSING_AGENT (400), AGENT_NOT_FOUND (404), USER_ERROR (400),
USAGE_LIMIT_EXCEEDED (429), TOOL_NOT_USED (400), QUALITY_FAILED (422),
UNEXPECTED_MODEL_BEHAVIOR (500), AI_SERVICE_ERROR (500), AGENT_ERROR (500).
Trace event types
Section titled “Trace event types”All four event types appear in the trace array:
| Type | Fields |
|---|---|
llm_call |
model (string), duration_ms (int) |
tool_call |
name (string), call_id (string), args (object) |
tool_result |
name, call_id, result (string or null), duration_ms (int), status ("completed" or "failed"), error (string, only on failure) |
text |
content (string) |
llm_call appears once per LLM API request; a multi-turn trace interleaves llm_call, tool_call, tool_result, and text entries. tool_result.result is always a string (str(return_value)[:500]), not the original object. requests in the usage dict counts LLM API calls; it may be null if the provider does not report it.
Long conversations: compaction
Section titled “Long conversations: compaction”Every agent run is kept inside the model’s context window automatically. The token threshold at which
compaction triggers is set by compaction_token_threshold in kirak.json (default: 150,000).
- Provider-native compaction. For
anthropic:...models, pydantic-ai’sAnthropicCompactionis used: when input tokens exceed the threshold, Anthropic summarises the earlier history server-side and the run continues from that summary. Foropenai:...andopenai-responses:...models,OpenAICompactionapplies the same threshold server-side. - Harness compaction (all other providers). For Gemini and any other provider, Kirak uses
TieredCompactionfrompydantic-ai-harness: when the conversation exceeds the threshold, it first clears old tool results (ClearToolResults), then slides the window to keep only the most recent messages (SlidingWindowCompaction). - Disabling compaction. Set
"compaction_token_threshold": nullinkirak.jsonto disable all compaction.
Kirak-native tools
Section titled “Kirak-native tools”A native tool names one Kirak operation as "owner.operation" (lowercase), and runs it as
the caller, so the model’s access rules and row-level security apply as for REST:
- A model’s CRUD operation:
fetch,search,count,exists,create,update,upsert,delete,destroyorrestore."customer.fetch"runskirak.fetch("customer", params). - A module operation, as
kirak modules NAME --jsonlists them:"vector.search"runskirak.vector.search(params). Only declared operations can be named, never another method of the module.
"customer.fetch" # kirak.fetch("customer", params)"vector.search" # kirak.vector.search(params)A ref that names no model or module operation fails schema validation at startup (kirak validate agents).
A ref that passes validation but the operation is unavailable fails when the agent calls the tool.
Custom tools with @kirak.agent().tool
Section titled “Custom tools with @kirak.agent().tool”from pydantic_ai import RunContextfrom kirak.ai import KirakDeps
@kirak.agent("billing-agent").toolasync def issue_refund(ctx: RunContext[KirakDeps], order_id: str, amount: float) -> dict: """Issue a refund for an order.""" ...
result = await kirak.ai.run_agent({ "agent": "billing-agent", "prompt": "Issue a refund of $49.99 for order ORD-123.",})Tools must accept ctx: RunContext[KirakDeps] as their first parameter to access per-request
context (ctx.deps.current_user, ctx.deps.kirak, ctx.deps.context). A tool registered
without RunContext is still accepted but generates a warning at startup and cannot access
ctx.deps.
The decorator takes no options. To set requires_confirmation or timeout (seconds the tool
may run) on a Python tool, register it with a call instead:
async def issue_refund(ctx: RunContext[KirakDeps], order_id: str, amount: float) -> dict: """Issue a refund for an order.""" ...
kirak.agent("billing-agent").tool(issue_refund, requires_confirmation=True, timeout=30)When requires_confirmation=True, the run pauses and returns:
{ "statusCode": 200, "status": "success", "data": { "status": "approval_required", "approval_id": "uuid" }}Resume with POST /ai/agents/{name}/resume:
{"approval_id": "uuid", "approved": true}agent()
Section titled “agent()”Bind to a named agent for tool registration or approval resumption:
from pydantic_ai import RunContextfrom kirak.ai import KirakDeps
# Register a tool scoped to this agent@kirak.agent("billing-agent").toolasync def get_invoice(ctx: RunContext[KirakDeps], invoice_id: str) -> dict: """Look up an invoice by ID.""" ...
# Run the agentresult = await kirak.ai.run_agent({"agent": "billing-agent", "prompt": "What is my balance?"})agent(name).tool registers the function as a tool scoped to that agent name. It is only
called for runs that reference agent_name; all other agents ignore it.
To run a named agent, use run_agent({"agent": name, "prompt": ...}) directly.
Declarative agents (agent JSON files)
Section titled “Declarative agents (agent JSON files)”Place .json files in agents/ (the directory name is fixed).
Kirak loads them at startup and registers per-agent HTTP endpoints automatically.
agents/customer-support.json:
{ "name": "customer-support", "description": "Handles customer support requests", "model": "anthropic:claude-haiku-4-5-20251001", "instructions": "You are a helpful customer support agent...", "tools": { "customer.fetch": {}, "order.fetch": {} }, "max_turns": 8, "access": "User"}Each file is checked at startup against the agent schema (kirak schema --print agent): an unknown key, a wrongly typed value or a malformed tool entry stops startup with a message naming it. Add "$schema": "../.kirak/agent.schema.json" after running kirak schema to get completion and checks in your editor; Kirak ignores that key.
Tool options
Section titled “Tool options”The tools map uses the ref as the key and an options object as the value. Use {} to take
all defaults. Python function tools are registered via @kirak.agent("name").tool, not in the
agent JSON file.
| Option | Default | Description |
|---|---|---|
name |
ref with _ for . |
Tool name the model sees (customer_fetch). The default replaces . because model provider APIs refuse dots in tool names. |
description |
operation description | What the tool does; written for the model. |
requires_confirmation |
false |
Pause the run for the caller to confirm before the tool is called. |
timeout |
none | Seconds the tool may run. |
parameters |
none | JSON Schema of the arguments the model may pass (see below). |
parameters declares the arguments the LLM may pass to the tool.
When omitted, the LLM receives an open schema (additionalProperties: true)
and relies entirely on the tool’s docstring for guidance. When set, pydantic-ai
validates the LLM’s arguments against the schema before dispatch.
{ "vector.search": { "parameters": { "type": "object", "properties": { "text": {"type": "string", "description": "What to search for"}, "index": {"type": "string", "enum": ["kb"]}, "top_k": {"type": "integer", "default": 5} }, "required": ["text", "index"] } }}A simpler example with a single required field:
{ "customer.fetch": { "parameters": { "type": "object", "properties": { "id": {"type": "string", "description": "Customer ID"} }, "required": ["id"] } }}The LLM sees these as typed arguments; pydantic-ai validates the call before
dispatch. Without parameters, the LLM receives additionalProperties: true
and must infer arguments from the docstring.
Agent-scoped tools
Section titled “Agent-scoped tools”Register a tool only for a specific agent using agent(name).tool in
on_kirak_ready (or any code that has the kirak instance):
from pydantic_ai import RunContextfrom kirak.ai import KirakDeps
async def on_kirak_ready(kirak): @kirak.agent("customer-support").tool async def get_order(ctx: RunContext[KirakDeps], order_id: str) -> dict: """Look up an order by ID.""" return {"id": order_id, "status": "shipped", "total": 49.99}Tools registered this way are injected at run time for that agent name and are not visible to other agents.
MCP server toolsets
Section titled “MCP server toolsets”An agent can use the tools of any MCP server by listing it in "toolsets".
This requires kirak[ai-mcp] alongside the model extra:
pip install "kirak[ai-openai,ai-mcp]"Kirak-own server (an mcp/<name>.json server in the same project):
{ "toolsets": { "knowledge_base": {} }}The URL is derived from base_url in kirak.json as {base_url}/mcp/knowledge_base/.
The caller’s bearer token is forwarded automatically – no credentials to configure.
base_url must be set in kirak.json; startup fails with a clear message if it is absent.
External MCP server (any other MCP endpoint):
{ "toolsets": { "github": { "url": "https://api.githubcopilot.com/mcp/" } }}The auth token is read from KIRAK_AI_TOOLSET_GITHUB_TOKEN in the environment at startup.
Leave the env var unset for public (unauthenticated) servers.
Both types can appear in the same agent file alongside "tools".
kirak validate agents checks that every Kirak-own toolset key has a matching
mcp/{name}.json. kirak schema --print agent shows the full schema including
the toolsets property.
Connection lifecycle: one connection is opened per agent run (not per LLM turn), so a 10-turn conversation makes one TCP connection to the MCP server.
Known limitation (v1): Kirak-own toolsets connect over HTTP to the same
process ({base_url}/mcp/{name}/). On a single-worker deployment this is a
local loopback; on multi-worker it routes through the load balancer.
Agent JSON fields
Section titled “Agent JSON fields”| Field | Default | Description |
|---|---|---|
name |
– | Required. Lowercase letters, digits, and hyphens; used in routes and as the agent() scope. |
model |
– | Required. "provider:name" string. |
instructions |
– | Required (or instructions_file). System prompt / agent instructions. |
instructions_file |
– | Path relative to the agent JSON file; YAML frontmatter (---...---) is stripped if present. |
description |
"" |
Shown in GET /ai/agents listings. |
tools |
{} |
Map of tool refs to options – see above. |
toolsets |
{} |
Map of MCP server names to options – see above. Requires kirak[ai-mcp]. |
input_schema |
{} |
JSON Schema validated against the incoming request body. |
output_schema |
none | Dotted path to a Pydantic class for structured output. |
conversation |
true |
Persist message history across calls (requires Redis or a database). |
max_turns |
10 |
Maximum LLM turns before stopping. |
max_budget_usd |
none | Cost cap; raises AI_SERVICE_ERROR (500) if exceeded. |
require_tool_use |
false |
Returns TOOL_NOT_USED (400) if the agent calls no tools. |
access |
"User" |
Who may call the agent’s endpoints: "Guest" (anyone, no sign-in), "User" (any signed-in caller), "Admin" or "System". |
rate_limit_per_minute |
60 |
Per-agent request rate limit. |
goal_condition |
none | Natural language condition evaluated by reviewer_model; the agent runs again until the reviewer judges the condition met. Has no effect without reviewer_model. |
max_iterations |
1 |
Maximum runs while goal_condition is not met. |
reviewer_model |
none | "provider:name" model that scores the output against quality_rubric. |
quality_rubric |
none | Rubric text for the reviewer model. |
quality_threshold |
2 |
Minimum reviewer score (0-3); lower fails with QUALITY_FAILED. |
temperature |
provider default | Sampling temperature (0 = most deterministic). |
requires_auth is not an agent field: a file that has it fails to load. Use access ("Guest" for a public agent).
Structured input and output
Section titled “Structured input and output”input_schema validates the body of POST /ai/agents/{name}/run before the agent runs.
A body that breaks it (a missing required field, or an extra field when the schema sets
additionalProperties: false) is a 400 VALIDATION_ERROR with the failures in
details.errors, and the LLM is not called. The keys a run takes for itself (stream,
conversation_id, new_conversation, instructions, context, max_budget_usd) are not
checked against it.
{ "name": "order-lookup", "model": "openai:gpt-4o-mini", "instructions": "Look up and summarise the customer's recent orders.", "input_schema": { "type": "object", "properties": { "message": {"type": "string"}, "customer_id": {"type": "string"} }, "required": ["message", "customer_id"], "additionalProperties": false }, "tools": { "order.fetch": { "parameters": { "type": "object", "properties": { "customer_id": {"type": "string"}, "limit": {"type": "integer", "default": 10} }, "required": ["customer_id"] } } }}output_schema constrains the response to a Pydantic model. Specify the
dotted Python import path to the class. data.output becomes a validated dict
instead of a plain string.
from pydantic import BaseModel
class OrderSummary(BaseModel): total_orders: int recent_order_id: str total_spent: float{ "name": "order-lookup", "model": "openai:gpt-4o-mini", "instructions": "Return a structured order summary.", "output_schema": "myapp.schemas.OrderSummary"}Response when output_schema is set:
{ "statusCode": 200, "status": "success", "data": { "output": { "total_orders": 12, "recent_order_id": "ORD-789", "total_spent": 549.99 }, "model": "openai:gpt-4o-mini", "usage": {"input_tokens": 95, "output_tokens": 30, "total_tokens": 125, "requests": 1, "cost_usd": 0.000008}, "trace": [ {"type": "llm_call", "model": "openai:gpt-4o-mini", "duration_ms": 245}, {"type": "text", "content": "..."} ] }}Auto-registered endpoints:
GET /ai/agents -- list all agentsGET /ai/agents/customer-support -- agent metadataPOST /ai/agents/customer-support/run -- start a runPOST /ai/agents/customer-support/resume -- resume after approvalRun via HTTP:
curl -X POST /ai/agents/customer-support/run \ -H "Authorization: Bearer <token>" \ -d '{"message": "Where is my order ORD-456?"}'Streaming
Section titled “Streaming”Pass "stream": true in the body of a named-agent run request. The response is
an SSE stream:
curl -X POST /ai/agents/customer-support/run \ -H "Authorization: Bearer <token>" \ -d '{"message": "Where is my order?", "stream": true}'data: {"type": "text_delta", "content": "Hello"}data: {"type": "text_delta", "content": " world"}data: {"type": "done", "usage": {"input_tokens": 5, "output_tokens": 3, "cost_usd": 0.000001}, "conversation_id": "uuid"}conversation_id is only present in the done event when conversation_id was supplied in the request body.
Error event:
data: {"type": "error", "message": "..."}Run history
Section titled “Run history”Every run_agent call is recorded by the monitoring module ("monitoring" in kirak.json modules): agent, model, tokens, cost, duration, status and the tools it called (with a failed tool’s error, emails and secrets masked) – never the prompt or the output. Read it with the monitoring bearer token (KIRAK_MONITORING_INGESTION_KEY), not a user’s JWT:
GET /monitoring/ai/runs?limit=20GET /monitoring/ai/runs/{run_id}See Monitoring for the summary, cost, per-agent, per-model and per-tool endpoints.
kirak.json reference
Section titled “kirak.json reference”{ "ai": { "default_model": "openai:gpt-4o-mini", "conversation_ttl_seconds": 3600, "compaction_token_threshold": 150000 }}| Key | Default | Description |
|---|---|---|
default_model |
none | pydantic-ai model string used when model is not passed per-call. |
conversation_ttl_seconds |
3600 |
Redis TTL for conversation session history. |
compaction_token_threshold |
150000 |
Token count at which compaction triggers. Set to null to disable. |
HTTP endpoints
Section titled “HTTP endpoints”| Method | Path | Auth | Description |
|---|---|---|---|
| POST | /ai/prompt |
required | One-shot text generation |
| GET | /ai/agents |
public | List registered agent definitions |
| GET | /ai/agents/{name} |
public | Agent metadata |
| POST | /ai/agents/{name}/run |
per-agent | Run a named agent (supports "stream": true) |
| POST | /ai/agents/{name}/resume |
per-agent | Resume a named agent run |