Skip to content

Monitoring

Reference for the built-in monitoring module’s endpoints and configuration. For how the system works conceptually, see Observability.

Install:

Terminal window
pip install "kirak[monitoring]"

Enable:

app = create_kirak_app(models_path="./models/", modules=["monitoring"])
KIRAK_MONITORING_INGESTION_KEY=a-long-random-secret

Every endpoint except /monitoring/ping requires:

Authorization: Bearer <KIRAK_MONITORING_INGESTION_KEY>

Compared with constant-time comparison. A request without a bearer token gets 401. If KIRAK_MONITORING_INGESTION_KEY is not set, a request with a token gets 503 (fail closed) rather than unauthenticated access. Endpoints that read monitoring data also return 503 until the module has started.


Method Path Auth Description
GET /monitoring/ping none Liveness check only.
GET /monitoring/health bearer SQLite reachability, queue depths, dropped-record counts, DB size, max_db_size_mb, Python version.
GET /monitoring/metrics bearer OpenMetrics-format text export.
GET /monitoring/summary bearer RED aggregate plus trend deltas vs. the prior window. Query param: window_seconds.
GET /monitoring/trace/{request_id} bearer Full timeline for one request: the request, its steps, errors, and logs.
POST /monitoring/alerts/evaluate bearer Evaluate a batch of threshold checks. Body: {"checks": [...]}.
GET /monitoring/requests bearer Per-route request breakdown.
GET /monitoring/requests/timeseries bearer Bucketed time series. Query param: bucket_seconds.
GET /monitoring/requests/latency-histogram bearer Fixed-edge latency buckets.
GET /monitoring/requests/recent bearer Recent individual requests. Query params: status (2xx-5xx filter), limit (max 200).
GET /monitoring/ai/summary bearer AI runs in the window: run, completed and failed counts, error rate, total and average cost, input/output/total tokens, average and p95 duration. Query param: window_seconds.
GET /monitoring/ai/cost-timeseries bearer AI cost over time, in buckets. Query params: window_seconds, bucket_seconds.
GET /monitoring/ai/agents bearer AI runs broken down by agent. Query param: window_seconds.
GET /monitoring/ai/models bearer AI runs broken down by model. Query param: window_seconds.
GET /monitoring/ai/tools bearer Agent tool calls broken down by tool. Query param: window_seconds.
GET /monitoring/ai/runs bearer Recent agent runs, newest first: agent, model, provider, tokens, cost, duration, turns, tool calls, status, error code, conversation and user id. No prompt or output text is recorded. Query params: window_seconds, limit (max 200), status (completed, failed, pending_approval).
GET /monitoring/ai/runs/{run_id} bearer One run (run) and its tool calls (tool_calls: tool, duration, status, and for a failed call error_message – the tool’s error with emails and secrets masked, at most 300 characters).
GET /monitoring/mcp/summary bearer Per MCP server and tool: calls, errors, confirmation forms sent, average and p95 latency. Query param: window_seconds.
GET /monitoring/mcp/calls bearer Recent MCP tool calls, newest first: server, tool, operation ref, caller, status (ok, error, asked), error code, confirmation outcome, duration. Query params: window_seconds, limit (max 200), server, status.

POST /monitoring/alerts/evaluate evaluates a batch of threshold checks in one call:

{
"checks": [
{"id": "api_errors", "metric": "5xx_rate", "op": "gt", "value": 0.05, "window_seconds": 300, "routes": ["posts", "/webhook/stripe"]},
{"id": "ai_spend", "metric": "ai_cost_rate", "op": "gt", "value": 0.01, "window_seconds": 3600},
{"id": "alive", "metric": "health"}
]
}
Field Required Description
id yes Your identifier, echoed back in the result.
metric yes One of the metrics below.
op yes, except health gt or lt.
value yes, except health Threshold to compare against.
window_seconds yes, except health How far back to look.
routes no Limit the check to these routes. Each entry is a model name ("posts", its whole CRUD group) or a literal route ("/webhook/stripe"). HTTP metrics only.
Metric Value
5xx_rate, 4xx_rate Share of requests with that status class, 0 to 1.
5xx_count Number of 5xx responses.
request_rate Requests per second.
p95_latency, p99_latency Latency in milliseconds; null (never breaching) when there were no requests.
ai_cost_rate AI spend in USD per second, from recorded agent runs.
ai_error_rate Share of agent runs that failed, 0 to 1.
ai_token_rate AI tokens per second.
health Breaching when the monitoring SQLite store is unreachable.

A window with no traffic reads 0 for the rate and count metrics, so an lt check on request_rate fires on silence. The response is {"results": [{id, breaching, value, sample_count}, ...]}, one entry per check. A malformed check (unknown metric or op, missing field) rejects the whole batch with 422. Delivery (Slack, email, etc.) is up to you.


Environment variables (secrets, KIRAK_MONITORING_ prefix)

Section titled “Environment variables (secrets, KIRAK_MONITORING_ prefix)”
Variable Description
KIRAK_MONITORING_INGESTION_KEY Bearer token required on every endpoint except /ping.
KIRAK_MONITORING_LOG_SHIP_SECRET Bearer token used when shipping logs to an external collector (see below).

kirak.json settings ("monitoring" section)

Section titled “kirak.json settings ("monitoring" section)”
{
"monitoring": {
"db_path": "var/kirak/metrics.db",
"batch_interval_seconds": 5,
"retention_days": 7,
"sample_rate": 1.0,
"slow_request_threshold_ms": 500,
"max_db_size_mb": 200,
"queue_max_size": 10000,
"log_capture_level": "WARNING",
"log_queue_max_size": 5000,
"log_ship_url": null,
"log_ship_level": "WARNING",
"log_ship_batch_size": 100
}
}
monitoring key Type Description
enabled boolean Not read: has no effect. Monitoring runs when “monitoring” is in modules.
db_path string Path of the SQLite metrics file. Default “var/kirak/metrics.db”.
batch_interval_seconds integer Seconds between writes of collected data. Default 5.
retention_days integer Days metrics are kept. Default 7.
sample_rate number Share of requests whose per-operation timings (request steps) are recorded, from 0 to 1; every request itself is recorded. Default 1.0.
slow_request_threshold_ms integer Operations taking at least this many milliseconds have their timing recorded whatever sample_rate. Default 500.
max_db_size_mb integer Size cap of the metrics file in megabytes. Default 200.
queue_max_size integer Maximum records (requests, audit, AI) waiting in memory to be written. Default 10000.
log_capture_enabled boolean Not read: has no effect. Log records at log_capture_level and above are always captured.
log_capture_level string Lowest level of captured log records. Default WARNING.
log_queue_max_size integer Maximum log records waiting to be written. Default 5000.
log_user_id boolean Not read: has no effect.
log_ship_url string or null URL captured log records are also sent to. Default none.
log_ship_level string Lowest level of shipped log records. Default WARNING.
log_ship_batch_size integer Log records sent per request to log_ship_url. Default 100.

Log shipping (log_ship_url) requires httpx, already a base Kirak dependency – it degrades silently (no shipping, no error) if the collector is unreachable.


Monitoring data lives in tables inside the SQLite file at db_path: http_requests, request_steps, request_errors, logs, ai_runs, ai_tool_calls and mcp_calls (one row per MCP tool call, when the mcp module is on too). These are internal to the monitoring module – query them through the HTTP endpoints above rather than directly, since the schema is not a public contract.