Skip to content

Observability

Kirak includes a built-in monitoring subsystem: request-level RED metrics (rate, errors, duration), structured request tracing, log capture, and threshold-based alert checks – all served from the runtime’s own endpoints, with no third-party APM required.

This page covers how it works. For the exact endpoints and configuration keys, see Monitoring Reference.


Monitoring is a module like payments or storage – it does nothing until you enable it:

app = create_kirak_app(models_path="./models/", modules=["monitoring"])

MonitoringMiddleware is actually added to every Kirak app unconditionally, but it’s a true no-op until the module is enabled – so there’s zero overhead if you never opt in.


Monitoring data (requests, timing, errors, logs) is written to its own SQLite database (var/kirak/metrics.db by default), completely isolated from your application’s own database pool. This means monitoring can’t contend with your app for connections, and it works the same way regardless of whether your app runs on MySQL or PostgreSQL.


  • Every HTTP request – via MonitoringMiddleware: method, path, status code, duration.
  • CRUD step timing – via hooks, sampled (see below), broken down by pipeline stage (access check, validation, SQL execution, serialization).
  • Unhandled exceptions – captured with a full stack trace, linked to the request that triggered them. Exceptions your code raises deliberately as KirakException/HTTPException are not treated as monitoring errors – they’re expected control flow, not incidents.
  • Application logs – captured up to log_capture_level (default WARNING) and linked to the request that produced them.

To keep overhead low on high-traffic apps, step-level timing is sampled rather than captured on every single request – controlled by sample_rate (default 1.0, i.e. no sampling). slow_request_threshold_ms (default 500 ms) flags requests as slow in the aggregated summary view without needing a separate alert rule.


Every request gets a request ID (the same X-Request-ID used for log correlation elsewhere in Kirak – see Middleware). Pull the full timeline for one request – the request itself, its CRUD steps, any errors, and any logs it produced – from GET /monitoring/trace/{request_id}. This is the primary way to debug a single problematic request: read the trace, don’t add print statements.


POST /monitoring/alerts/evaluate evaluates a batch of threshold checks (5xx rate, p95/p99 latency, request rate, and more) against the stored metrics and returns which ones are breaching. Kirak does not ship Slack/email/PagerDuty delivery, firing-state tracking, or rule storage – it’s pure computation. Wire the response into whatever notification path you already have (a scheduled job that calls this endpoint and posts to Slack on a breach, for example).


Two-tier cleanup, run daily (via the scheduler module if enabled, else a bare asyncio loop):

  1. Time-based – rows older than retention_days (default 7) are deleted across every monitoring table.
  2. Size-cap safety valve – if the database is still over max_db_size_mb (default 200) after step 1, the oldest 15% of the timeline is dropped regardless of age.