Observability¶
AiFlow exposes a small set of read-only endpoints for monitoring a deployment: a liveness check, a deeper credential check, and operational counters in both Prometheus and plain JSON form.
The liveness check is open, because a load balancer has to reach it before
anything else works. The metrics endpoints are not: queue depth, in-flight
calls, and failure counts describe how the business is doing, and a scraper
is the only thing that needs them. They are also licensed, so a deployment
without the metrics right does not mount them at all and they return
404.
Liveness: GET /health¶
The cheapest possible check: no credentials verified, no database query. Point an uptime monitor or load balancer health check at this.
Deep health: GET /health?deep=true¶
Checks the deployment's Gemini, Twilio, and Resend credentials against the live services. Use it to tell "the API is running" apart from "the API is running but misconfigured".
Each check reports one of three states:
ok, the credential works.error, the credential is set but failed a live check.not_configured, this deployment has not set that integration up. A normal state, not a failure.
{
"status": "ok",
"checks": {
"gemini": { "status": "ok", "detail": null },
"twilio": { "status": "ok", "detail": null },
"resend": { "status": "not_configured", "detail": null }
}
}
Overall status is "degraded" if any individual check came back
"error". A "not_configured" integration on its own doesn't degrade
the overall status.
Licence mode: the X-AiFlow-Mode header¶
Every response carries X-AiFlow-Mode whenever this deployment is running on
anything other than a verified, bought licence. A deployment in normal
operation sends no such header, so its presence is the signal.
| Value | Meaning |
|---|---|
| absent | A verified licence. Normal operation. |
demo |
No key configured, or a key that has never verified. The free evaluation tier. |
grace |
A key stopped verifying recently. Running on its last known-good entitlements, for 14 days from the first failure. |
dev |
Licensing is bypassed (AIFLOW_ENVIRONMENT=development), or the licence is signed by a test trust key rather than by MeridFlow. |
Worth alerting on. grace means something is wrong with the key, the clock,
or the environment, and that there is a deadline attached: after 14 days the
deployment drops to demo and licensed surfaces stop being mounted. dev
appearing on something you believe is a production deployment means it is not
running on a real licence at all.
GET /api/v1/features reports the same thing as a mode field, alongside the
resolved plan, entitlements, and usage.
Instances running on this licence¶
The same response reports how many instances of this deployment are running and whether the one answering may start new sessions:
| Field | Type | Meaning |
|---|---|---|
deployments_permitted |
integer | Production instances the licence allows. 0 means unmetered. |
instances_live |
integer | Instances that have sent a heartbeat in the last 90 seconds. |
instance_holds_licensed_slot |
boolean | Whether the instance answering this request may start new sessions. |
deployment_id |
string or null | A stable id for this deployment, so two of them can be told apart. |
Instances sharing a database count each other, which is what running several
behind a load balancer means. One beyond the limit keeps serving this endpoint
and its dashboard, and finishes any session it already has, but starts no new
ones. Watch instance_holds_licensed_slot if you scale automatically: a
false there is the signal that a replica is running unlicensed, and it is
also the reason a request may be turned away while usage still looks well
under the concurrency cap.
instances_live only counts instances sharing this deployment's database. A
separate deployment on the same licence is not visible here, by design: seeing
it would mean reporting somewhere, which this product does not do.
Metrics: GET /metrics¶
Prometheus text format. Point a Prometheus scraper at this endpoint. If you don't run Prometheus, it's simply unused: nothing else about the deployment depends on it being scraped.
Set METRICS_TOKEN in your environment and have the scraper send it as a
bearer token. Generate one with openssl rand -base64 32. With the variable
unset the endpoint returns 503 rather than serving openly, so scraping is
something you switch on deliberately.
In a Prometheus scrape config:
scrape_configs:
- job_name: aiflow
authorization:
type: Bearer
credentials_file: /etc/prometheus/aiflow-metrics-token
static_configs:
- targets: ["api.your-domain.com"]
It exposes four gauges:
| Metric | Meaning |
|---|---|
aiflow_outbound_queue_depth |
Outbound calls waiting to be dialed |
aiflow_outbound_calls_in_flight |
Outbound calls currently ringing or in progress |
aiflow_calls_failed_total |
Calls that ended in a failed state |
aiflow_context_documents_indexing_failed_total |
Context documents whose knowledge-base indexing failed |
Metrics summary: GET /api/v1/metrics/summary¶
The same four counts as plain JSON, for anyone who wants them without running a Prometheus server. This one takes an admin token rather than the metrics token: it backs the Metrics panel on the dashboard's Settings page, and the dashboard already has a session.