Skip to content
MeridFlow AiFlow v8.x • self-hosted

Observability

AiFlow exposes a small set of read-only endpoints for monitoring a deployment: a liveness check, a deeper credential check, and operational counters in both Prometheus and plain JSON form.

The liveness check is open, because a load balancer has to reach it before anything else works. The metrics endpoints are not: queue depth, in-flight calls, and failure counts describe how the business is doing, and a scraper is the only thing that needs them. They are also licensed, so a deployment without the metrics right does not mount them at all and they return 404.

Liveness: GET /health

The cheapest possible check: no credentials verified, no database query. Point an uptime monitor or load balancer health check at this.

curl https://api.your-domain.com/health
{ "status": "ok" }

Deep health: GET /health?deep=true

Checks the deployment's Gemini, Twilio, and Resend credentials against the live services. Use it to tell "the API is running" apart from "the API is running but misconfigured".

Each check reports one of three states:

  • ok, the credential works.
  • error, the credential is set but failed a live check.
  • not_configured, this deployment has not set that integration up. A normal state, not a failure.
curl "https://api.your-domain.com/health?deep=true"
import httpx

response = httpx.get("https://api.your-domain.com/health", params={"deep": "true"})
response.raise_for_status()
health = response.json()
const response = await fetch(
  "https://api.your-domain.com/health?deep=true",
);
const health = await response.json();
{
  "status": "ok",
  "checks": {
    "gemini": { "status": "ok", "detail": null },
    "twilio": { "status": "ok", "detail": null },
    "resend": { "status": "not_configured", "detail": null }
  }
}

Overall status is "degraded" if any individual check came back "error". A "not_configured" integration on its own doesn't degrade the overall status.

Licence mode: the X-AiFlow-Mode header

Every response carries X-AiFlow-Mode whenever this deployment is running on anything other than a verified, bought licence. A deployment in normal operation sends no such header, so its presence is the signal.

Value Meaning
absent A verified licence. Normal operation.
demo No key configured, or a key that has never verified. The free evaluation tier.
grace A key stopped verifying recently. Running on its last known-good entitlements, for 14 days from the first failure.
dev Licensing is bypassed (AIFLOW_ENVIRONMENT=development), or the licence is signed by a test trust key rather than by MeridFlow.

Worth alerting on. grace means something is wrong with the key, the clock, or the environment, and that there is a deadline attached: after 14 days the deployment drops to demo and licensed surfaces stop being mounted. dev appearing on something you believe is a production deployment means it is not running on a real licence at all.

GET /api/v1/features reports the same thing as a mode field, alongside the resolved plan, entitlements, and usage.

Instances running on this licence

The same response reports how many instances of this deployment are running and whether the one answering may start new sessions:

Field Type Meaning
deployments_permitted integer Production instances the licence allows. 0 means unmetered.
instances_live integer Instances that have sent a heartbeat in the last 90 seconds.
instance_holds_licensed_slot boolean Whether the instance answering this request may start new sessions.
deployment_id string or null A stable id for this deployment, so two of them can be told apart.

Instances sharing a database count each other, which is what running several behind a load balancer means. One beyond the limit keeps serving this endpoint and its dashboard, and finishes any session it already has, but starts no new ones. Watch instance_holds_licensed_slot if you scale automatically: a false there is the signal that a replica is running unlicensed, and it is also the reason a request may be turned away while usage still looks well under the concurrency cap.

instances_live only counts instances sharing this deployment's database. A separate deployment on the same licence is not visible here, by design: seeing it would mean reporting somewhere, which this product does not do.

Metrics: GET /metrics

Prometheus text format. Point a Prometheus scraper at this endpoint. If you don't run Prometheus, it's simply unused: nothing else about the deployment depends on it being scraped.

Set METRICS_TOKEN in your environment and have the scraper send it as a bearer token. Generate one with openssl rand -base64 32. With the variable unset the endpoint returns 503 rather than serving openly, so scraping is something you switch on deliberately.

curl -H "Authorization: Bearer $METRICS_TOKEN"   https://api.your-domain.com/metrics

In a Prometheus scrape config:

scrape_configs:
  - job_name: aiflow
    authorization:
      type: Bearer
      credentials_file: /etc/prometheus/aiflow-metrics-token
    static_configs:
      - targets: ["api.your-domain.com"]

It exposes four gauges:

Metric Meaning
aiflow_outbound_queue_depth Outbound calls waiting to be dialed
aiflow_outbound_calls_in_flight Outbound calls currently ringing or in progress
aiflow_calls_failed_total Calls that ended in a failed state
aiflow_context_documents_indexing_failed_total Context documents whose knowledge-base indexing failed

Metrics summary: GET /api/v1/metrics/summary

The same four counts as plain JSON, for anyone who wants them without running a Prometheus server. This one takes an admin token rather than the metrics token: it backs the Metrics panel on the dashboard's Settings page, and the dashboard already has a session.

curl -H "Authorization: Bearer $ADMIN_TOKEN"   https://api.your-domain.com/api/v1/metrics/summary
{
  "outbound_queue_depth": 0,
  "outbound_calls_in_flight": 0,
  "calls_failed_total": 0,
  "context_documents_indexing_failed_total": 0
}