Skip to content
MeridFlow AiFlow v8.x • self-hosted

Errors & rate limits

The error envelope

Almost every non-2xx response on this API is a small JSON body with a single detail field:

{ "detail": "A human-readable message" }

Every endpoint on this site uses that shape for authentication, authorization, and not-found errors. Treat detail as something to log, not something to branch on. Its wording can change; the status code will not.

Two responses don't follow this shape, and both are worth knowing about specifically:

  • A request that fails schema validation (a missing required field, a field of the wrong type) gets a 422 whose detail is a list of validation errors rather than a single string. Each entry names the field that failed. See the interactive schema on API Reference for the exact shape an endpoint expects.
  • A 429 doesn't use detail at all. See Rate limits below.

Status codes you'll encounter

Status Meaning
401 Missing, invalid, or expired credential, an API key or an admin session token.
403 The credential is valid but doesn't grant this: an API key missing the required scope, or scoped to a different agent than the one in the URL.
404 The agent or resource in the URL doesn't exist, or this deployment's licence does not include the capability behind that path. See below.
402 A licence capacity limit is already reached. See Licence limits below.
422 The request body doesn't match the expected schema. See the note above, this one has a different detail shape.
429 Rate limited. See Rate limits below.
500 An unexpected server error, an issue on the deployment's side, not something your request caused.

Licence limits

A deployment's licence caps how much of each resource it can hold at once: agents, concurrent live sessions, admin seats, knowledge base documents and megabytes, skills, custom tools, MCP connections, Orchestrators, federation links, triggers, webhook subscriptions, and API keys.

Creating one more of something already at its cap returns 402, with a detail naming the limit:

{
  "detail": "This licence allows at most 3 agents. Upgrade at https://meridflow.com/pricing to add more."
}

402 rather than 403 on purpose: a licence ceiling is a billing matter, not a permissions one, and the two want different handling. Retrying will not help. Either delete something to free a slot, or move to a plan with a higher cap.

Three cases behave differently because they are not create calls:

  • Concurrent live sessions are capacity at an instant rather than a stored row. An outbound call waits in the queue for a slot instead of failing. An inbound caller gets the usual spoken hold. A widget visitor gets a 503 saying the assistants are busy, since a member of the public should not be shown a message about somebody else's licence.
  • An instance without a licensed slot turns new sessions away exactly the same three ways, for a different reason: more instances of this deployment are running than the licence permits. Worth knowing when debugging, because a 503 from this cause looks identical while usage still reads well under the concurrency cap. Check instance_holds_licensed_slot on GET /api/v1/features to tell them apart. The instance keeps serving its API and dashboard and finishes sessions already in flight.
  • A capability the licence does not include at all returns 404, not 402 or 403. Its routes are never mounted, so an unlicensed surface looks like one that was never built. Confirming that a feature exists but is withheld is itself information.

GET /api/v1/features reports every limit alongside current usage, which is the reliable way to see how close a deployment is to any cap.

Rate limits

Rate limiting is in-memory and per-client-IP, on by default. Two endpoints are rate-limited today:

Endpoint Default limit
Event ingestion (Sending events) 120 requests/minute
Widget session creation (Embedding the widget) 30 requests/minute

Both defaults are configurable by whoever operates your deployment, so the exact number can vary. Don't hardcode 120 or 30 into your integration: treat a 429 response itself as the signal to slow down.

A 429 doesn't use the {"detail": ...} envelope described above: it comes directly from the rate limiter:

{ "error": "Rate limit exceeded: 120 per 1 minute" }

There's no guaranteed Retry-After header on this response by default. Treat a 429 as "try again shortly" and back off with your own delay, rather than trying to parse one.

Backing off on a 429

status=$(curl -s -o /tmp/response.json -w "%{http_code}" \
  -X POST https://api.your-domain.com/api/v1/agents/1/events \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"event_type": "abandoned_cart", "payload": {}}')

if [ "$status" = "429" ]; then
  sleep 2
  curl -X POST https://api.your-domain.com/api/v1/agents/1/events \
    -H "Authorization: Bearer $API_KEY" \
    -H "Content-Type: application/json" \
    -d '{"event_type": "abandoned_cart", "payload": {}}'
fi
import time

import httpx


def send_event(payload: dict) -> httpx.Response:
    url = "https://api.your-domain.com/api/v1/agents/1/events"
    headers = {"Authorization": f"Bearer {api_key}"}

    response = httpx.post(url, headers=headers, json=payload)
    if response.status_code == 429:
        time.sleep(2)
        response = httpx.post(url, headers=headers, json=payload)

    response.raise_for_status()
    return response
async function sendEvent(payload: unknown): Promise<Response> {
  const url = "https://api.your-domain.com/api/v1/agents/1/events";
  const post = () =>
    fetch(url, {
      method: "POST",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
      },
      body: JSON.stringify(payload),
    });

  let response = await post();
  if (response.status === 429) {
    await new Promise((resolve) => setTimeout(resolve, 2000));
    response = await post();
  }
  return response;
}

A single retry after a short, fixed delay is enough for most integrations. If you're sending events in a tight loop (a bulk import, a batch job), prefer spacing requests out under the limit in the first place over relying on retries to absorb bursts.