Skip to content
MeridFlow AiFlow v8.x • self-hosted

Orchestrators: multi-model, multi-agent delegation

Orchestrators are a core, first-class part of AiFlow, on equal footing with Agents and Triggers. Everything on this page is configurable both from the admin dashboard and directly through the API, using an admin session token (see Authentication). This page covers the concepts, the Guides section has the full endpoint-by-endpoint reference. If your deployment uses Orchestrators, it changes what your Agent integration can actually do, so it's worth knowing what's happening behind the scenes before jumping into the API reference.

What an Orchestrator is

An Agent (see Agents & channels) is a realtime, Gemini Live-powered voice or text persona. An Orchestrator is a different kind of entity entirely: a non-realtime, tool-enabled task-delegation layer. An Agent can hand a task to an Orchestrator mid-conversation and get back a real, model-generated answer, which the Agent then speaks or types back in its own voice, the same way it would relay an answer from another Agent it delegates to.

What sets an Orchestrator apart from an Agent-to-Agent delegation is capability: an Orchestrator runs a full, bounded, tool-using agentic loop of its own, with its own model choice, its own connected tools, and optionally a graph of its own sub-Orchestrators beneath it, rather than producing a single one-shot response.

Multi-model, not just Gemini

Every Agent's realtime voice and video engine is Gemini-only, permanently: that's a hard technical constraint, since no other provider currently offers an equivalent bidirectional realtime audio API. Orchestrators are where AiFlow breaks out of that constraint. An Orchestrator's model field accepts either a bare Gemini model name or a provider/model string (LiteLLM's own convention), which unlocks over 100 providers with no AiFlow-specific integration work:

Model string Provider
gemini-3.5-flash-lite Google (native)
anthropic/claude-sonnet-5 Anthropic
openai/gpt-5 OpenAI
ollama/llama3.1 A self-hosted Ollama instance
openrouter/<any model> Any model OpenRouter proxies

In practice, this means an Agent on a phone call, talking to a caller in real time over Gemini Live, can delegate a task mid-call to an Orchestrator running on Claude, GPT, or a locally hosted model, then speak that answer back, all inside one conversation.

Local or remote execution

Each Orchestrator independently runs one of two ways:

  • Local: the Orchestrator runs in-process, inside your AiFlow deployment, by default, with no separate infrastructure required.
  • Remote: proxies to an Orchestrator deployed externally (Vertex AI Agent Engine, Cloud Run, or anywhere else your team hosts it), over the A2A protocol, an open standard for one agent to call another. AiFlow never runs deployment tooling on your behalf here: your team deploys and hosts it themselves, then points the Orchestrator at the resulting endpoint.

Tools an Orchestrator can use

An Orchestrator reaches tools from three sources: other Orchestrators linked beneath it (as a callable sub-task, or a full hand-off of control), MCP servers it's connected to, and a curated set of built-in tools. Those built-ins mix AiFlow's own (send_email, send_whatsapp, send_telegram) with several general-purpose ones (web search, URL fetching, private-data search, code execution, and a few others), each scoped to Gemini models specifically. Any of AiFlow's own three can also be switched to require a human's approval before it actually runs, see Orchestrator tools. A queued request can announce itself by email, Telegram, WhatsApp, or a phone call, so it is not left waiting on somebody opening the dashboard.

Not supported: arbitrary custom Python function tools. Any business logic an Orchestrator needs is meant to live behind an MCP server instead, the same declared-integration boundary the rest of AiFlow holds to.

Built-in templates

Configuring an Orchestrator from scratch, its instruction, its tools, and any sub-Orchestrators beneath it, is real setup work. A template skips that: it instantiates a whole ready-made multi-Orchestrator setup, a top-level Orchestrator plus two sub-Orchestrators, already linked and tooled, given only a slug and a name. Each one is adapted from a real Google ADK sample agent, narrowed to what needs nothing beyond Gemini or Vertex AI and never pauses mid-run waiting on a human decision, the same two constraints every Orchestrator in AiFlow is built around. The result is an ordinary Orchestrator: edit it afterward exactly as you would one built by hand. See Built-in templates for the full catalog and the instantiate endpoint.

Why delegation can't form a loop

The call graph is one-directional by construction: a native Agent may delegate into an Orchestrator, and an Orchestrator may delegate into its own sub-Orchestrators, but an Orchestrator can never call back into the native-Agent layer. Combined with cycle detection on the sub-Orchestrator graph itself, an infinite delegation loop is structurally impossible, not merely something a runtime check happens to catch. Every Orchestrator run is also bounded by a step budget and a wall-clock timeout, independent of that structural guarantee.

Whether your deployment has this turned on

Orchestrators are enabled per deployment (an admin setting, off by default). If you're unsure whether the Agent you're integrating with can delegate to an Orchestrator, call GET /api/v1/features rather than assume either way. See Agent-to-Orchestrator delegation. When the setting is off, every Orchestrator endpoint (creating one, configuring its tools, linking sub-Orchestrators, connecting MCP servers, granting an Agent delegation permission) simply doesn't exist on your deployment: a request to any of them gets a plain 404.

Where to go next

This page is the concepts; the full API reference for everything described above lives under Guides: