Orchestrators: multi-model, multi-agent delegation¶
Orchestrators are a core, first-class part of AiFlow, on equal footing with Agents and Triggers. Everything on this page is configurable both from the admin dashboard and directly through the API, using an admin session token (see Authentication). This page covers the concepts, the Guides section has the full endpoint-by-endpoint reference. If your deployment uses Orchestrators, it changes what your Agent integration can actually do, so it's worth knowing what's happening behind the scenes before jumping into the API reference.
What an Orchestrator is¶
An Agent (see Agents & channels) is a realtime, Gemini Live-powered voice or text persona. An Orchestrator is a different kind of entity entirely: a non-realtime, tool-enabled task-delegation layer. An Agent can hand a task to an Orchestrator mid-conversation and get back a real, model-generated answer, which the Agent then speaks or types back in its own voice, the same way it would relay an answer from another Agent it delegates to.
What sets an Orchestrator apart from an Agent-to-Agent delegation is capability: an Orchestrator runs a full, bounded, tool-using agentic loop of its own, with its own model choice, its own connected tools, and optionally a graph of its own sub-Orchestrators beneath it, rather than producing a single one-shot response.
Multi-model, not just Gemini¶
Every Agent's realtime voice and video engine is Gemini-only, permanently:
that's a hard technical constraint, since no other provider currently
offers an equivalent bidirectional realtime audio API. Orchestrators are
where AiFlow breaks out of that constraint. An Orchestrator's model field
accepts either a bare Gemini model name or a provider/model string
(LiteLLM's own convention), which unlocks over 100 providers with no
AiFlow-specific integration work:
| Model string | Provider |
|---|---|
gemini-3.5-flash-lite |
Google (native) |
anthropic/claude-sonnet-5 |
Anthropic |
openai/gpt-5 |
OpenAI |
ollama/llama3.1 |
A self-hosted Ollama instance |
openrouter/<any model> |
Any model OpenRouter proxies |
In practice, this means an Agent on a phone call, talking to a caller in real time over Gemini Live, can delegate a task mid-call to an Orchestrator running on Claude, GPT, or a locally hosted model, then speak that answer back, all inside one conversation.
Local or remote execution¶
Each Orchestrator independently runs one of two ways:
- Local: the Orchestrator runs in-process, inside your AiFlow deployment, by default, with no separate infrastructure required.
- Remote: proxies to an Orchestrator deployed externally (Vertex AI Agent Engine, Cloud Run, or anywhere else your team hosts it), over the A2A protocol, an open standard for one agent to call another. AiFlow never runs deployment tooling on your behalf here: your team deploys and hosts it themselves, then points the Orchestrator at the resulting endpoint.
Tools an Orchestrator can use¶
An Orchestrator reaches tools from three sources: other Orchestrators
linked beneath it (as a callable sub-task, or a full hand-off of control),
MCP servers it's connected to, and a curated set of built-in tools. Those
built-ins mix AiFlow's own (send_email, send_whatsapp,
send_telegram) with several
general-purpose ones (web search, URL fetching, private-data search, code
execution, and a few others), each scoped to Gemini models specifically.
Any of AiFlow's own three can also be switched to require a human's
approval before it actually runs, see
Orchestrator tools.
A queued request can announce itself by email, Telegram, WhatsApp, or a
phone call, so it is not left waiting on somebody opening the dashboard.
Not supported: arbitrary custom Python function tools. Any business logic an Orchestrator needs is meant to live behind an MCP server instead, the same declared-integration boundary the rest of AiFlow holds to.
Built-in templates¶
Configuring an Orchestrator from scratch, its instruction, its tools, and any sub-Orchestrators beneath it, is real setup work. A template skips that: it instantiates a whole ready-made multi-Orchestrator setup, a top-level Orchestrator plus two sub-Orchestrators, already linked and tooled, given only a slug and a name. Each one is adapted from a real Google ADK sample agent, narrowed to what needs nothing beyond Gemini or Vertex AI and never pauses mid-run waiting on a human decision, the same two constraints every Orchestrator in AiFlow is built around. The result is an ordinary Orchestrator: edit it afterward exactly as you would one built by hand. See Built-in templates for the full catalog and the instantiate endpoint.
Why delegation can't form a loop¶
The call graph is one-directional by construction: a native Agent may delegate into an Orchestrator, and an Orchestrator may delegate into its own sub-Orchestrators, but an Orchestrator can never call back into the native-Agent layer. Combined with cycle detection on the sub-Orchestrator graph itself, an infinite delegation loop is structurally impossible, not merely something a runtime check happens to catch. Every Orchestrator run is also bounded by a step budget and a wall-clock timeout, independent of that structural guarantee.
Whether your deployment has this turned on¶
Orchestrators are enabled per deployment (an admin setting, off by
default). If you're unsure whether the Agent you're integrating with can
delegate to an Orchestrator, call GET /api/v1/features rather than
assume either way. See
Agent-to-Orchestrator delegation.
When the setting is off, every Orchestrator endpoint (creating one,
configuring its tools, linking sub-Orchestrators, connecting MCP servers,
granting an Agent delegation permission) simply doesn't exist on your
deployment: a request to any of them gets a plain 404.
Where to go next¶
This page is the concepts; the full API reference for everything described above lives under Guides:
- Orchestrators: the full CRUD reference, every field, every endpoint.
- Built-in templates: the ready-made multi-Orchestrator setups and how to instantiate one.
- Orchestrator tools: the built-in tool catalog and how to enable one.
- Orchestrator links: building a graph of sub-Orchestrators.
- Orchestrator MCP servers: per-Orchestrator and shared MCP connections.
- Agent-to-Orchestrator delegation: granting a native Agent permission to delegate into an Orchestrator.
- Console: chatting with an Orchestrator directly, for debugging and exploration.