Design the Control Plane Before Agents Call Tools
October 9, 2026 · 7 min read

Multi-agent systems rarely fail because one agent cannot produce a plausible plan. They fail at the boundaries: two agents issue the same refund, a stale task overwrites newer work, or an approval is treated as permission to call every tool.
The central engineering decision is therefore not which orchestration pattern sounds most intelligent. It is where probabilistic reasoning ends and deterministic control begins.
For production systems, agents should propose work. A control plane should decide whether that work may run, with which credentials, against which version of state, and under what operational limits.
Treat orchestration as a distributed system
Once agents delegate tasks and call external services, you have a distributed workflow engine with unusual workers. Those workers are nondeterministic, may misunderstand instructions, and can emit malformed or duplicated requests.
The same concerns that apply to payment processors and job queues apply here:
- Explicit ownership of each task
- Durable state transitions
- Idempotent operations
- Bounded retries and timeouts
- Concurrency control
- Least-privilege credentials
- Traceable inputs, decisions, and side effects
A conversational transcript does not provide these guarantees. It is useful context for a model, but it is a poor system of record. Keep authoritative workflow state in normal infrastructure: a database, queue, workflow engine, or event log.
Represent each unit of work with a task ID, status, owner, input references, dependency list, deadline, attempt count, and output location. Agents may interpret that state, but they should not invent an alternative version inside their context windows.
Choose topology based on coordination cost
Three topologies cover most practical implementations.
A supervisor-and-workers design routes all delegation through one coordinator. It is easy to understand and gives you a clear place to enforce policy. Its weakness is concentration: the supervisor accumulates context, latency, and responsibility. Use it when workflows are bounded and centralized decisions genuinely matter.
A pipeline assigns a stable stage to each agent, such as intake, extraction, review, and publication. Pipelines are simpler to operate because dependencies are explicit. They work well when the business process is already sequential. Do not add a general-purpose supervisor when a state machine can advance the work.
An event-driven mesh lets specialized agents consume events and publish results. It supports parallelism and independent scaling, but ownership becomes harder to see. Use it only when domains are genuinely decoupled and your team already operates event-driven systems competently.
The right question is not “Which topology is most autonomous?” Ask: “Where can two components disagree about ownership, order, or completion?” Choose the topology that makes those disagreements easiest to prevent.
In many systems, the strongest design is hybrid: a deterministic workflow controls stage transitions, while an agent chooses among approved actions inside a stage.
Make every tool call a typed command
A model-generated tool call is untrusted input, even when the model runs inside your network. Treat it like an API request from an external client.
Each tool should expose a narrow, versioned contract. Avoid generic tools such as run_sql, call_api, or execute_shell. They move too much authority into model-generated arguments. Prefer business-level commands such as get_customer_summary, draft_refund_request, or schedule_approved_maintenance.
A robust command envelope should include:
- Tool name and schema version
- Workflow, task, and invocation IDs
- Actor and delegated identity
- Structured arguments
- Idempotency key
- Expected resource version
- Timeout and retry classification
- Policy decision reference
- Reason for the call
Validate syntax and semantics before execution. A date can match the schema and still violate a booking window. An account ID can be valid and still belong to another tenant.
Separate read tools from write tools. Reads can often tolerate bounded automatic retries. Writes need idempotency, conflict detection, and sometimes approval. Never let the model decide whether an operation is safe to retry; encode that property in the tool registry.
Separate planning, authorization, and execution
A common design mistake is allowing the same agent to decide an action, authorize it, and execute it. Splitting those responsibilities creates inspectable checkpoints.
A practical call path looks like this:
- An agent proposes a typed command.
- A validator checks the schema and business invariants.
- A policy service evaluates identity, scope, data sensitivity, and environment.
- An approval service intervenes when thresholds require it.
- A tool gateway injects short-lived credentials and executes the command.
- The result and side-effect receipt are stored durably.
- The orchestrator advances or compensates the workflow.
Agents should never receive broad service credentials in prompts or runtime environments. The gateway should exchange the agent's delegated identity for credentials limited to one tool, tenant, operation, and short time window.
Approval is also not a blanket capability grant. If a person approves a refund of $200 for order 123, bind that approval to the exact command or a cryptographic digest of its material fields. Replanning must not silently turn it into a $500 refund for another order.
Control concurrency explicitly
Parallel agents can reduce latency, but they also create races. Two workers may update the same ticket, reserve the same capacity, or derive actions from different snapshots.
Use established controls instead of asking agents to coordinate through prose:
- Optimistic locking for versioned records
- Leases for exclusive task ownership
- Idempotency keys for duplicate delivery
- Compare-and-swap transitions for workflow state
- Transactional outboxes for reliable event publication
- Sagas or compensating actions for multi-system writes
Do not assume exactly-once execution. Design for at-least-once delivery and make effects idempotent where possible. Where they cannot be idempotent, require a unique business key and record the external system's receipt before retrying.
Conflicts should return structured errors such as VERSION_MISMATCH or LEASE_EXPIRED, not free-form text. The orchestrator can then reload state, abandon the action, or request a new plan.
Bound failure instead of debating with it
Agent loops become expensive when every error is fed back to the model for another attempt. Many failures do not require reasoning.
Classify failures before deciding what happens next:
- Transient infrastructure errors: retry with backoff within a fixed budget.
- Invalid arguments: return field-level errors for one constrained repair attempt.
- Policy denials: stop; do not let the agent rephrase its way around policy.
- State conflicts: reload authoritative state and replan.
- Ambiguous business decisions: escalate to a person or designated decision service.
- Partial side effects: run a deterministic compensation path.
Set budgets per workflow, not merely per agent. A chain of five agents with three retries each can create fifteen expensive calls and repeated side effects. Cap elapsed time, model calls, tool calls, and monetary cost at the workflow level.
Circuit breakers should disable a tool or route when error rates spike. The fallback should be operationally explicit: queue for review, switch to read-only mode, or pause the workflow. “Ask another agent” is not a resilience strategy.
Observe commands, not just conversations
Prompt logs explain what agents saw, but production diagnosis depends on reconstructing causality.
Create one trace across delegation, model inference, policy checks, approvals, tool execution, state transitions, and compensation. Record task IDs and command IDs consistently so an operator can answer:
- Which agent proposed the action?
- Which policy allowed it?
- What resource version did it read?
- Was the call retried or duplicated?
- What external side effect occurred?
- Which later decision consumed the result?
Track operational measures such as conflict rate, duplicate suppression, policy denials, approval latency, compensation frequency, abandoned tasks, and tool-call fan-out. These expose orchestration defects that model-quality metrics will miss.
Redact sensitive arguments, but do not discard structural evidence. Store hashes, references, schema versions, and decision metadata when full payload retention is inappropriate.
Takeaway
Multi-agent orchestration is safest when agents remain replaceable reasoning components, not the owners of execution semantics. Put durable state, authorization, concurrency, retries, and side effects behind a deterministic control plane. Then choose the simplest topology that preserves clear task ownership and makes every tool call accountable.