When building autonomous AI systems, the control plane is your window into the machine. It's the console where you monitor workflow executions, track token costs, and occasionally intervene when an agent gets stuck. But what happens when the orchestrator's database pool is exhausted, or the execution engine crashes? If your control plane relies on the same infrastructure, you are flying blind right when you need visibility the most.
To solve this, I designed a strict "Two-Surface Architecture" for my production control plane, Bastion. The core principle is absolute decoupling: workflow observability and session control must operate on entirely independent tracks.
Here is how separating the read path from the control path creates a resilient control plane.
Surface One: Workflow Observability
The first surface handles everything you need to observe the system: monitoring the live directed acyclic graph (DAG) of the workflow, tracking token usage, and reviewing errors.
In Bastion, this track acts as a read-only observer of the Python orchestrator. It polls the orchestrator's PostgreSQL events table to reconstruct live state. A few critical rules govern this surface:
- Never Write: The control plane never writes to the database or triggers orchestrator-side state. It only consumes.
- Lazy Connections: The database connection pool is lazy and opened only on demand. It is never initialized at startup.
- Pinned Data Contracts: The orchestrator guarantees a stable JSON structure for the
eventstable, allowing the control plane to reliably parse the workflow graph without needing relational tables for every node state.
This surface is rich and detailed, but it is inherently fragile. If the database goes down, observability goes down with it. That is why the second surface exists.
Surface Two: Process and Session Control
The second surface handles action: attaching to an agent's terminal, sending commands, or killing a runaway process.
This surface must always work. To guarantee this, Bastion's session management track completely bypasses the database. Instead, it shells out directly to tmux.
When I run a command to attach to an agent, or send a prompt to an idle shell, the control plane communicates with tmux using standard shell invocations.
There is zero database infrastructure involved. The session commands run synchronously, and if the output contains a malformed line, it degrades gracefully with a warning rather than failing fatally.
By keeping session control infrastructure-free, I ensure that my ability to intervene is never blocked by a crashed application layer.
The Payoff: Graceful Degradation
The true value of the Two-Surface Architecture reveals itself during an outage. If the orchestrator's database becomes unreachable, Bastion's observability commands (like checking status or costs) will return a clean "service unreachable" diagnostic. But the session commands? They remain perfectly operational. I can still attach to the agent's tmux pane, see the raw standard output, and manually kill the process.
When building your own AI control planes, treat observability and control as two different domains. Let your dashboards rely on the database, but keep your kill switches and terminal attachments tied directly to the operating system.
When the database inevitably falls over, you'll be glad your control plane didn't go down with it.