Skip to main content

Blog Post

The Two-Surface Architecture for AI Control Planes

•3 min read•By Brandon

ArchitectureControl PlanesAgentic SystemsRust

When building autonomous AI systems, the control plane is your window into the machine. It's the console where you monitor workflow executions, track token costs, and occasionally intervene when an agent gets stuck. But what happens when the orchestrator's database pool is exhausted, or the execution engine crashes? If your control plane relies on the same infrastructure, you are flying blind right when you need visibility the most.

To solve this, I designed a strict "Two-Surface Architecture" for my production control plane, Bastion. The core principle is absolute decoupling: workflow observability and session control must operate on entirely independent tracks.

Here is how separating the read path from the control path creates a resilient control plane.

Surface One: Workflow Observability

The first surface handles everything you need to observe the system: monitoring the live directed acyclic graph (DAG) of the workflow, tracking token usage, and reviewing errors.

In Bastion, this track acts as a read-only observer of the Python orchestrator. It polls the orchestrator's PostgreSQL events table to reconstruct live state. A few critical rules govern this surface:

  • Never Write: The control plane never writes to the database or triggers orchestrator-side state. It only consumes.
  • Lazy Connections: The database connection pool is lazy and opened only on demand. It is never initialized at startup.
  • Pinned Data Contracts: The orchestrator guarantees a stable JSON structure for the events table, allowing the control plane to reliably parse the workflow graph without needing relational tables for every node state.

This surface is rich and detailed, but it is inherently fragile. If the database goes down, observability goes down with it. That is why the second surface exists.

Surface Two: Process and Session Control

The second surface handles action: attaching to an agent's terminal, sending commands, or killing a runaway process.

This surface must always work. To guarantee this, Bastion's session management track completely bypasses the database. Instead, it shells out directly to tmux.

When I run a command to attach to an agent, or send a prompt to an idle shell, the control plane communicates with tmux using standard shell invocations.

There is zero database infrastructure involved. The session commands run synchronously, and if the output contains a malformed line, it degrades gracefully with a warning rather than failing fatally.

By keeping session control infrastructure-free, I ensure that my ability to intervene is never blocked by a crashed application layer.

The Payoff: Graceful Degradation

The true value of the Two-Surface Architecture reveals itself during an outage. If the orchestrator's database becomes unreachable, Bastion's observability commands (like checking status or costs) will return a clean "service unreachable" diagnostic. But the session commands? They remain perfectly operational. I can still attach to the agent's tmux pane, see the raw standard output, and manually kill the process.

When building your own AI control planes, treat observability and control as two different domains. Let your dashboards rely on the database, but keep your kill switches and terminal attachments tied directly to the operating system.

When the database inevitably falls over, you'll be glad your control plane didn't go down with it.

I taught before I built, and it still shapes how I explain this work. I build production agentic AI systems and write about what I learn doing it.

Not sure if this is for you?

Take the readiness check on the practice site — see if it's a fit before you book anything.

two minutes

Check if it's a fit