Skip to main content

Blog Post

Build Log: An Event-Driven Orchestration Framework in Python

•3 min read•By Brandon

Build LogAI EngineeringAgentic SystemsPython

The hardest part of building multi-agent systems isn't the LLM calls — it's knowing what went wrong and where. I built this framework because I kept watching agent runs fail in ways I couldn't diagnose. This is the first in a series of build logs on it: here's what I built, here's why, here's what broke.

What It Is

I'm building an agentic orchestration framework from scratch in Python. The core is a node-based workflow engine: you wire up LLM calls, tools, parallel branches, and routing logic the same way you'd describe a process on a whiteboard.

A workflow is a graph of nodes. Each node does one thing — call a model, run a tool, route to one of several branches, fan out into parallel work — and passes a typed context object to the next. Nothing is implicit. When something goes wrong, you can point at the exact node where it happened.

Why Event-Driven

Most agent frameworks hide the control flow. You hand them a goal and a bag of tools, and they loop internally until they decide they're done. That's fine for a demo. It's miserable in production, because when it fails you have no idea where it failed.

I went the other way. The engine is explicit and inspectable:

  • Typed context. Every node receives and returns a TaskContext. The shape is known at every step.
  • A validator that runs before execution. A WorkflowValidator checks the graph is well-formed before a single model call is made — no dangling routes, no orphan nodes.
  • Real router and parallel primitives. RouterNode for branching, ParallelNode for fan-out, both as first-class, tested building blocks rather than ad-hoc loops.

The design goal is that you can read a workflow and know what it does, and when it breaks you can see the break.

What's Built So Far

The foundation is in place and tested. So far that includes:

  • The core engine — workflow schema, validator, task context, and the base node types.
  • A ToolUseNode built directly on the raw model SDK, so tool calls aren't buried under a wrapper.
  • A pgvector-backed data layer with shared services for embedding, transcript extraction, article extraction, search, and chunking — the reusable parts every workflow needs.
  • A full unit suite. Right now the project runs at 210 passing tests with a clean linter.

What Broke

The honest part. Bringing the shared services online surfaced a class of problems that had nothing to do with the model and everything to do with plumbing: import-time side effects that made tests order-dependent, a repository method that returned ghost rows, and a router that coupled its keys too tightly to one caller. None of these are glamorous. All of them would have bitten me in production. Fixing them first is the whole point of building the boring foundation before the interesting agents.

Next: Build Log #2 — the multi-agent SDLC harness that builds this framework

I'm documenting all of this in English and Portuguese at learn-agentic-ai.com. If you're building something similar — or want this kind of work done — reach out at [email protected].

I taught before I built, and it still shapes how I explain this work. I build production agentic AI systems and write about what I learn doing it.

Not sure if this is for you?

Take the readiness check on the practice site — see if it's a fit before you book anything.

two minutes

Check if it's a fit