In Build Log #1 I described the orchestration framework I'm building. This one is about the thing that builds it: a multi-agent SDLC harness. The framework and the harness develop each other. The system is building itself.
What It Is
The harness runs a task through the full software development lifecycle with AI agents at each stage:
plan → implement → test → review → document
Each stage is a distinct agent with a distinct job. The planner decomposes a task spec. The implementer writes the code. The tester runs the real validation gates — lint, type-check, unit tests, build — and reports actual output, not a summary it invented. The reviewer checks the work against the acceptance criteria and returns a verdict. The documenter updates the docs only after the work has passed.
There are real checkpoints between stages. A stage can fail, and failure is a first-class outcome, not an exception to paper over.
Isolated by Default
Every task runs in its own git worktree. That single decision unlocks parallelism: multiple tasks can run at the same time without stepping on each other's files, each on its own branch, each with its own clean checkout. When a task passes review, its branch merges back in dependency order.
The orchestrator drives a whole spec as dependency-ordered waves of parallel tasks, with bounded retries and an escalation path when an agent can't recover on its own.
What Broke
The interesting failures were not in the agents — they were in the merges.
When I ran a wave of parallel tasks that each appended to the same auto-generated doc files, git couldn't auto-reconcile them. Every task had legitimately edited the same lines. Same story with a package's __init__.py: each service rewrote it with only its own export, so the merges collided. The agents all did their jobs correctly and the integration still failed.
There was also a subtler one: a dependency-cycle guard aborted a run because the analysis step had marked an append-only file as a hard dependency. Logically task 4 depended on task 7, but the conflict-serialization saw "7-after-4" and called it a cycle.
The fixes were unglamorous and exactly right: treat append-only files (__init__.py, generated docs/*.md) as additive by default so a union merge is safe, and hand-author the execution plan when the auto-analysis is too conservative. The lesson is that multi-agent systems fail at the seams. The agents are the easy part now. The integration is the engineering.
This post was itself produced by the harness — planned, implemented, tested against the content validator and a production build, and reviewed before it shipped.
Next: Build Log #3 — what I learned letting AI agents run the dev lifecycle
I'm documenting all of this in English and Portuguese at learn-agentic-ai.com. If this is the kind of system you want built, reach out at [email protected].