Skip to main content

Blog Post

Build Log: What I Learned Letting AI Agents Run the Dev Lifecycle

•4 min read•By Brandon

Build LogAI EngineeringAgentic SystemsClaude

In Build Log #2 I described the mechanics of the SDLC harness: the pipeline stages, the isolated worktrees, the merge conflicts that broke parallel runs. This one is about the agents themselves — what I learned about how AI agents reason when you give each one a distinct, bounded role and hold them to real output.

The One-Job Principle

The most important decision I made was giving each agent exactly one job:

  • The planner decomposes a task spec into atomic sub-steps. It does not implement.
  • The implementer executes a plan. It does not review its own work.
  • The tester runs the real validation commands and reports the exact output — lint results, test failures, build errors — verbatim. It does not summarize.
  • The reviewer checks output against acceptance criteria and returns a verdict. It does not fix.
  • The documenter updates docs after the work passes. Not before.

What I found: when an agent is scoped to one job, it's honest. When the same agent both implements and reviews its own work, it finds the work acceptable almost every time. That's not a model limitation — it's the same problem you'd have asking a person to QA their own code on a deadline.

The Verifier Problem

The single most important agent is the tester, and the single hardest thing to get right is getting it to report real output rather than output it believes should be there.

Early versions would run a command, infer what the result probably was, and report a cleaned-up version. The lint passed. The tests passed. The build passed. And then review caught that the actual commands were never run.

The fix was specific: the tester's prompt requires verbatim command output. Exact error messages. Exact line numbers. If the commands weren't run, the output section stays empty and the verdict is FAILED. Agents that must show their work do their work.

How Specs Become Agent Instructions

The other thing that changed how I work: writing task specs is now the primary engineering activity. A spec is an instruction to an AI agent. If the spec is vague, the agent invents scope. If it's contradictory, the agent picks one interpretation and doesn't say which. If it's missing acceptance criteria, the agent decides when it's done.

The spec format I landed on: one sentence that states the goal, an itemized list of exactly what gets built, and a checklist of acceptance criteria that are falsifiable. "The UI looks good" fails. "The build passes with no TypeScript errors on the target routes" passes.

What Breaks Silently

The honest failure mode is optimistic reporting. An agent will mark a task complete when it has made a good-faith attempt — not when the work actually passes. The verifiable output gates exist to catch this. But the subtler version is a reviewer that grades on a curve: it notices the implementation is mostly correct and rates it PASS instead of flagging the part that isn't.

I added explicit instructions to the reviewer: if any acceptance criterion is not met, the verdict is PARTIAL, not PASS. PARTIAL triggers a fix pass with a specific list of what failed. This sounds obvious. It's not enforced by default.

The Meta-Angle

I'm building this harness inside Claude Code — the same production AI tool I'm writing about. Every task in the orchestration framework gets planned, implemented, tested, and reviewed by Claude agents running in git worktrees on my Mac Mini. The system is building itself, and when it breaks, I'm the human in the loop deciding whether to retry, escalate, or fix the spec.

That's the honest shape of agentic development right now: not autonomous, not manual, but something in between. You write the specs. The agents do the work. You review the seams.

Next: Build Log #4 — self-hosting on a Mac Mini

I'm documenting this in English and Portuguese at learn-agentic-ai.com. If you're building agent infrastructure or want someone who has shipped these systems, reach out at [email protected].

I taught before I built, and it still shapes how I explain this work. I build production agentic AI systems and write about what I learn doing it.

Not sure if this is for you?

Take the readiness check on the practice site — see if it's a fit before you book anything.

two minutes

Check if it's a fit