Skip to main content

Blog Post

Building AI Agents in Pure Python with the Anthropic SDK

•9 min read•By Brandon

AI EngineeringPythonAgent DevelopmentAnthropic SDKProduction AI

Before I built my Python orchestration framework, I spent time trying to understand every agent framework that existed. LangChain, AutoGen, CrewAI — I evaluated all of them. And every time I ran into the same problem: when something broke, I had no idea what was actually happening underneath.

The turning point was deciding to strip everything back and build directly against the Anthropic API. No abstractions, no magic. Just Python and import anthropic. What I found was that the "framework" I thought I needed was mostly just a few clean patterns repeated with discipline.

This post walks through those patterns — the same ones I use in production today.

Why pure Python? Direct API usage means you understand every token flowing in and out of your system. When latency spikes, when costs run high, when an agent gets confused — you have the full picture. Frameworks can come later; understanding comes first.

What you need to get started

Three things: a Python environment, the anthropic package, and an API key.

That's it. No vector databases, no framework config files, no DAG definitions. We'll build up from there.

Pattern 1: The direct API call

The foundation of everything is a single messages.create call. Before you write any agent code, you should be completely comfortable with this:

A few things to note about this API shape. The model parameter takes an exact string — claude-opus-4-8 — no date suffixes. The messages list carries the conversation history. And the response arrives in response.content[0].text for a basic text reply.

System prompts go in a top-level system parameter, not inside the messages list:

This separation matters. The system prompt sets the agent's persona and operating constraints. The messages carry the conversation. Keep them distinct.

Pattern 2: Structured outputs with Pydantic

Raw text responses are useful, but for agent workflows you almost always want structured data you can route, validate, and feed into the next step. Pydantic + the Anthropic SDK makes this clean.

The model returns valid JSON when you ask for it clearly in the system prompt and provide a schema. Parse it with Pydantic and you have type-safe structured data moving through your pipeline.

For stricter guarantees, you can pass response_format in the request and use client.messages.parse() to get automatic schema validation. I often prefer the explicit system prompt approach because it's easier to debug — the model's reasoning is visible in the raw JSON output.

Pattern 3: Tool use

Tool use is where agents get their real power. You define functions in Python, describe them to the model, and the model decides when and how to call them. Your code executes the function and returns the result. The model synthesizes everything into a final response.

Here's how to define a tool:

The key insight: the model doesn't execute your function. It generates a structured call request. Your code runs the function and feeds the result back into the conversation. The loop looks like this:

Call the model, check whether it wants to use a tool, execute the tool, append the result, call again. Repeat until stop_reason == "end_turn". That's the whole loop.

Pattern 4: Multi-turn conversation memory

The Anthropic API is stateless — every request is independent — so you maintain conversation history in your code. The simplest approach is a list of message dicts that you append to and pass on each call:

For long-running conversations, you'll want to prune the history to avoid hitting context limits. A simple strategy: keep the last N message pairs and any messages you've explicitly flagged as important.

Pattern 5: Sequential pipeline

Most production agent work isn't free-form conversation — it's a sequence of processing steps where each step's output feeds the next. Here's how I structure pipelines in my orchestration system:

Each step has a focused system prompt and passes relevant context from prior steps. The result object collects everything so downstream code can access any step's output.

Pattern 6: Routing

When you have multiple specialized agents, you need a router that classifies incoming requests and dispatches them. This is exactly what I use in my orchestration system — a lightweight classifier that keeps specialized agents focused:

Notice I'm using claude-haiku-4-5 for the classifier — it's faster and cheaper for a task that just needs to return one word. The specialized handlers use claude-opus-4-8 where the quality matters.

Putting it together: a triage and draft agent

Here's a minimal but real example — a support triage agent that classifies a ticket, searches documentation, and drafts a response:

Under 50 lines. Classifies, searches, drafts. That's a real agent.

When to reach for a framework

Pure Python is the right default for production systems where you need observability, control, and long-term maintainability. The patterns above handle the vast majority of agentic use cases I've encountered in practice.

Frameworks become worth considering when you genuinely need something they provide that would take significant engineering effort to build yourself — persistent agent state with managed infrastructure, complex multi-agent orchestration with automatic retry and failover, or pre-built integrations with dozens of external tools.

The rule I follow: if I can describe what the agent needs to do in a single paragraph, pure Python is probably the right tool. If the description fills a page and involves branching coordination between multiple specialized agents, I revisit.

Start with the patterns here. Build something real. The framework decision gets much easier once you understand what the frameworks are actually abstracting.


If you want to go deeper, the Anthropic Python SDK covers streaming, extended thinking, batch processing, and the newer managed agents surface. The patterns in this post are the foundation — everything else builds on them.

I taught before I built, and it still shapes how I explain this work. I build production agentic AI systems and write about what I learn doing it.

Keep going

Continue to the next module in this path.

Next moduleOr get updates by email