When I started building the support-automation system at my last engagement, I ran into a problem that I suspect a lot of AI engineers eventually hit: the system needed to do too many things, and every approach I tried made it worse.
The workflow was genuinely complex. Incoming support tickets needed to be read and classified. Urgent issues had to trigger immediate notifications to the right team. Bug reports needed cards created in the project management tool. Routine questions needed routing to a queue, and specific ticket types needed to go to specific squads. Each of those steps touched a different platform, a different API, a different data shape.
My first instinct was to build one big agent that handled all of it. That approach lasted about two weeks before it collapsed under its own weight. Every prompt edit to improve classification broke the notification logic. Every API change rippled through a monolithic function that was hard to test and harder to reason about. I needed a different architecture.
MCP — Model Context Protocol — turned out to be exactly the right tool. Here's what I built, and why the distributed design made everything more maintainable.
Why MCP fits this problem
The core insight behind MCP is that AI agents work best when their tools are well-defined and composable. Instead of one agent with access to everything, you define focused MCP servers — each responsible for a specific integration domain — and let the orchestrating agent decide which tools to call.
For a support-automation workflow, this maps naturally:
- One MCP server owns the support platform integration (reading tickets, updating status, moving to queues)
- One MCP server owns notifications (Slack messages, team alerts)
- One MCP server owns project management (creating bug cards, assigning issues)
- The AI agent reads tickets through one server and coordinates responses through the others
This separation meant I could test each server independently, update the Slack integration without touching the triage logic, and reason clearly about what each piece was responsible for.
Building the triage MCP server
The support platform server is where tickets enter the system. It exposes both resources (ticket data the agent can read) and tools (actions the agent can take). The typed schema constants from the MCP SDK are the key to getting this right:
Using the typed schema constants (ListResourcesRequestSchema, CallToolRequestSchema, etc.) is important — they give the SDK proper type safety for routing and response validation. Early in my prototyping I was using string-based handlers, which worked but lost the type information and made debugging harder.
Resources expose the ticket queue as readable data:
Tools expose the actions the agent needs to perform after making a triage decision:
The notification server
The Slack notification server follows the same pattern — same typed schemas, focused scope. Its job is simple: send the right message to the right channel when something needs human attention.
The triage agent: Anthropic SDK connecting everything
The agent that ties everything together uses the Anthropic SDK for reasoning and the MCP SDK to connect to the servers. For this kind of classification-and-routing task, claude-haiku-4-5-20251001 was the right model choice — fast, accurate for structured classification, and inexpensive enough to run on every incoming ticket.
The tool definitions the agent sees come from the MCP servers. In the production system I wrote an adapter layer that translates MCP's ListToolsResponse into Anthropic's tool definition format — the schemas are compatible, it's mostly field renaming. The agent picks which tool to call based on its analysis; the MCP server validates the input and executes it.
What the architecture enabled
After the refactor, a few things got noticeably better.
Testing became tractable. Each MCP server could be tested in isolation with mock data. The triage agent could be tested with a mock MCP server that returned predictable responses. Before the refactor, testing required standing up the entire system.
Prompt iteration didn't break integrations. When I wanted to improve the agent's classification accuracy for billing-adjacent technical tickets, I only edited the triage prompt. The Slack notification logic and the project-management server were unaffected.
Failure modes became localized. When the project management API had an outage, triage and notification kept working. The agent logged that bug card creation failed and moved on. A monolithic design would have been much harder to make that resilient.
The tradeoff is initial complexity — setting up three separate servers with their own transport and initialization is more boilerplate than a single service. But for a workflow that runs continuously and needs to evolve as team needs change, that investment paid off quickly.
The practical pattern
If you're building a multi-platform support automation workflow, here's the approach I'd recommend:
- Draw a boundary around each external system. Each boundary becomes an MCP server.
- Keep each server focused — resources expose the data the agent needs to read, tools expose the actions it can take. Don't let one server do everything.
- Use the typed schema constants (
ListToolsRequestSchema,CallToolRequestSchema) from the start. The type safety pays off when you're debugging why a tool call returned the wrong shape. - Choose your model size deliberately. For structured classification on known categories, a smaller, faster model is usually the right call. Reserve the heavy reasoning for workflows where the agent needs to synthesize novel responses.
The support-automation system this came out of handled production traffic, and the architecture held up under real load. MCP's real strength is that it gives you a principled way to give AI agents access to external systems without turning your codebase into a tangle of API calls mixed with prompt logic. That separation is worth the setup cost.