Skip to main content

Blog Post

How I Built a Knowledge Graph for My Company Brain

•6 min read•By Brandon

AI EngineeringPythonKnowledge ManagementAgentic Systems

A few months into running my solo agentic engineering practice, I hit a problem I hadn't anticipated: my documentation was growing faster than I could navigate it.

I had a CLAUDE.md for each project, a planning/ directory with decision records and status files, a blog content backlog, cross-project architectural decisions, and a changelog. All of it was accurate. None of it was connected. When I needed to answer "why did I choose pgvector over Pinecone for the RAG pipeline?" I had to grep through four directories across two repos.

The solution I built is what I call the Company Brain knowledge graph — a documentation-ingestion pipeline that reads all my planning files, extracts entities and relationships, and stores them in a structured graph that any agent in my system can query.

The problem with flat documentation

Flat Markdown files are easy to write and easy to store. But they don't scale well when the interesting information lives in the connections between documents.

My support-automation system references a de-identified customer help-desk. That same system is the subject of several blog posts, a project card on my portfolio site, and a learning module about MCP servers. If I'm writing a new blog post and I want to know every place I've described that system publicly, a flat file search gives me filenames. A knowledge graph gives me a typed relationship: (BlogPost "mcp-customer-support") -[DESCRIBES]-> (System "support-automation").

That distinction matters when you're building agents that reason about your own work. An agent that can traverse relationships instead of just scanning text is a fundamentally more useful assistant.

Architecture of the ingestion pipeline

The Company Brain knowledge graph has three layers:

  1. Document ingestion — read Markdown and JSON files, extract structured content
  2. Entity extraction — use Claude to identify entities (systems, decisions, posts, projects) and their relationships
  3. Graph storage — persist edges and nodes for retrieval

I built this on top of my Python orchestration framework, which already had a task-queue model I could reuse. Each file in the planning directory becomes an ingestion task.

Here's the core ingestion loop:

Handling cross-repo relationships

My practice spans two repos: learn-ai (the portfolio site) and python-orchestration-system (the agentic framework). The Company Brain repo holds cross-project context that references both.

The challenge is that a decision in python-orchestration-system/planning/decisions/ might reference a feature visible only in learn-ai/content/projects/. My ingestion pipeline needs to resolve these references without merging the repos.

I solved this with a canonical ID scheme: every entity gets a prefixed ID that encodes its repo and path. A system called support-automation in the Python orchestration repo becomes python-orch:system:support-automation. The portfolio's project card for the same system becomes learn-ai:project:ai-chatbot-rag. The Company Brain stores a SAME_AS edge connecting them.

Graph storage and querying

I store the graph in a structured JSON file that any agent can read. For my current scale — a few hundred nodes across two projects — this is fast enough. If the graph grew significantly I'd move to a proper graph database, but right now the overhead isn't worth it.

The key thing I built is a relationship traversal function that lets agents ask questions like "what does this decision affect?" or "which blog posts describe this system?":

How I use this in practice

The knowledge graph powers a few specific workflows in my day-to-day:

Before writing a blog post, I query the graph for all entities related to the post's topic. This surfaces prior decisions, existing content on the same subject, and project cards I should link to — without needing to remember where each piece lives.

When creating a new decision record, the graph tells me which existing systems the decision affects and flags any conflicts with prior decisions on the same topic.

When an agent is working on spec tasks, it can query the graph to understand the broader context of the code it's touching. A spec for the MCP learning path, for instance, can automatically pull in the project card for the support-automation system and the prior learning module decisions.

The real payoff is that I stopped losing context. When you're running multiple long-running projects in parallel, the knowledge graph becomes the connective tissue that keeps everything coherent — not just for me, but for every agent working on my behalf.

What I'd do differently

The current extraction prompt works well but produces inconsistent relationship types. "USES", "DEPENDS_ON", and "REFERENCES" end up describing very similar things in different documents. I'm planning to define a strict ontology upfront and pass it as part of the system prompt so Claude picks from a controlled vocabulary.

I also haven't tackled incremental updates yet. Right now the pipeline re-ingests everything on each run. The next step is diff-based ingestion: track file modification timestamps and only re-process files that changed since the last run.

If you're running a multi-project practice or an agentic system with significant planning documentation, building a knowledge graph is worth the setup time. The compounding benefit of connected context scales with every document you add.

I taught before I built, and it still shapes how I explain this work. I build production agentic AI systems and write about what I learn doing it.

Not sure if this is for you?

Take the readiness check on the practice site — see if it's a fit before you book anything.

two minutes

Check if it's a fit