CCE SRE Mission Control
An agent-ready knowledge platform where every published fact traces back to its source.
- Role
- Founder and primary engineer
- When
- 2026 – present
- Stack
- Python · Claude · Databricks · Confluence · Jenkins · GitHub · Slack · Zoom
The knowledge platform for Nike's Cloud Core Engineering (CCE) Site Reliability Engineering (SRE) team. Team knowledge lives as reviewed Markdown in git, flows in through an LLM-assisted pipeline, publishes to Confluence, and is served to people and agents alike, including the production Slack agent CCE-SRE-Expert.
Problem
Operational knowledge for an SRE team decays quickly. Runbooks drift from reality, wiki pages contradict each other, and nobody can say where a “fact” came from. An LLM assistant built on top of that would only repeat the drift with confidence.
So when I created CCE SRE Mission Control in May 2026, I started from one premise: the team’s knowledge lives as Markdown in git, where people and agents can review it, and Confluence is the published view, not the source.
Approach
The platform is four layers, each built on the one below:
- A source of truth in two layers. Handbook, knowledge vault, and incident records as reviewed Markdown: a vault structured for LLMs, and a wiki for people, kept in sync by a round-trip converter with contract tests.
- A knowledge pipeline. Capture scripts turn a Slack channel, a Zoom transcript, or a Confluence space into a structured summary; an LLM proposes edits; a person accepts or rejects each one.
- A context layer. Per-repository
AGENTS.mdfiles, rules, skills, and a context-budget audit on every push, so any agent, in the IDE or in Slack, finds the right knowledge without flooding its context. - An agent runtime. CCE-SRE-Expert, the team’s production Slack agent, lives in the same repository and retrieves from both layers.
The first three layers would be valuable with no bot at all. The fourth depends on them entirely, which is why the platform, not the bot, is the product.
The payoff is the questions nobody planned for. During an audit of our most critical network devices, we couldn’t explain why a set of them had no instrumentation. The vault could: it traced the answer to a meeting recorded a month earlier, which the pipeline had captured.
What I got wrong along the way. Early on, I let the vault copy measured figures into its entries. They pushed the context budget into the red and kept being cited as current long after they weren’t, so the vault changed shape to pointers. And the validator I trusted most was the one hiding the most: it had been rewriting malformed metadata so it could keep running, which meant three quarters of the vault was invisible to anything reading it.
Architecture
A four-stage knowledge pipeline with an audit trail
-
Step 1: Capture
- Source: Captured knowledge Slack, Zoom, Confluence
-
Step 2: Propose
- Model: LLM-proposed edits
-
Step 3: Review
- Person: Human review
-
Step 4: Publish
- Data store: Markdown vault 196 entries, for agents
- Tool or service: Wiki 160+ pages, for people
Round-trip Confluence ⇄ Markdown with contract tests; drift detection on every publish.
- Source
- Model
- Person
- Data store
- Tool or service
Key decisions
-
Two layers, one source
A vault structured for LLMs and a wiki written for people, generated from the same reviewed Markdown so they can't drift apart.
-
LLMs propose, people approve
Nothing reaches the wiki without review, and every fact keeps a trail back to the message it came from, so reviewers can check it and the agent can cite it.
-
Pointers, not values
The vault holds what a service is and where its live data lives, not copies of the data. A copied value goes stale silently; a pointer fails visibly.
-
Checks that can fail
A lenient validator had hidden broken metadata in 135 of 177 files by quietly fixing it. Every check now reads files exactly as written.
-
Guard what agents write
Agents produce data faster than reviewers read it, so a sensitive-data scanner blocks every push, gates publishing, and runs again in CI.
Results
- used by the whole ~15-person team, for nearly everything
- Daily
- from an empty scaffold to the team's source of truth
- ~4.5 months
- wiki pages from 196 vault entries
- 160+