CCE-SRE-Expert
CCE SRE Mission Control's production agent: it answers with citations, debugs as well as it retrieves, and writes back only through a person.
- Role
- Architect and primary engineer
- When
- 2026 – present
- Stack
- Python · Claude · Databricks Model Serving · Databricks Vector Search · MCP · Slack · Jira · GitHub
The agentic Slack assistant built into CCE SRE Mission Control for Nike's Cloud Core Engineering (CCE) Site Reliability Engineering (SRE) team, on Claude via Databricks Model Serving. It retrieves its own context from the platform's vault and wiki, cites its sources, and acts on Jira, Confluence, GitHub, and the repo through a human-in-the-loop write path.
Problem
The team’s knowledge was spread across Confluence, repositories, tickets, and people’s heads. Answers to operational questions had to be grounded and cited, not plausible. And an assistant was only worth building if it could also do things, like file tickets, fix docs, and open pull requests, without becoming a liability.
The agent is the runtime layer of CCE SRE Mission Control: its code lives in the same repository, its knowledge comes from the platform’s vault and wiki, and its write path is how teammates update the platform from Slack.
Approach
- Simple first, then measured. The first version loaded 113 documents at startup with keyword search. Semantic retrieval and an eval harness followed two weeks later, and the failure analysis drove the move to a tool loop: vector search over an index rebuilt on every merge, plus whole-file reads, grep, and directory listing, so the agent can look again until it has enough. That’s also why it can debug a customer’s code instead of pointing them at a document.
- Live data over copies. Instance-type and policy questions are answered live from the company’s cloud-policy MCP server over OAuth client credentials, not from files someone has to keep current. If a live call fails, the agent falls back to the enforced policy, never to a copy.
- Writes a person approves. Ask in Slack, and the agent files a Jira ticket, stages the edits, shows a diff, and opens a draft pull request on “ship it.” It never merges, and runtime re-checks, path block-lists, and per-thread caps back up the schema.
- Every miss becomes a test. Each production miss becomes an eval case, and any change the agent proposes to its own code has to come with a test.
What I got wrong along the way. I fixed the same misbehavior with prompt changes four times before accepting that a prompt is a strong suggestion, not a guarantee, and moving the rule into the harness. And when the agent kept stopping short of finishing a fix, I raised its tool-call limit step by step before discovering that the limit actually firing was wall-clock time: the error message named the wrong one. The agent now reports which limit it hit.
Architecture
Ask in Slack, answer with citations, ship through a draft PR
-
Step 1: Ask
- Person: Slack thread
-
Step 2: Reason
- Model: Agentic tool loop Claude via Databricks Model Serving
- Check or gate: Permission-scoped tools write tools only in the approved channel
-
Step 3: Retrieve
- Data store: Vector search BGE-Large · ~2,200 chunks
- Tool or service: File reads + grep whole documents
-
Step 4: Check
- Check or gate: Post-answer verifier
-
Step 5: Act
- Person: Cited answer
- Tool or service: Jira → diff → draft PR on "ship it"
The model re-queries until it has enough; the index is rebuilt on every merge.
- Person
- Model
- Check or gate
- Data store
- Tool or service
Key decisions
-
Agentic loop over single-shot RAG →
On the same 34 questions, every failure traced to the bot being unable to look again. A tool loop raised citation F1 from 0.448 to 0.498, modest and measured.
-
Safety by schema →
Write tools are withheld outside the approved channel. A tool the model never sees can't be misused, and the smaller schema is cheaper and caches better.
-
Single agent until the data says otherwise →
Our own six domain experts behind a router modelled at ~12× the token cost, for none of the conditions where multi-agent pays.
-
Guardrails in the harness →
When a prompt rule fails more than once and the harness can check it mechanically, it moves out of the prompt and into code.
-
Leave out what doesn't belong in the channel
Of six tools on the cloud-policy MCP server, the agent gets five. The sixth returns per-account cost data, and a test pins its absence.
Results
- of customers helped correctly without an SRE stepping in (33 of 40 over 60 days)
- 82%
- used by internal customers
- Almost daily
- tools, with a human-in-the-loop write path
- 20+