Decision · Jul 2026
From single-shot RAG to an agentic tool loop
Every eval failure traced to one gap, that the bot couldn't move around inside what it retrieved. Retrieval became tools the model calls, and citation F1 went from 0.448 to 0.498 on the first 34 cases.
Context
CCE-SRE-Expert began like most internal assistants: embed the question, retrieve the top chunks, answer. On July 8, 2026 it got real semantic retrieval (Databricks Vector Search, BGE-Large embeddings, a chunker that splits Markdown at headings, rebuilt on every merge), and the eval harness got a Slack-bot surface to measure it.
Evidence
The single-shot baseline scored a citation F1 of 0.448 across 34 cases. The shape of the failures was more useful than the number:
- Procedural questions: right answer, adjacent citation. The bot picked a sibling document in the same directory.
- Plane-local questions: it cited the people manifest instead of the entry the manifest pointed to.
- Cross-plane questions: it answered from one chunk when the question needed two or three.
- Repo-wide questions: it cited scaffolding files instead of the document that actually held the synthesis.
Every one traced to the same gap: the bot couldn’t move around inside what it had retrieved. It couldn’t re-query, follow a link, read a whole file, or check its answer. A coding agent in the IDE does all four implicitly through its own tool loop.
Options considered
- Keep single-shot RAG and tune chunking and ranking.
- An agentic tool loop: give the model retrieval tools and let it decide when it has enough.
Decision
On July 9 I wrote a design note and made the change the same day. Retrieval became four tools the model calls when it decides it needs them: vector search, file reads, grep, and directory listing. Vector search stayed as one tool of four rather than being replaced. The old single-shot pipeline stayed as a runtime fallback if the loop fails or hits its limits, and the loop had limits on tool calls, input tokens, and wall-clock time from the start.
The rerun scored 0.498: a modest improvement, and a measured one.
Consequences
- Over the next day the loop gained Jira, GitHub, and Confluence tools.
- A loop costs more per question than one retrieval, which is part of why the cost analysis ranks ten levers, with prompt caching alone estimated to cut input cost 70–85%.
- The harness also stopped a feature: an MCP server over the same index was gated on semantic search clearly beating keyword search, and since that comparison was never run, it stayed gated.
- Production later showed that limits need honest reporting: one failure blamed the tool-call limit when wall-clock time was the one firing.
Where it applies
- Nike In production
CCE-SRE-Expert
CCE SRE Mission Control's production agent: it answers with citations, debugs as well as it retrieves, and writes back only through a person.
2026 – present