Skip to content
James Kirkman , home

Decision · Jul 2026

Build vs. buy for an enterprise agent platform, decided by evals

Proposed At work

An internal platform overlapped almost completely with what I'd built. Instead of defending my own system, I wrote down the eval result that would make me abandon it.

Context

In mid-July 2026 an internal, agent-first chatbot platform came onto my radar: a team supplies a configuration and gets a Slack assistant backed by hosted RAG and MCP tools. The overlap with what I’d built was nearly complete. We’d independently picked the same model endpoint, the same vector search engine, the same embedding model, and the same Slack integration pattern. That was reassuring and uncomfortable in equal measure.

Build-vs-buy decisions like this usually get made on feature matrices, demos, and who built what.

Options considered

  • Buy: move the assistant onto the platform.
  • Build: keep evolving the in-house agent.
  • Measure: run both against the same eval cases before deciding.

Decision

I wrote the comparison up honestly. The platform would remove the hosting and ingestion work I owned and add things I didn’t have, such as escalation hooks into the ticketing system. We’d give up the bot’s direct access to the repository’s files, custom message handling, control over releases, and some confidentiality over where the knowledge is indexed.

I recommended a side-by-side eval run: add the platform as a surface in the existing harness, run the same cases against both, and apply written decision gates (adopt, hybrid, don’t adopt, or revisit in ninety days) depending on the scores and on privacy and service-level answers. Whatever the outcome, I recommended publishing the knowledge vault as an MCP server, so other teams’ assistants could use it whichever runtime wins.

Consequences

  • The decision criterion is the same one used everywhere else: citation quality, hallucination rate, and cost per configuration.
  • Exposing knowledge over MCP decouples the knowledge investment from the agent investment. The platform could have replaced the agent’s runtime, but not the knowledge platform underneath it, which is why publishing the vault made sense whichever way the eval went.
  • Nike In production

    CCE-SRE-Expert

    CCE SRE Mission Control's production agent: it answers with citations, debugs as well as it retrieves, and writes back only through a person.

    2026 – present

  • Nike In production

    CCE SRE Mission Control

    An agent-ready knowledge platform where every published fact traces back to its source.

    2026 – present