Skip to content
James Kirkman , home

Decision · May 2026

Single agent plus structured context, until the data says otherwise

Adopted At work

A router-to-six-experts topology modelled at ~12× the token cost of a single agent with retrieval, so multi-agent became a gated phase, not a starting point.

Context

Multi-agent architectures are an easy sell: a router, a handful of domain experts, an orchestrator stitching answers together. For a site reliability engineering (SRE) assistant covering several domains, a “router to six experts” design was the obvious candidate.

Options considered

  • Single agent with retrieval and structured context. One model, one tool loop, domain knowledge supplied as context.
  • Router to per-domain experts. A classifier routes each question to one of six specialist agents.
  • Orchestrator–worker. An orchestrator fans out to subagents and merges their results.

Evidence

  • Cost model first. A dry-run estimate put a router fanning out to six plane experts plus a synthesizer at roughly 12× the token cost of a single agent with retrieval.
  • Published guidance. Anthropic’s guidance on building effective agents makes the same point: a single agent with good retrieval usually beats an orchestrator–worker design at a fraction of the cost, and multi-agent shapes pay off on parallel, breadth-first work rather than tightly coupled work.
  • No context problem yet. Paying 12× to solve a problem I didn’t have would have been a guess.

The decision wasn’t just “no.” Three instruments were built the same week, so it could be reopened on evidence instead of opinion:

  • A context-budget audit that runs on every push and estimates how much of the context window a realistic session uses, with thresholds informed by the “Lost in the Middle” research on long-context degradation. Its first reading was 13.8% of the day-to-day budget.
  • A six-configuration eval harness: two surfaces (Cursor SDK and Anthropic API) and three topologies (single agent, routed, orchestrator with plane experts), starting with 25 versioned cases in four categories. It has since grown past 200 cases.
  • Structured retrieval instead of more agents: per-directory AGENTS.md files, a vault index, and a routing skill that loads the right domain’s context into one agent rather than handing the question to another.

One honest footnote: when it was built, the harness couldn’t run, because no inference surface was available to our accounts yet. It waited, built and ready, until the bot’s hosting decision in June gave it a model to call.

Decision

Adopt the principle “single agent plus structured context until the data says otherwise.” Capture multi-agent as a phased, reversible roadmap: single agent with structured context → selective RAG → subagent fan-out → per-domain expert agents. Each phase is gated by a measurable trigger, not by enthusiasm.

When multi-agent does pay

Saying no to this design isn’t saying no to multi-agent. There are four situations where more agents earn their cost:

  • Parallel, independent work. Subtasks that share little state and merge cheaply, like an IDE spawning a subagent to research while another edits, or the parallel lanes building my Unreal game. Tightly coupled work just turns into coordination and merge conflicts.
  • Context isolation. A subagent can read a 10,000-line log or a whole repository and hand back three sentences, keeping the caller’s context clean.
  • Independent review. An agent that didn’t write the work checks it without the author’s assumptions. In a research benchmark I’m building, a second agent from another vendor found spec mismatches that the first agent’s passing tests couldn’t, because those tests encoded the first agent’s reading of the spec.
  • Ownership and permission boundaries. Another team’s agent holds the permissions, data, and domain knowledge, so you call it like a service instead of copying what it knows. It’s least privilege, applied to agents.

The router-to-six-experts design had none of them. Questions weren’t parallel work, the context-budget audit read 13.8%, and my team would have owned all six experts, so there was no boundary to respect. It would have paid the multi-agent cost for none of the benefit.

Later, domain teams began building their own expert agents. That’s the same pattern with the boundary made real, and calling an agent another team owns is a design I’d support. Same topology, opposite verdict, because the boundary changed.

Consequences

  • Topology changes now need a cost model and an eval delta before they’re proposed.
  • The roadmap keeps multi-agent available, so the decision is reversible and not dogma. Each phase has a measurable trigger, a scope, and an exit criterion.
  • So far the answer is still no. I run one agent per task, several tasks in parallel, each in its own worktree, and use sub-agents inside a task only when the work genuinely fans out.
  • Nike In production

    CCE-SRE-Expert

    CCE SRE Mission Control's production agent: it answers with citations, debugs as well as it retrieves, and writes back only through a person.

    2026 – present