Skip to content
James Kirkman , home

Decision · Aug 2026

Guardrails move from the prompt to the harness when prompting has plateaued

Adopted At work

After four prompt fixes for the same failure, the fifth was a post-answer check the model can't argue with, plus a rule for deciding which other guardrails should follow.

Context

CCE-SRE-Expert kept telling people to post in the team’s support channel while it was already in that channel. Between late July and late August 2026 I fixed it with prompt engineering four times: a phrase blacklist, a categorical rule, and a stronger rule that held even when a cited source said otherwise. On August 27 it did it again, with every one of those instructions in the prompt.

The rule wasn’t under-specified. The model’s own probabilities still put weight on “here’s a channel that handles this,” and no wording reliably drives that to zero.

Options considered

  • A fifth rewrite of the instruction.
  • A post-answer verifier. The harness knows the current channel before the model writes a word, so it can check the answer and retry if the rule was broken.

Decision

Ship the verifier, and stop treating guardrails one incident at a time. I inventoried every guardrail in the roughly 830-line system prompt and scored each on whether it can be checked mechanically, whether prompting has plateaued, and how bad the failure is:

  • Move to verifiers: manifest follow-through, coverage of every incident in a thread, Slack formatting.
  • Stay in the prompt: semantic rules, such as “cite a source or stay silent.”
  • Move to the harness’s state machine: rules about the order of actions, such as “ticket first” and “confirm before writing.”

The rule: a guardrail migrates when it’s deterministic and has failed more than once in the prompt.

Consequences

  • Prompt rules stay the cheap first resort; verifiers are for known leaks.
  • The bot’s engineering is now described on four layers: prompt, context, harness, and loop. The loop is a single ReAct-style loop on purpose, with no planner, critic, or sub-agent dispatch until an observed failure earns one.
  • Nike In production

    CCE-SRE-Expert

    CCE SRE Mission Control's production agent: it answers with citations, debugs as well as it retrieves, and writes back only through a person.

    2026 – present