Decision · Sep 2026
Tiered intelligence, from deterministic code to a human
Route each development decision to the cheapest tier that answers it reliably. On 36 real escalation decisions: rules alone 56%, a fast decision model 94%, rules then model 97% with no missed escalations, for US$0.0015 in total.
Context
A game built by coding agents makes thousands of small development decisions: accept or revise, which test to run next, whether a result looks suspicious, whether a change needs the creative director. Sending all of them to a frontier model is slow and expensive. Hard-coding all of them is brittle.
This is about the development process, not the game. Nothing in the game calls the decision layer.
Decision
Route each question to the cheapest tier that answers it reliably:
- Deterministic code for anything the rules can settle. (“A file exists” is code. “The packaged build launches” is a test.)
- A fast, bounded decision model for the rest.
- A frontier reasoning model when the bounded one isn’t enough.
- A human for canon, scope, and art.
Interfaces are model-neutral: the repository depends on a decision interface, not on any vendor. Providers for deterministic rules, recorded replays, the bounded model, and a human are implemented. The frontier-model slot is defined but not built yet.
Evidence
The first task measured was the one that matters most: does this change need the creative director? The 36 cases came from decisions actually recorded in the project’s briefs, 22 made by agents and 14 by me.
| Provider | Coverage | Accuracy | Missed escalations |
|---|---|---|---|
| Rules alone | 22% | 56% | 0 |
| Decision model | 100% | 94% | 0 |
| Rules, then model | 100% | 97% | 0 |
Both of the model’s errors were false escalations, the safe direction. The model’s answers cost US$0.0015 for all 36 cases, at a median 198 ms each, and they’re stored with the corpus, so the eval replays offline for free.
Consequences
- “No missed escalations” is a pass/fail criterion of its own, separate from accuracy.
- The next day, the project standardized on Claude for its coding agents, and the decision model was switched off by default. The tiering principle and the interface stayed; the model can be switched back on per task.
- An agent benchmark protocol for comparing coding agents was written on day one and never run. The vendor question was settled by use instead, and the decision record says so.
Where it applies
- At home Active
Ascendant's Archipelago: the game
An Unreal Engine 5 open-world RPG built by coding agents, with me as creative director.
Sep 2026 – present