Skip to content
James Kirkman , home
At home Active

Ascendant's Archipelago: the game

An Unreal Engine 5 open-world RPG built by coding agents, with me as creative director.

Role
Creative director (agents implement)
When
Sep 2026 – present
Stack
Unreal Engine 5 · C++ · Python · Blender · Git worktrees · Claude Code · Whisper

An exploration-driven, systemic open-world RPG. Claude coding agents implement; I direct, playtest, and decide canon, scope, and art. In thirteen days, 760 commits from parallel agent lanes on two machines, held together by verify gates the agents can't skip and a process that measures its own time, flakes, and token cost.

Problem

Could one person direct a game with real systemic depth, using coding agents as the implementation team, without the codebase or the canon falling apart?

Approach

  • Clear roles. The repository’s instructions tell agents I’m “the creative director, not the implementation engineer.” Agents implement and decide technical matters; canon, scope, and art come to me, and canon lives in a separate worldbuilding vault the game must follow.
  • Agents that check their own work. One automation entry point, playtest scenarios that drive the real game, and verify gates a merge can’t skip. My feedback arrives the same way: one key records a screenshot, the game state, and my voice, transcribed locally and triaged into tasks with an acceptance check. The first session found the whole map was mirrored, which no test could see because every test shared the same wrong frame.
  • Measure the factory. Timing everything overturned my assumptions twice: start-up was slow because of per-line log flushing, not rendering, and the limit of three concurrent Unreal processes I’d asked for was costing 55 minutes a day while six showed no slowdown. A full verify fell from about 45 minutes to 13, then climbed back to a median of 44 as the suite grew to 49 scenarios. Every round also reports tokens per agent task.
  • Parallel lanes, made cheap enough to use. Worktree lanes are the parallel-work case where multi-agent earns its cost, but merging them was expensive. One command now merges a batch, regenerates whatever the merge touched, runs one full verify, and pushes only on a pass. A passing verify is stored on its commit as a git note, so the laptop verifies and the desktop merges on its record in about 3 minutes instead of a median 44.
  • Two machines, one pusher, and a heartbeat. Twice a sleeping laptop looked like slow work, so it now pushes a heartbeat every ten minutes and the desktop alarms when it goes stale. After a crash with four Unreal instances running, Unreal is handed out in slots.
  • Failures become tooling. A silent microphone cost a whole session of notes, so recording now warns when the input stays silent. Problems that only appeared when lanes met became automatic steps in the merge.
  • Intelligence, measured and then switched off. A tiered decision layer for development decisions scored 97% with no missed escalations on 36 real cases. When the project standardized on Claude, the fast model was switched off by default; the tiering principle and the interface stayed.

Architecture

A playtest becomes merged, verified work

  1. Step 1: Direct

    • Person: Playtest notes voice + screenshot + game state
  2. Step 2: Plan

    • Model: Coordinator agent triage into tasks and issues
  3. Step 3: Build

    • Model: Claude agent lanes per round, or per big feature
  4. Step 4: Verify

    • Check or gate: Quick verify only what a change can affect
    • Check or gate: Full verify 49 playtest scenarios, once per batch
  5. Step 5: Merge

    • Data store: Integration branch one pusher, two machines

Every round ends with a check of time, flakes, and tokens; the process changes when the numbers say so.

  • Person
  • Model
  • Check or gate
  • Data store
I playtest and record notes by voice with a screenshot and the game state. A coordinator agent triages the notes into tasks and GitHub issues. Claude agents build them in worktree lanes, a short-lived lane per playtest round and long-lived lanes for big features, across two machines. Every change passes a quick verify; one command merges a batch of lanes under one full verify of headless and rendered playtest scenarios, with one pusher at a time. Passing verifies are shared between machines as git notes. Every round ends with a check of time, flakes, and token cost.

Key decisions

  • Agents decide technical matters; I decide canon

    Agents may tidy the canon vault but change canon only after approval. The vault changes first, and the game follows.

  • Measurement before every other change

    Two days of artifacts showed 933 playtest runs and nothing recording durations. Timing came first, and every optimization after it was measured.

  • Lanes, merged small and often →

    Lanes first cost a verify each and a sync after every merge. Once one command merged a batch under one verify, and verifies were shared between machines, 12 lanes merged 49 times in five days.

  • Merge what passed, hold what didn't

    A higher jump broke the city lane's dock fight. The city merged without it, and the jump shipped hours later with the ledge grab that made it safe.

  • Test the real input path

    Quick save passed every playtest and failed in play, because my laptop's F5 and F6 never reached the game. A key logger found it.

Results

commits in 13 days, from 12 parallel agent lanes
760
playtest scenarios gating every merge
49
architecture decision records
42
a merge's full verify, reusing a lane's shared record
~44 → 3 min