Skip to content

Status: Work in progress. S0–S2 sealed. S3 oracle in development: its first full S3 run succeeded and replays byte for byte (Oct 4, 2026); no S3 gate yet. Premiere goal: Tail Cave through collection of the Full Moon Cello.

Track:TourOperatorBuilder

War stories

A blameless postmortem is an incident write-up that asks “what happened and how do we prevent it?” rather than “whose fault was it?” Every story here comes from the project’s own task reports. Where the records don’t say something, this page says so instead of filling the gap.

A few terms used throughout:

  • Executor: the code that turns a chosen move into button presses.
  • Oracle: the program that supplies the correct move.
  • Tracker: the running summary of what has happened in the game.
  • Phase: one numbered task in the project’s sequence, such as Phase 4b.

In one line: a safety timeout added in the wrong layer turned out to refuse the expert’s own correct route. It took four phases to move it where it belonged.

What happened. When the game hands over an item, it runs a multi-step “receipt” sequence. To stop the executor from waiting forever, a clamp (a timeout) refused to continue once a receipt step had lasted 600 frames. A frame is one screen update, about a sixtieth of a second, so 600 frames is about ten seconds.

  • Phase 3b-2 found the counter was wrong. It started when the executor was created, not when the wait began, and it counted per step instead of per wait.
  • Phase 3b-3 stopped (a STOP under the rules). The fix it was asked for needed a tracker value that didn’t exist. The existing receipt age kept counting across steps, which contradicted the design’s “reset each stage.”
  • Phase 3b-4 fixed the tracker with a two-line change and re-checked every baseline. Then it found the real problem: four receipt steps in the teacher’s own route last longer than 600 frames. The longest S3 one was 1,725 frames. Following the expert exactly would have been refused 600 frames after entering the room.
  • Phase 3b-5 stopped again. The newly decided rule, “consecutive not-ready frames within a step,” couldn’t be computed from existing values. Its audit also showed the teacher’s route would never trip that rule.
  • Phase 3b-6 removed the clamp from the executor entirely. Receipt timeouts became the oracle’s job, as per-step caps in its stage tables.

How it was detected. Re-running the teacher’s full recorded route through the new code, the full-log shadow audit. Real data exposed what the spec missed.

The fix. No receipt logic remains in the executor. Afterward, the audit showed zero receipt-related refusals.

The lesson. The records don’t state a one-line lesson. The arc suggests one: test a safety limit against known-good production traffic before trusting it, and put a limit in the layer that understands the situation.


2. A replay that would have passed for the wrong reason

Section titled “2. A replay that would have passed for the wrong reason”

In one line: a check proved the right code was loaded, but the code itself read files from the wrong place. A second check caught it.

What happened. In Phase 3b, the task was to re-run an old campaign from a clean exported copy of the old code, not from the live working folder. The import proof confirmed that all 11 tool modules loaded from the export, and none from the live repository.

How it was detected. A different check, the source membership check, refused before any worker started. The task report then explained why the import proof’s pass meant nothing. The old code had the live repository’s path written into it. So even code loaded from the export would have checked and used the live folder’s files. In the report’s words, a replay from an export was “indistinguishable from a replay at HEAD.”

The fix. None in that phase; it listed three options and stopped. A later task (campaign-v3b) built replays with an explicit root: the worker is pointed at the export. It also added an open guard, which fails a shard if it opens any file in the live repository. In that later work, deliberately broken export copies were all caught.

The lesson. Proving where code came from is not the same as proving what it read. Defense in depth applies to evidence too.


In one line: a task prompt included a design call that contradicted the written design. The agent stopped and said so, and the call was withdrawn.

What happened. In Phase 3b-2, the prompt’s decision “D2” would have allowed medium-length “hold still” moves outside of normal play, such as during screen transitions. The prompt itself said to stop if the design documents restricted those moves to normal play.

How it was detected. The agent checked the committed action-schema spec. It says medium moves are eligible only when the game is world-ready (normal play), and that waits during transitions, receipts and menus use the single-frame wait. Applying D2 would also have made the committed spec’s text wrong. The agent noted honestly that no document used the prompt’s exact wording, and asked for a ruling.

The fix. D2 was withdrawn. The commit message records it as “design restricts neutral holds to WR; D2 withdrawn.”

The lesson. The written design outranks a convenient instruction, and the rulebook backs that up: if a prompt conflicts with the rules or the design, stop and ask. The records do not say who proposed D2.


In one line: a commit list left out three changed files. (A commit saves a set of changes into the project’s history. The commit list is the agent’s list of exactly which files to include.) A two-way check against git status now makes that impossible to miss.

What happened. In Phase 4b, the agent’s first COMMIT_PATHS.txt omitted three files it had legitimately changed under the repair rule: the oracle’s decision code and two route tables. The protected-file check had already classified them correctly. They were just missing from a hand-maintained list in the finishing script.

How it was detected. The owner, reviewing the first report. It was raised as a follow-up before commit.

The fix. The files were added. The finishing script also gained a completeness check: every path git reports as changed must be on the list, every listed path must exist, and every listed path must actually be changed. Exceptions are listed, never silently dropped. The final result was 185 paths on each side, PASS. Every later task runs this check.

The lesson. Don’t rely on a hand-kept list when the system can tell you the truth. Compare against git status, in both directions.


In one line: the oracle’s first real run of S3 pressed against a cave door for 1,500 decisions. The cause was a 4-pixel alignment tolerance.

6× speedVerified replay: game state hash-checked every frame

Iteration 0, the run this postmortem is about. The oracle's first S3 attempt reached the cave door and then kept stepping back and forth beneath it until its 1,600-decision budget ran out, shown at 6× speed.
Provenance
Item
iter0-stuck: War story: stuck at the cave door
Source run
Phase 4b oracle drive, iteration 0 (4a tables), repeat-1: FAILURE, DECISION_BUDGET_EXHAUSTED at f15668
Range
frames 13059–15668 (2,610 frames), drive ticks 143–2752
Playback
6×: one of every 6 emulator frames, played at 59.7275 frames per second; 7.283 s
Verification
Re-rendered from one boot-state load by input replay only; all 15,662 emulator frames from frame 7 to frame 15668 matched the recorded run's per-frame digests. Start-up loads: 1. Other saved-state loads: 0. Memory writes: 0.
Verified range digest
c9a79030609dbba9d0823170246bc43da5ac2951c461473830db28d855cdd646
Files (SHA-256)
  • iter0-stuck.webm, 191,250 bytes: 3f6080b73de63317645a6d4b5910a3bfded2d735b474e1fb4acc07ab6805623e
  • iter0-stuck-poster.png, 9,016 bytes: 0922adaad4a79e49951c867c353a406868fa9ad3b4fa808aa545453194afeca0
  • iter0-stuck.gif, 272,575 bytes: 8b1cd1e85c5c49fe8b001c6457f78ec998d48a15e62d3c58a7fd129173d4d246

How the footage is made

What happened. In Phase 4b, iteration 0, the oracle reached the cave door at decision 88. Then it made no progress for 1,512 decisions, until the 1,600-decision budget ran out.

How it was detected. The budget failure, followed by a decision-by-decision analysis. The mechanism ran in two stages:

  1. At 4 pixels off the door’s center (inside the allowed 4-pixel tolerance), pressing up was blocked by the wall. Recovery never kicked in, because recovery only turned on when the character was not at the door position.
  2. Later, standing exactly in line, the character stepped into the door tile itself. The tracker then said “not at the door position.” The oracle walked 8 pixels back down to line up again, and the cycle repeated: 1,308 press decisions and 137 walking decisions.

The fix. The tolerance went from 4 pixels to 0. Doors now carry their exact center pixel, the door tile counts as a valid pressing position, and the oracle steers to the exact pixel. The executor, the finish tests and the budgets were not changed.

Result. Iteration 1 succeeded: S3 done in 1,043 decisions and 1,660 frames, with full health and zero refusals. It was byte-identical across two processes, and it replayed cleanly as a one-shard campaign.

The lesson. The records don’t state one. The data shows a classic edge case: a tolerance that seems harmless can create a state that no rule handles.


In one line: clean code exports don’t include large files git ignores, so a replay from an export needs them supplied separately, by copy and never by link.

What happened. Campaigns fingerprint a large frame log from the route-scoping work. That file is ignored by git, so git archive doesn’t include it in an export.

How it was detected. Before any code was written. The agent raised it as the first of four questions at the start of campaign-v3b. No incident happened.

The fix. The replay procedure gained a step: copy every untracked bound input into the export, with its fingerprint checked. If the step is skipped, the replay refuses with a clear “frozen file missing” message. A symbolic link instead of a copy trips the open guard. Both behaviors were tested: the copied frame log gave 33/33 matches, and the linked one failed before any frozen code ran.

The lesson. “Everything is in git” is never quite true. List what isn’t, and make the tooling refuse when it’s missing.


7. The S2 gate re-lock: 99% to 97%, commit-bound to content-bound

Section titled “7. The S2 gate re-lock: 99% to 97%, commit-bound to content-bound”

In one line: the first S2 gate lock had both the wrong bar and a brittle binding. It was replaced before it ever ran.

What happened. The first S2 lock set a 99% bar (292 of 294). According to the author’s notes, that bar came from an advisor error: S1 had actually used 97%. The lock also pinned exact repository commits. That would have broken on any later commit.

How it was detected. By review before the gate ran. The 99% lock was never run as a gate; only its rehearsal ran.

The fix.

  • By the owner’s decision, the S2 gate was re-locked at the standing 97% rule (at least 286 of 294).
  • The new lock is content-bound. The locked base commit must be an ancestor of the current one, any changed files must be inside the gate’s own folders or match their locked bytes, and the tree must be clean.
  • A new gate run --check-only mode runs every check without executing anything. Before the commit, it correctly refused, with “repository is dirty” as the only reason. After the commit it must pass before the real run.
  • The S2 gate then ran and passed: 294 of 294 primary.

The lesson. Choose the threshold from the record, not from memory. Bind a lock to what matters (file contents), not to something that changes for unrelated reasons (the latest commit ID). A later, unrelated bug in a diagnostic flag was corrected in a separate folder, and the sealed verdict stayed unchanged.


One more clip was planned for this page: a fall through crumbling floor (floor tiles that give way under the character). It is not shown yet. Part of that run was not recorded frame by frame, so it cannot be replayed and checked the way every clip on this site is. The clip needs a later re-render.

Reserved for a later war story.

Gameplay footage from The Legend of Zelda: Link’s Awakening DX, captured from the author’s own emulator runs for technical commentary. The game and its imagery are © Nintendo. This project is not affiliated with or endorsed by Nintendo. How the footage is made.