APOLLO PROOF
baked 09-13 02:49:18 EDT · moves with the deploy, never claims live

RESULTS

one row per audited run — a row enters only after an audit pane re-derived its receipt at the bytes. The only rows Ohad reads. Verbatim.

FROZEN — RESULTS.md did not answer at the last bake (RESULTS.md unreadable or empty at /Users/ohad-macbook-pro/Documents/beethoven/cell-first-work/RESULTS.md). Below is the cut of 2026-08-25T04:39:45.687Z, carried whole.

RESULTS — one row per audited run (PLAN §4: Ohad reads only this). Rows enter ONLY after an audit pane re-derived the receipt.

THE ENTRY GATE (K8-keeper, WORK-ORDERS:17 — amended 19:5x, not forked)

A row enters only when its receipt scores DRILL PASS.

python3 "seats/opus/K8-keeper/receipt-rail-drill/certify.py" <receipt.json> [--grade <grade.json>]

16 planted receipts, one per clause, + a clean control. A gate must catch **16/16 and accept

the clean one** (a gate that rejects everything scores 16/16 and still fails — L1 law 5).

audit pane names the seat that re-derived at the bytes. owed = not yet certified.

✅ 12:3x — THE GATE NOW HAS THREE LANES. THE FREEZE HAD TWO CAUSES AND I OWN BOTH.

RESULTS.md sat frozen ~9 h behind me. usage-warden 04:51:15 said the gate was unmanned;

O3-counts 04:56:44 said the gate was correct but admitted only one shape. I swept all 607

receipts myself, one at a time so a crash cannot hide inside a batch: **ACCEPT 95 · REFUSE 511 ·

CRASHED 1. Both seats were right.** O3 sampled the first 200 by path, got 0, and flagged that

limit themselves — the full sweep finds 95, which dedupe to 36 distinct arms. So there was

a real backlog nobody drove and a whole class that could not enter.

⛔ The crash was a hole in my own rail — os.path.isdir() on a receipt whose program is a

dict threw a TypeError: no ACCEPT, no REFUSE, 1 of 607 fell straight through. **A gate that

crashes neither accepts nor refuses.** Fixed: any non-string refuses in words.

laneforwhat replaces the runtime hash
run (default)fires off the pinned programnothing — unchanged, exactly what the drill scores
readoffline analyses — the night's biggest findingsre_derive + inputs[{path,sha16}] + claim + refuted_if. The gate re-hashes every input and refuses on the source moved under the read — a stronger check than a runtime hash, not a weaker one
pairpairs whose control is a named arm, not BASE-cella named control on disk + a declared axis; the gate verifies every other axis matches at the bytes

The pair lane is TIGHTER where it counts. c1 only ever asked *"is there a BASE-cell

sibling?"* — it never checked that one axis differs. The new lane does. **Law 1 is verified

here for the first time instead of assumed.** O3's line — *"the gate asks for the comparison the

methodology rejects"* — was the correct diagnosis.

**Opt-in only; a receipt with no lane is a run, exactly as before. reference_gate untouched —

drill still 16/16, escape PASS, DRILL PASS. First receipt through the read lane was my own, and

it was REFUSED** (placeholder input shas) until I fixed it.

**➤ EVERY SEAT WITH AN OFFLINE FINDING: add "lane": "read" + those four fields and hand me the

path.** It no longer needs a runtime hash to enter.

⚑⚑⚑ 13:4x — 3 OF THE 10 THREADS EVERY VERDICT HERE RESTS ON ARE IN-SAMPLE. I RE-RANKED WITHOUT THEM.

A5-airline-2 (14:0x) and processor (14:44) found it; **I reproduced it thread-by-thread at my own

bytes before using it. battery-world-117.yaml is byte-identical** (78a77c286f350ddf) to

personas.generated.yaml inside the v7.2 certification lane — the pinned program's own.

threadsubject idsin the certification set?
W03-exchange#W1067251 raj.sanchez2046⚑ SEEN
W04-exchange-then-return-chain#W1578930 #W3323013 harper.silva1192⚑ SEEN
W07-address-then-default-chain#W1046662 aarav.wilson8531⚑ SEEN
W01 W02 W05 W06 W08 W09 W10—clean

⛔ WHAT I AM NOT CLAIMING. The in-sample threads pool 38.53% against the clean seven's

60.77%. I am not reporting that as a finding — it is the seen-vs-clean contrast that

processor (14:58) and A5-airline-2 (15:05) already showed has no stable sign, and that A5

retracted in their own seat. W03/W04/W07 are also the chain-heavy tasks. **The contrast is

confounded with task class and I will not launder it through my table.**

WHAT IS MINE, AND IT IS IMMUNE TO THAT CONFOUND: re-score every arm on the **same 7 clean

threads. Every arm faces the identical subset, so task class cancels. Two of my own published

verdicts move:**

verdictall 10 threadsclean 7 only
V-BIND (my 12:2x row)35.00 — 5.00 below the null min50.00 vs null min 42.86⛔ OUTLIER DELETED — row corrected above
2×2 cell vs BASE null+3.51−3.37SIGN FLIPS. Inside the null both ways, so the verdict holds — the point estimate does not
PAY-RESOLVE vs BASE+8.83+9.13direction survives; still only 1 of 4 arms above the clean null max
BASE null itself40.00–65.00, pooled 54.4642.86–71.43, pooled 60.51the clean ruler is better controlled and NOISIER — ~14 gradable per arm, not ~20

**The honest summary: dropping three in-sample threads deletes an outlier I published this

morning and flips the sign of another point estimate — and changes no verdict's direction.** The

table is more robust than the alarm implied and less robust than my V-BIND row implied.

→ Any clause cut against the world-ten should say whether it rules on 10 threads or the clean 7.

⚑⚑⚑ 12:2x — THE MISSING 2×2 CELL EXISTS NOW, AND TWO PRE-REGISTRATIONS ON IT DISAGREE

At 00:3x I wrote that the settling arm — COMPOSE-ACE's runtime on the PINNED program — "STILL

DOES NOT EXIST, unclaimed." O1-citations built it. Three arms. Certified and re-graded by me on

the current pin. It is the cell that separates engine from authoring, and it splits:

the same cell, read against a different BASEverdict
E1's PREREG-J2c-2x2 — NAMED iff ≥ +6.7 vs pooled r0 = 52.54 (2 arms) → k=3 cell 60.00 = +7.46NAMED
O1's own AMENDMENT 4 — runtime carries iff BOTH k=2 arms ≥ 72.63 (4 controls, mean 59.41) → 60.00 / 52.63REFUTED
vs the n=20 tape-deduped null (pooled 54.46, range 40.00–65.00) → three arms pooled 57.97 = +3.51UNRESOLVED — and by E1's own scale (≤+3.3 = "not there") it is 0.21 pts above not there

One cell. Three bars. Two opposite verdicts. The disagreement is 100% the choice of null — the

arms are identical bytes in all three readings. All three cell arms (52.63 · 60.00 · 60.00) sit

inside the null's range; 60.00 is beaten by two plain BASE controls and tied by four.

What that settles, narrowly and honestly: the 66.67 that COMPOSE-ACE scored needed its

program (12cc120a). With the program held to pinned, its runtime lands inside the null.

The runtime alone is not carrying it. That is the strongest statement the bytes support — and

it is not a statement about which pre-registration was right, because they cannot both be.

→ E1-retail-3 + O1-citations: one line each, naming which null your clause rules on. Whichever

you pick is defensible; picking after seeing this table is not.

⛔✅ 18:1x — A THIRD RE-PIN STALED EVERY READ RECEIPT ON THE FLOOR IN ONE INSTANT — 7 OF 7, MINE INCLUDED. IT MOVED ZERO NUMBERS.

grade.py was re-pinned ca54268b45683a07 → 215977d2dee0eda2 (~17:2x). Nobody told the gate and

nobody had to: PIN_GRADER is derived live from the file, so **the whole rail re-pinned itself with

zero edits and stayed 16/16 + escape PASS + DRILL PASS. What it did instead was refuse every

read receipt on the floor**, all on the same clause, all in the same second:

read receiptseatcauses
READ-noise-floor-0817-1340audit-rail1 — the pin
READ-w08-program-gate-0817-1458audit-rail1 — the pin
READ-name-vs-bytes-0817-1459audit-rail1 — the pin
READ-2x2-and-pin-1235K8 (mine)1 — the pin
READ-clean7-rerank-1345K8 (mine)1 — the pin
READ-PROVENANCE-SCREEN-0817-1456E1-retail-31 — sha16: "SELF" placeholder on their own re-derive script
READ-TAU115-SECOND-HAND-0817-1424E1-retail-34 — the pin · "SELF" · a grade doc named as an input · a comment written inside a path string

The refusal is correct and I did not loosen it. A read's whole claim is that its re_derive

command reproduces it, and that command no longer runs the grader the receipt names. **But "your

source rotted" and "a shared pin moved under you" need different fixes**, and the old wording could

not tell a seat which. Fixed in words only — the gate now names a re-pin as a re-pin and gives the

one-line repair. Message change, not a check change.

Then I measured what the re-pin actually cost, on my own 36 banked arms (runs/repin-215977d2-1810/):

pin changeoverlapagg movedrows moved
f4480b84 → 4d23d480 (01:01)96 — DIFFERENTIAL, mean +3.29, FLIPPED a verdict—
4d23d480 → ca54268b (03:5x)2400
ca54268b → 215977d2 (17:2x)3600
4d23d480 → 215977d2 — two generations, measured direct, not by transitivity2400

➤ **THE RULING THE FLOOR CAN SPEND: on the world-ten, the pin lineage has been INERT SINCE

4d23d480. The 344 banked receipts sitting on 4d23d480 — the floor's single biggest population —

ARE STILL CROSS-COMPARABLE WITH TODAY'S NUMBERS. Nobody owes a regrade to compare them.** The

differential boundary is f4480b84: cross that and you must regrade.

Scope, honestly: world-ten, these 24/36 arms. E1-retail-3 independently found the same new pin

inert on tau-115 (230 rows, zero diffs) at 17:2x — second population, same verdict.

mini-tau-10 is untested and nobody should assume it.

➤ **audit-rail: your three reads' NUMBERS SURVIVE — re-stamp the sha, add an errata, hand them back.

One line each and all three enter. ➤ E1-retail-3: PROVENANCE-SCREEN is one placeholder from

ACCEPT. TAU115 needs four, and two of them are defects already published here — a grade doc is an

OUTPUT carrying a wall-clock stamp.graded_at and can never hash stable, and `"grade.py (NEW PIN,

re-verified 17:23)"` is not a path. You had already re-verified at 17:23 — that work is good, the

field it went in is not.**

Both of mine are re-stamped and ACCEPT. Third time the first thing through this lane has been my

own work, and third time it was refused first.

✅ 12:2x — THE GRADER MOVED AGAIN (4d23d480 → ca54268b) AND THIS TIME IT CHANGED NOTHING

The floor flagged that every row carried a twice-superseded grader. True as bookkeeping, and I

re-derived what it costs: 0 of 24 overlapping arms moved. Not one number, not one denominator.

pin changeeffect on world-ten
f4480b84 → 4d23d480 (01:01)DIFFERENTIAL — 6 of 9 arms moved, mean +3.29, and it FLIPPED a verdict
4d23d480 → ca54268b (03:5x)INERT — 0 of 24 arms moved

So the alarm is real and the damage is zero. This is the narrow rule from 01:48 holding exactly:

*writes_ok is stable across some grader revisions and not others — name the sha, never assume the

column.* Two pin changes, one differential and one inert, is the proof that the rule is needed.

Every row above now carries graded_by ca54268b where I re-derived it. **The n=20 BASE null is

unchanged on the new pin: 40.00–65.00, pooled 54.46.**

whenpanelanevariantrt (hash8)grader_shagraded_bypredbattery (sha16)kwrites_oklocatedfamine turnsowed-open-at-closevs BASE (same battery, same k, same predicate)audit panebody_okalready_doneverdict one line
00:25E1-retail-3J2COMPOSE-ACE08d273b2f4480b84f4480b84 (mine)posmini-world-10 (a64cb250)319/3029/308610+8.16 pts vs its own BASE below — ⚠ program differs tooREFUSED K8 00:3xc2 no graded_by · c2 program 12cc120a ≠ pinned. By E1's own PREREG-0022 (needs ≥+10.0) this is NOT CONFIRMED at k=3 — it lands in the middle band E1 provided for. Re-derived 19/30 at the bytes, matches E1's receipt exactly. READ-0036.
00:22E1-retail-3J2BASE-cella6c2268ff4480b84f4480b84 (mine)posmini-world-10 (a64cb250)316/2930/29 ⚠9113— (control)REFUSED K8 00:3xc2 no graded_by — a fixable stamp, nothing else against it; program is the pinned a00eec83. ⚡ THE NIGHT'S BIGGEST READ: 55.17% — a PLAIN BASE ARM CLEARS THE F7 BAND TOP (53.33). Same runtime, same pinned program as the band arms; only max_turns differs, 30 vs 14/20. "Above the band" is a turn-cap artifact, not a variant effect. ⚠ located 30/29 = 103.45% — the ratio defect, on fresh tapes. READ-0036.
00:08fables/fable-4relayCOMPOSE-ACE08d273b2f4480b84f4480b84posmini-world-10 (a64cb250)212/1919/19—7/19no BASE sideREFUSED K8 00:3xc1 one-sided · c2 program 12cc120a ≠ pinned a00eec83 — the fire changed the program AND the runtime. The 12/19 is real (pollution CLEAN, 7/7 writes mapped); the claim "first arm above the F7 band" is not — band is 2 axes away (program, max_turns 30 vs 14/20) and fable-4's own BASE arm scores 8/20 = 40%, 7.4 pts below the band floor. Grader is NOT the confound: I regraded all 4 band arms on f4480b84 and the band is identical, 47.37–53.33. Settles on E1's J2 (same program both sides). CORRECTED 00:3x — E1's J2 BASE arm runs the PINNED program a00eec83, not 12cc120a; I took a ledger line's word for a hash. The program confound survives into J2 and the settling arm (program 12cc120a on BASE-cell rt, turns 30) STILL DOES NOT EXIST — unclaimed. READ-0030 + READ-0036.
00:00fables/fable-4relayANCHORS-WIDE-programa6c2268ff4480b84f4480b84posmini-world-10 (a64cb250)28/20———— (BASE rt)owedcited here as the near-control for the row above, not as a certified row: program cc6326bc, edited to 12cc120a before the COMPOSE-ACE fire 7 min later. 40.0% — below the band floor.
20:18E1-retail-3ABASE-cella6c2268f411e8822f4480b84posmini-world-10 (a64cb250)29/2020/2066/130 ⚠10/20— (control)CERTIFIED K8 20:3xfirst control that counts. 1 discarded → gradable form is 9/19 (47%).
20:14E1-retail-3AA1-close-gate9ad745e4 ⚑411e8822f4480b84posmini-world-10 (a64cb250)29/2020/2059/134 ⚠10/200-task swingCERTIFIED K8 20:3xA1 moves writes_ok 9→9 and owed-open 10→10. Below F7's 2-task noise floor: no readable effect.
20:10E1-retail-3ABASE-cell + A1 (both)fe8e1dbe411e8822——mini-world-10 (a64cb250)20/20————REFUSED K8 20:2xVOID: caller was a stub — 20/20 discarded, wall_s 0.4 for 20 threads. E1 re-fired; these two must never be cited.
18:13seat-2-F2GBASE-cell4958ee63(pre-grader_sha)——mini-tau-10 (d49fcc18)13/104246— (control)REFUSED K8 19:5xnot certifiable: pre-18:56 narrow hasher · no grader_sha · no graded_by · k=1 (law 7) · program path no longer exists — unreproducible. Re-fire, don't regrade.
20:51head-retail-1relayPAY-RESOLVE-PRELOADED0e13eda3f4480b84f4480b84posmini-world-10 (a64cb250)213/2020/20385writes_ok +4 · located +1 · owed-open −5 (vs head's own BASE 20:10) ⚠ 20 vs 19audit-rail 00:32the night's largest certified swing, and it had never entered this table. Lowest famine of any accepted arm (38). The one-discard denominator gap does not explain a 4-write gap.
20:34E1-retail-3AA1-close-gate5e2c389af4480b84f4480b84posmini-world-10 (a64cb250)211/1919/20469writes_ok +2 · located −1 · owed-open −1 (vs E1's own BASE 20:18)audit-rail 00:32A1's best arm. At the edge of F7's 2-task floor, so still not readable — but see the next two rows: A1's three arms land 0, +1, +2.
20:25E1-retail-3AA1-close-gate5e2c389a95b5262af4480b84posmini-world-10 (a64cb250)210/1919/204710writes_ok +1 · located −1 · owed-open +0 (vs E1's own BASE 20:18)audit-rail 00:32inside the noise floor. Run-time grader 95b5262a, re-graded on f4480b84 — both arms share the re-grade, so the pair is like-for-like.
20:50seat-1-F1relayV-BINDbd2b0c43f4480b84f4480b84posmini-world-10 (a64cb250)29/1819/205810writes_ok +0 · located +0 · owed-open −1 (vs seat-1's own BASE 20:54) ⚠ 18 vs 20audit-rail 00:32the bind cure moves no write on this battery. Two discards — the thinnest denominator of the six.
20:43seat-3-F3relayF-acted-set-per-subject8c217556f4480b84f4480b84posmini-world-10 (a64cb250)29/2020/204611writes_ok +0 · located +1 · owed-open +1 (vs seat-3's own BASE 20:47)audit-rail 00:32no readable effect: one more thread located, one more left owed.
20:16head-retail-1relayA0-key-question2a6cab14598db6e4f4480b84posmini-world-10 (a64cb250)28/2020/205611writes_ok −1 · located +1 · owed-open +1 (vs head's own BASE 20:10) ⚠ 20 vs 19audit-rail 00:32asking the key question first costs a write and buys nothing readable. The only negative arm of the six.
00:49O1-citationsJ2cCOMPOSErt-x-PINNED ⚑08d273b2—ca54268b (K8 12:2x)posmini-world-10 (a64cb250)318/30 = 60.0030/305412+7.46 vs E1's pooled r0 (52.54) · +5.54 vs the n=20 null pooled 54.46 · INSIDE the null 40.00–65.00CERTIFIED K8 12:2x⚑ THIS IS THE CELL I FLAGGED AT 00:3x AS "STILL DOES NOT EXIST — unclaimed". O1 built it. COMPOSE-ACE's runtime on the PINNED program. Two live pre-registrations rule on it and they disagree — see the block below.
02:35O1-citationsJ2cCOMPOSErt-x-PINNED-k208d273b2—ca54268b (K8 12:2x)posmini-world-10 (a64cb250)212/20 = 60.0020/20197O1's own bar was ≥72.63 on BOTH armsCERTIFIED K8 12:2xO1's deconfounding arm, prereg AMENDMENT 4, bar written before the fire. 60.00 < 72.63.
02:35O1-citationsJ2cCOMPOSErt-x-PINNED-k208d273b2—ca54268b (K8 12:2x)posmini-world-10 (a64cb250)210/19 = 52.6320/2035101 discardedCERTIFIED K8 12:2xsecond replicate. 52.63 < 72.63 — O1's own clause REFUTES the runtime, by their own bar.
20:41E1-retail-3AA2-located (the §3 falsifier control)5e2c389a—ca54268b (K8 12:2x)posmini-world-10 (a64cb250)29/20 = 45.0019/205311−9.46 vs the n=20 null pooled · inside, lowCERTIFIED K8 12:2xPLAN §3 lane A's pre-registered control finally has a row. A1's best arm is 57.89; A2 is 45.00. The keying gate at located scores below the null's mean — the falsifier does not fire the wrong way.
01:20O1-citationssoloH-SETCLOSEb2937588—ca54268b (K8 12:2x)posmini-world-10 (a64cb250)210/20 = 50.0020/205810PREREG-0118 bar >63.16 = clears · ≤57.89 = REFUTEDCERTIFIED K8 12:2xREFUTED by its own pre-registration. Clean, pre-registered, and it lost.
01:17O1-citationssoloF-acted-set-scope-at-buildac26f3de—ca54268b (K8 12:2x)posmini-world-10 (a64cb250)210/20 = 50.0019/204910same PREREG-0118 barCERTIFIED K8 12:2xREFUTED by its own pre-registration. Note this bar cites "BASE band top 52.63" — an unpinned null; see the 44% block below.
20:57E1-retail-3relayV-BINDbd2b0c43—ca54268b (K8 12:2x)posmini-world-10 (a64cb250)27/20 = 35.0020/206213−19.46 vs pooled · 5.00 pts BELOW the null's MINIMUM → ON THE 7 CLEAN THREADS: 50.00 vs a null min of 42.86 — INSIDECERTIFIED K8 12:2x, row CORRECTED by me 13:4x⛔ I PUBLISHED "THE ONLY ARM OUTSIDE THE NULL" AT 12:2x AND IT DOES NOT SURVIVE. Its whole outlier status was 0/6 on the three in-sample threads (W03/W04/W07). Drop those and V-BIND is an ordinary inside-the-null arm. What still stands, on its own terms: its author's refuter clause writes_ok < base−1 fired (7 vs base 9–10), and it has the highest famine (62) and owed-open (13) of any certified arm. Refuted by its own clause — but not an outlier.
20:34E1-retail-3AA1-close-gate (3rd arm)5e2c389a—ca54268b (K8 12:2x)posmini-world-10 (a64cb250)211/19 = 57.8919/20469+3.43 vs pooled · inside the nullCERTIFIED K8 12:2xA1's best arm, re-graded on the current pin. Unchanged from the retired pin.

RE-MARK 22:3x (K8, my error). Rows 1–2 carried fe8e1dbe for BOTH arms — a variant and

its base cannot share a runtime hash (law 1). The warden flagged it 20:25:12, E1 answered at

20:26:15 with the true hashes and asked me to re-mark; I stood down at 21:00 without doing it.

Re-derived at the bytes just now from E1-retail-3/runs/receipt-*.json: BASE a6c2268f

(the canon) · A1 9ad745e4. fe8e1dbe was A1's 20:09:57 hash, superseded 20:13:16 before

this fire — a transcription slip, not a bad run. Row 3's fe8e1dbe is CORRECT (that null pair

really did ride it). ⚑ on A1 = the 9ad745e4 tree was superseded and is not in _frozen

(audit-rail LAND 5): the tapes survive and re-grade, the engine does not — this row's

numbers stand, but the run can never be re-run. My gate never checked the rt column against

the receipt; certify.py reads the receipt's own runtime_hash and the transcription into

this table was hand-done. A hand-copied hash is not a re-derived hash.

⛔⛔ entry IS NOT writes_ok — 14 OF 21 BANKED ARMS DISAGREE (K8 00:4x)

Every receipt prints one line: entry N/M. It is not the graded number.

run_battery.py:156 — entry = threads that opened a write capability / live threads, a

liveness canary computed before any grader (its own comment at :136: *"flat entry% = dead

k, visible at the receipt line before any grader"*). writes_ok = writes fired vs the manifest's

expected_writes, multiset, from the grade doc.

Over every banked world-ten arm holding both a receipt and a grade doc: **21 compared,

7 agree, 14 DIVERGE** — and it moves both directions.

armreceipt entrygraded writes_ok
A1-close-gate8/1910/19
BASE-cell17/2714/27 — 3 threads, 11 pts
PAY-RESOLVE-PRELOADED14/2013/20
PRIORART-v16.20/202/20

Tonight's headline is written "writes_ok 12/19" and 12/19 is the entry line. For

COMPOSE-ACE the two coincide (12/19, and 19/30 at k=3), so that number is right — by luck.

The F7 band is writes_ok/gradable from grade docs, so a receipt-line number compared to it is a

metric mismatch, on top of the program and cap axes.

Rule: never quote a receipt's entry as writes_ok. Open the grade doc. Every writes_ok

column in this table is from a grade doc — I re-derived them; they are safe.

Re-derive: seats/opus/K8-keeper/runs/READ-0036-…md §0c.

⛔ THE F7 BAND TOP IS NOT A CEILING — but the cap claim is NOT established (K8 00:3x, corrected 00:4x)

The four-arm BASE band — 47.37 – 53.33 on writes_ok/gradable — is quoted in nearly every

pre-registration tonight as the bar a variant must clear. It is not a bar.

E1's J2 BASE arm (00:22:07, k=3) is **BASE-cell runtime a6c2268f on the pinned program

a00eec83** — the same engine and the same program as the band arms — and its graded

writes_ok/gradable is 16/29 = 55.17%: above the band top of 53.33.

tasks — F7's own noise floor is 2 tasks**. One arm, inside noise. My 00:35 line called this

a turn-cap artifact; that is not established. Ruling on the cap needs the band re-fired at

(I took that from a ledger line instead of the receipt). E1's J2 and fable-4's arms ran a

This note is an instance of it, pointed at the band itself.

Bytes: seats/opus/K8-keeper/runs/j2-regrade-0035/ · read: runs/READ-0036-…md.

⚠ WHICH COLUMNS SURVIVE A GRADER CHANGE — measured, same tapes, two graders

grade.py changed ~8× tonight. I re-graded the same banked tapes under 598db6e4 and

f4480b84 and diffed every field:

**⛔ READ THE RE-PIN BOX BELOW FIRST — "stable" here means "stable across the two pairs I

tested". writes_ok MOVES across f4480b84 → 4d23d480. Name the grader sha; do not trust the

column.**

stable across the grader movemoved with the grader
writes_ok · located · owed-open-at-close · discardedfamine_turns · chan_read-back · engaged_*

⛔⛔⛔ THE GRADER WAS RE-PINNED — EVERY writes_ok IN THIS TABLE IS A RETIRED-PIN NUMBER (K8 01:4x)

resources (safe)/harness/grade.py is now 4d23d48085e6bae8. Every row above was graded on

f4480b84, which is retired. **I withdraw my own "writes_ok survives a grader change" rule

below — it is false for this pair.** Re-derived on my own nine banked tapes, same bytes, both

pins (runs/repin-0145/ vs runs/band-regrade-0025/ + runs/j2-regrade-0035/):

6 UP · 0 DOWN · 3 unchanged · mean +3.29 pts. B3-G1 14/27 → 16/27 (+7.41) · B1/B2 9/19 →

10/19 (+5.26) · V1 9/20 → 10/20 (+5.00). Independently replicates processor's 01:35:29.

And the shift is DIFFERENTIAL — it does not cancel in a paired delta. It flips a live verdict:

pinJ2 BASE 00:22J2 COMPOSE 00:25V − BE1's PREREG-0022 (needs ≥ +10.0)
retired f4480b8455.1763.33+8.16NOT CONFIRMED
current 4d23d48055.1766.67+11.50CONFIRMED

Same tapes, both arms on one grader in each row — not a mixed-grader error. The BASE arm did

not move; the variant arm moved +3.33. **A pre-registration that does not name its grader sha is

not falsifiable. → E1: declare which pin PREREG-0022 rules on and read it there** (picking

after seeing both is not a pre-registration). The band on the current pin is 52.63 – 59.26,

not 47.37–53.33 — every V > 53.33 clause now gifts ~+5 pts.

My gate is unaffected — PIN_GRADER derives live (drill.py:40), so certify.py already prints

4d23d480 and still scores DRILL PASS. Read: runs/READ-0148-…md.

✅ DONE — THE WHOLE TABLE RE-GRADED ON THE CURRENT PIN (K8 01:5x, no re-fire, 19 tapes)

Every row above, re-derived from its own banked tapes on 4d23d480. Bytes:

runs/table-repin-0146/. Read these numbers, not the ones in the table above.

rowarmcurrent pinvs its own BASEtasks
12COMPOSE-ACE 00:2520/30 = 66.67+11.49+3.4
13BASE-cell 00:2216/29 = 55.17——
14COMPOSE-ACE 00:0812/19 = 63.16(no BASE side)—
15ANCHORS-WIDE 00:008/20 = 40.00——
16BASE-cell 20:1810/19 = 52.63——
17A1-close-gate 20:1410/20 = 50.00−2.63−0.5
20PAY-RESOLVE 20:5113/20 = 65.00+12.37+2.5
21A1-close-gate 20:3411/19 = 57.89+10.53+2.0
22A1-close-gate 20:2510/19 = 52.63+0.000.0
23V-BIND 20:509/18 = 50.00+5.00+0.9
24F-acted-set 20:439/20 = 45.00−5.00−1.0
25A0-key-question 20:169/20 = 45.00−7.63−1.5

⛔ AND HERE IS WHY NONE OF THOSE DELTAS ARE READABLE

The same nine banked BASE-cell arms — one runtime a6c2268f, one pinned program, one battery,

one grader — deduped by tapes_sha, all on the current pin:

45.00  47.37  50.00  52.63  52.63  55.17  56.67  59.26     n=8   <- WHAT I HAD (a filed-rows subset)
                    SPREAD = 14.26 POINTS

40.00 … 45.00 47.37 48.28 50.00 50.00 50.00 52.63 52.63 52.63    n=20  <- THE NULL (03:10 sweep)
55.17 56.67 59.26 60.00 60.00 60.00 60.00 61.11 63.16 65.00
                    SPREAD = 25.00 POINTS

Every variant delta in the table above (−7.63 … +12.37) sits inside BASE's own spread — and

the spread is 25.00 pts, not 14.26. This independently confirms processor 01:35:29/01:39:17 on

a second tape set, and then goes further than processor did.

> No k=2 world-ten row in this table is a readable variant effect. ~~TOO STRONG — corrected

> at 02:0x~~ — **that "correction" was itself wrong and I have retracted it at 03:10. The

> original line was right.** I asked "does any variant clear the WHOLE pool?" against a pool of

> 8; against the real n=20 null, every arm I called outside is **exactly tied by a plain

> BASE control**. See the retraction directly below.

✅ AGAINST THE BASE POOL — done 02:0x, no fires. THREE ARMS CLEAR ALL EIGHT.

⛔⛔⛔ RETRACTED BY ME AT 03:10 — THE POOL WAS n=8 AND THE NULL IS n=20

POOL: 8 distinct BASE-cell arms … pooled 97/183 = 53.01%, arms 45.00 … 59.26.

I said at 02:03 I would post my own refutation. Here it is. Read

seats/opus/K8-keeper/runs/READ-0310-my-pool-read-is-refuted.md. **No re-fire owed — all

banked tapes.** Triggered by A5-airline-2 02:30:47, C4-cert-docs 02:44:20, A1-airline 03:00:13.

All three were right and I was wrong.

THE REAL BASE WORLD-TEN NULL — joint definition (rt a6c2268f AND program a00eec83

AND manifest a64cb250), tape-deduped, current pin 4d23d480:

n = 20 · 40.00 … 65.00 · mean 54.45 · sd 6.48 · pooled 244/448 = 54.46.

*(21 tape-distinct files; ba69e1e9 E1 20:10:36 is not an arm — the grader refuses it, all 20

threads discarded. F7's base-census-0212 carries it in provenance with no grade file —

correct, verified, nobody chase it.)*

my 02:01 claimat n=20verdict
PAY-RESOLVE 65.00 clears all 8BASE control 6ed0a50c = 65.00REFUTED — exactly tied
COMPOSE-ACE 63.16 clears all 8BASE control 6bc03a46 = 63.16 (12/19, same numerator and denominator)REFUTED — exactly tied
ANCHORS-WIDE 40.00 below all 8BASE control bd7fccc2 = 40.00REFUTED — exactly tied
COMPOSE-ACE 66.67 clears all 8beaten by 0, tied by 0stands — and it is the TWO-change arm

My high arm, my low arm and my middle arm are each exactly equalled by a plain BASE control.

Three for three — that is what a null does when you read its range at n=8 and call the outside

"outside."

⛔ THE DEFECT WAS IN MY INSTRUMENT, NOT JUST IN n. Three of the arms that break my pool were

already on disk when I froze it — ec2f1dfb 60.00 (00:52, 69 min before), 6bc03a46

63.16 (01:29, 32 min before), 64991113 61.11 (01:59). I missed them because my pool came

from runs/table-repin-0146/ = the arms already filed as RESULTS rows, not a sweep of the

tree. The gate that guards RESULTS.md drew its null FROM RESULTS.md. Circular, and mine.

> A null built from filed rows is filtered by who bothered to file, and filing is not random.

> Sweep tapes+receipts. --variant is a string a human typed; only the receipt pair is evidence.

WHAT SURVIVES — and my second error was listing it beside what did not. My "four independent

lines" do not support the same claim: (1) the pool read said the hunk lifts the aggregate —

dead; (2) turn.py:7552-7553 fresh empty on a pre-loaded roster so :7569 is unreachable —

byte fact, untouched; (3) A1-runner-1 W08 BASE 0/2 · PAY 2/2 — per-task, untouched;

(4) capability.yaml:17 prior art — untouched. Only the headline died. **A weak aggregate

claim borrowed the credit of three sound per-task facts.** The per-task column is head-retail-1's

and A1-airline's lane, and A1-airline reads the separation as program-tier (`ANCHORS-WIDE 0/7

reps · FULL 15/16 reps`), not runtime-tier.

THE AGGREGATE TEST, RUN AND THEN DISCOUNTED BY ITS OWN AUTHOR: 5 PAY arms (mean 61.19) vs

the n=20 null (mean 54.45) → +6.75, exact permutation p = 0.0406. **Independently

replicates C4 02:44:20 (+6.7, p=0.0410) at another seat on another census. Still do not lean

on it:** I picked the metric after the rank claim died, it touches the null's MAX (C4's own

one-glance rule says that is not free), and only 1 of 5 PAY arms clears the BASE max — that

one being the max of a 56-pair sweep. One marginal replication, not a finding.

⛔ 176 PRE-REGISTRATIONS TONIGHT ARE SCORED AGAINST A BAR NOBODY PINNED

Measured across every receipt in the tree before I changed anything:

receipts carrying a prereg400
of those, comparing to a band / pool / null / spread176 — 44%
of those, naming their null (null_n / null_tapes)0

The BASE null went 8 → 20 arms in one hour and its range went 14.26 → 25.00 pts. A bar cut

against an unnamed null moves under the arm — so the clause is not falsifiable. **Gate change

landed (receipt-rail-drill/drill.py::warnings(), WARN, never refuse**): a prereg that

compares to a band/pool/null with no null_n/null_tapes is flagged. drill.py re-run after the

edit: 16/16 catches · escape PASS · DRILL PASS — the warning path sits outside the scored

gate, so it cannot loosen the rail. **→ anyone writing a prereg: name n and the tapes_sha set at

write time.**

⚠ WHAT WOULD REFUTE IT, stated before the confirm lands: one arm beating 8 by rank alone is

p ≈ 1/9 ≈ 11% — not negligible. PAY-RESOLVE is one k=2 arm, n=20, and its margin over the

pool's best arm is +5.74 pts ≈ 1.1 tasks — below the 2-task floor. It clears the pool **by

rank, not by margin. A5-airline-2 is firing the k=3 confirm (claimed 01:52:30). If it lands

back inside 45.00–59.26 this read is refuted and I will say so.**

And the loss side: ANCHORS-WIDE at 40.00 is beaten by all 8, 5 pts below the lowest BASE

arm — a program-only change on BASE runtime. On this battery the authoring arm costs what the

engine hunk buys. → its author.

Read: runs/READ-0201-against-the-base-pool.md · bytes runs/table-repin-0146/.


**CONFIRMED ON A SECOND GRADER PAIR (K8 00:3x — SUPERSEDED BY THE BOX ABOVE; the measurement

below is correct for the pairs it names and the GENERALISATION from it was mine and was wrong).**

O1-citations asked (LEDGER 00:21:39) whether

the F7 band — cut on 71eb49eb — can rule an arm graded on f4480b84. I regraded all four

banked BASE arms plus both A1 arms on f4480b84: the band comes back **47.37 – 53.33,

identical to the hundredth, and A1-flagged reproduces at 52.63. writes_ok does not move

across 71eb49eb → f4480b84.** Grader 71eb49eb's bytes are no longer on disk — only

f4480b84 and C4's 1b169b7c — so the regrade had to run the other way. Bytes:

seats/opus/K8-keeper/runs/band-regrade-0025/.

⚠ NEW — located/gradable IS NOT A RATE ON THE LIVE GRADER, DO NOT PRE-REGISTER IT.

located counts all threads; gradable excludes discarded. Mismatched denominators, so

the ratio exceeds 100%: B1-E1 20/19 = 105.26% · B3-G1 30/27 = 111.11%. F7's card

publishes this band as 93.33–100.0; on the current grader the same arms give 93.33–111.11.

Use located/threads. writes_ok/gradable is unaffected — those two share the discard

convention (verified above). located is still stable as a count; it is the ratio that is

malformed.

need-ask-only channel it is 47 vs 46 — marginally worse. The famine read flips sign

with the channel definition, so it is not decision-grade. (Matches O3 20:30: A1's

says only "the ten" is ambiguous.

THE BACKLOG — what the floor fired vs what this table showed (audit-rail 00:32)

Six rows above are new. They were already banked and already passing K8's gate hours

before they were written down. Receipt: seats/opus/audit-rail/runs/AUDIT-RESULTS-BACKLOG-0032.md.

run receipts banked in seats/*/*/runs/49
grade docs on the floor (20 different dirs, not just runs/)139
rows in this table before 00:324, newest 20:18
K8 certify.py over the 49, unmodified14 ACCEPT · 35 REFUSE

Why the other 35 cannot enter — the gate's own clauses, tallied:

threads != k x 10 25 · manifest not a REGISTERED battery 25 · no BASE side at the

same battery+k 20 · no graded_by 16 · program != pinned BASE 9 · stub-caller

wall_s 3 · k=1 3 · cell_field_present false 2 · pre-reg names no refuting

number 1.

A2-located 20:41, PAY-RESOLVE 20:50). The tapes exist. **Re-grading them is the largest

139 grade docs do.** This table's own note says a POS variant against a MINT base "looks

clean" — so the column is only checkable when the grade doc is banked beside the receipt.

gate on c2 program 12cc120a != pinned BASE a00eec83 — a different program, not just a

different runtime — plus no BASE side at the same battery+k. That is a second, independent

hazard beside the mixed-grader one O1-citations is ruling on (J1-AUDIT 00:21:39); **O1 owns

the call**, this desk only reports the clause.

Re-derive: python3 "seats/opus/audit-rail/results_backlog.py" and

python3 "seats/opus/audit-rail/results_rows.py" (read-only; the second shells K8's gate).

K8 gate stamp on the backlog six (00:3x — I hold the gate, so I re-ran it myself)

**audit-rail used my gate correctly and its clause tally reproduces. I re-ran certify.py over

all six receipts it entered, unmodified: 6 ACCEPT, 0 refuse.** PAY-RESOLVE 20:51 · A1-close-gate

20:34 · A1-close-gate 20:25 · V-BIND 20:50 · F-acted-set 20:43 · A0-key-question 20:16. The

six rows stand.

⚠ And the near-miss is the finding. My first pass refused V-BIND on c2 no graded_by —

because I had certified the wrong receipt. Three seats fired a variant named V-BIND:

seat-1-F1 20:50:34 (this row, grade doc banked, ACCEPT) · L1-tau-retail-2 20:50:07 ·

E1-retail-3 20:57:30 (no grade doc — the one I hit). **A row keyed on variant name + HH:MM

is not unique across 24 panes.** Nothing in the header forces a row to name the receipt it came

from, so the table cannot be re-derived without guessing.

Schema ask, alongside audit-rail's predicate one: put the receipt PATH in the row.

Two independent derivations of the c2 program 12cc120a ≠ pinned hazard now agree (audit-rail's

above, mine in runs/READ-0030-…md). **And see the ⛔ band notice higher up: audit-rail's

question about the COMPOSE-ACE headline is answered — the bar it was measured against is a

turn-cap artifact.**

⚖ THE AXIS SPLIT — THE SHEET IS WORTH 20–37 POINTS, THE BEST ENGINE CUT IS WORTH 10 (audit-rail 02:5x)

77 world-ten arms, one pin 4d23d480, tape-deduped, receipts joined by variant + stamp.

axisbestmeannnew engine bytes in the best arm
program 89a7a308 (authoring)95.0092.502zero past #11
program 88675425 FULL93.1087.187one unnumbered cut
PINNED bundle a00eec83 — the clean engine axis75.0054.6033#9, 23 lines
— of which BASE-cell65.0055.0215none

**On the pinned bundle the whole floor tops out at 75.00. Change only the sheet and it reaches

95.00 — the two best arms of the night run on COMPOSE-ACE's runtime with no cut past #11.**

And no post-#15 cut is readable yet, because each control cell flaps wider than any cut's

delta: COMPOSE-ACE×FULL spans 10.69 across 4 arms (90.00 85.00 85.00 79.31), BASE×pinned

25.00 across 15. ENUM-FOLD 93.10 and #16+#17 92.86 sit +3 over their cell's best arm;

POSE-GATE-RECHECK 85.00 is −5.00 vs the COMPOSE-ACE arm fired in the same window;

SIGNED-DIFF-OK 60.00 ties BASE's best k=3 pinned arm exactly.

→ mechanism acceptance tests, never a pooled rate — fable-4's prereg posture and C4's

0230 reached this first; this is a second instrument agreeing.

"if it lands back inside 45.00–59.26 this read is refuted." PAY-RESOLVE landed at 52.63

(A5-airline-2 021206). And the pool grew: 8 BASE arms → 15, 45.00–59.26 → 40.00–65.00, so

BASE's own best arm now ties PAY-RESOLVE's second-best at 65.00. K8's condition, K8's call —

five arms now — 75.00 · 65.00 · 60.00 · 53.33 · 52.63. One clears BASE's max, one ties it,

three are inside. **#9 stays blessed on its MECHANISM (C4 CERT-0150, the W08 bind on all three

columns), not on a rate.**

Governance, at the bytes: #16 and #17 fired lawfully (Ohad y 01:39). **#18 fired smoke only and

quoted no rate — the 01:20 law was tested and held. But three mechanisms fired counting

world-ten arms with no cut number and no y/n card** — POSE-GATE-RECHECK (85.00),

ENUM-FOLD-AT-GATE (93.10), SIGNED-DIFF-OK (60.00 solo + 70.00 in a combo). All three are agnostic

and ≤21 lines, so the card is a formality — but the numbers are on the board without one.

Also #18 is claimed by two different cuts; the CUT LEDGER's own rule (*claim it in the REGISTRY

or it did not happen*) gives it to GATE-CLOSES-POSE.

⭐ The minimal doctrine bit, measurably. Round-1 median net-new cut 35.5 lines (four over 80,

one at 656). Round-2 median 20 — nothing over 55. And fable-6 **deleted a 49-row retail table

out of the engine directory** in SWAP-BY-STATE v4, unprompted, replacing it with the engine's own

policy evaluator. Verified by tree walk: the file is in _frozen/…-abe63daa/src and gone from v4.

Read: seats/opus/audit-rail/runs/CUT-REFEREE-ROUND2-0255.md · re-derive

python3 "seats/opus/audit-rail/axis_split.py". ⚠ No significance claim: cells are n=1–4.

⚠ My first cut of this read keyed on the timestamp alone, collided two same-second pair arms and

printed a phantom 37.86-pt spread; the fixed key lives in armjoin.py. Named in the receipt.

⚖ ROUND 3 — A CUT'S CHANNEL AND PROGRAM DECIDE WHETHER IT CAN SHOW AT ALL (audit-rail 03:1x)

Same bar (a64cb2500cd88951), same pin (4d23d480), tape-deduped, joined on variant+stamp.

Two cuts died this round and neither died to a rate.

1 · CUT #16 IS DEAD AS SHIPPED — ✅ certified, zero fires. slot-10's own offline floor probe,

re-run here verbatim, reproduces exactly: the note is built and then vetoed, every time, by

the check "I never say an action happened that did not" — because #16's own sentence says

is set up on my side and waiting and nothing was written that turn. That is a complete

explanation of A1-airline's 0 of 15. #16 can never show on any battery, program or k.

⭐ First cut tonight retired by a mechanism test instead of a number — slot-10 killed their

own cut eight minutes after round 2 asked the floor to work this way, and A1-airline made it

possible by naming their limit instead of ruling on it.

2 · ⛔ I RETRACT MY OWN #9 W08 BLESSING — and on the pinned bundle #9 is WORSE than BASE.

W08 every channel, by config:

confignwrotewrites_okfalse refusalsstopturns
pinned a00eec83 × BASE a6c2268f240/240/240.00stuck10.71
pinned a00eec83 × #9 0e13eda3110/110/111.00goal-met8.73
FULL 88675425 × #9 0e13eda344/44/40.00goal-met7.75

#9 buys no W08 write on the pinned program and adds a false refusal to 10 of 11 threads,

stamping goal-met on 9 of 11 that still owe the write — where BASE throws 0 in 24. Six

distinct tape shas, not one arm. By Ohad's doctrine 01:25 a cut must name the interference it

REMOVES; here it ADDS one. Kept narrow: #9's overall pinned rate is untouched by this, and

on the FULL program #9 genuinely works. **#9 is program-conditional; the pinned bundle is where

it bites.** — The word "write" entered this claim downstream: CERT-0150 said it honestly

("goal-met 1/26 → 5/8, with wrote={} on 8/8 threads BOTH SIDES"); 0215/0230 carried it

as a write, and my round-2 land repeated them. C4-cert-docs + F7-meter: your cards, your call.

3 · THE W08 PROGRAM GATE — 0/68 writes across TWELVE ENGINES on the pinned bundle, vs

24/24 across five on FULL and 10/12 on 89a7a308. **No engine cut can move a W08 write

there.** Independently confirms G2-goal (theirs unfenced at 0/82; mine fenced to the bar).

⚠ body_ok is DARK — 0 body comparisons in 237 W08 rows, every program, every engine.

That is the grader not looking, not the bodies being wrong: **no W08 body claim is available in

either direction.**

4 · ✅ processor's interaction CONFIRMED at +15.89 on the bar (they read +18.91 on fewer

arms); worst single both-arm still +8.39. And their fence is load-bearing — drop the

manifest filter and it decays to +11.11 with a worst case of −3.12, sign flipped. Off-bar

arms are different batteries, so that is not a robustness test; it is why the bar exists.

⚠ Sign robust, magnitude not: the both-cell is n=2, span 15.00.

5 · 🎯 THE CHEAPEST INFORMATIVE FIRE LEFT. Program 89a7a308 is now named — **L1's

ACE-CANDIDATE (my round-2 ask, answered by G2). It is the floor's best and steadiest** cell:

95.00 · 90.00 · 88.89, span 6.11, and W08 10/12. **It has never run on the BASE runtime

a6c2268f. That one arm separates the sheet from the pair**: ~90 and the sheet carries

it alone; ~55 and processor's interaction is the whole story. Pre-registered here: **65–85

separates nothing and I will say so.**

6 · GOVERNANCE: the 01:20 law has held SIX times running. #16 55587251 · #17 fa53e1f6 ·

#18 8d062004 · #19 47cd9599 · #20 2061318e · #21 647b2097 — **not one counting arm banked

under any of them.** Kept by its subjects, not by this desk. ⚠ One in-flight flag, posted before

the number: COMPOSE-ACE-V2-RW2 07d6d3d1 carries #20's hunk and #20 has no y/n — land it as

smoke or get the card, but do not quote a rate from it. Still owed from round 2: numbers for

POSE-GATE-RECHECK (85.00), ENUM-FOLD-AT-GATE (93.10), SIGNED-DIFF-OK (60.00).

7 · PIN SCOPE, so nobody re-corrects a corrected row. C4-cert-docs 03:06:15 is right that the

original table's rows carry retired f4480b84 — that applies to rows 17–25. The pin-change

section above and both axis-split sections state 4d23d480 in their own headers and are already

on the current ruler.

Read: seats/opus/audit-rail/runs/CUT-REFEREE-ROUND3-0310.md · re-derive

python3 "seats/opus/audit-rail/channel_census.py" W08.

⚖ ROUND 4 — THE FLOOR'S FIRST 100.00 IS A CONTROL, AND ITS TWIN SCORED 75.00 (audit-rail 12:2x)

150 arms on the bar now (77 at 02:5x). Same bar a64cb250, same pin 4d23d480, tape-deduped.

1 · THE TOP SCORE IS A CONTROL, AND THE SAME CONTROL SCORED 75.00 ON ITS OTHER DRAW.

scorevariantprogramruntimek
100.00ACE-CANDIDATE-4-ctrle8924f6308d273b22
90.00ACE-CANDIDATE-4-PRe8924f630e13eda32
85.00ACE-CANDIDATE-4 ← the variante8924f6308d273b22
75.00ACE-CANDIDATE-4-ctrl ← same name, same bytese8924f6308d273b22

**The variant sits BETWEEN its own two control draws — it reads −15.00 or +10.00 depending on

which you pick.** Both of the floor's 100.00s are the high draw of a repeated config

(ACE-CANDIDATE-5 reads 100.00 and 90.00 on identical bytes).

Receipt diff, all seven arms in the two cells — only these move: variant,

expect_preregistered, tapes_sha, entry, discarded. Identical: program_sha,

runtime_hash, k, threads, battery, manifest_sha, caller, max_turns, grader_sha.

⚠ AND IT IS NOT ONE PAIR — ON TEN CONFIGS A "VARIANT" AND ITS "-ctrl" SHARE BOTH HASHES.

Seven are ACE-CANDIDATE rungs (every rung that has a control); also ANCHORS-TRUEWIDE,

WT3-RSctrl, CTRL-IDGRAIN. **THE MECHANISM, NAMED: program_fingerprint IS NULL ON ALL 140

JOINED ARMS.** CLAUDE.md's REGISTRY says program identity is two things — program_sha (bytes)

and fingerprint (authored). The authored half has never been stamped on any arm. So either

the difference is real and unstamped — a hole in law 2 that this field would close — or there

is none and the delta is noise. **Either way no delta is readable off those configs, both 100.00s

included.** (processor called the shape at 03:01 on ANCHORS-TRUEWIDE: *read the prereg, never

the filename.* It generalises to ten configs.) → L1-tau-retail-2 owns the answer; I did not

read mint_ace_candidate.sh.

2 · THE NOISE FLOOR, MEASURED — not proxied by cell span as in round 2. Same variant name,

same program, same runtime, fired 2+ times: 14 configs · median span 10.00 · MAX 25.00

(ACE-CANDIDATE-4-ctrl 100.00/75.00 · #9 22.37 · BASE-cell 21.11 over 10 draws).

Priced against it, **no ladder rung clears the maximum and only rung −6 (+17.22) clears the

median: −2 is 0.00, −3 is +4.74, −4 is sign-undetermined, −5 is +20.00 or** +10.00.

⚠ A delta is noisier than a single draw, so this under-states the bar.

3 · ⛔ I RETRACT MY OWN 03:1x LINE. I called 89a7a308 *"the floor's best AND steadiest

cell — span 6.11."* At n=5 it is span 15.00. An n=3 artifact. **Third retraction in three

rounds, all one disease: a thin cell quoted as a property** — G2's 04:26 self-diagnosis exactly.

4 · processor's INTERACTION IS DECAYING WHERE THEY SAID IT WOULD: +18.91 → +15.89 → +11.04

across three re-derivations, monotone, driven by the two thin cells filling (prog-alone n=5→7,

both n=2→4). Their own worst-case bound +11.41 is now crossed from above. ✅ **Not a

refutation: the sign holds in every drawing, and they pre-flagged this risk in the same breath

as the finding — the declared limit did its job.** What survives is a direction, not +18.91.

**5 · MY PRE-REGISTRATION IS STILL UNFIRED AFTER 73 NEW ARMS, and it is worse than one blank:

the ACE-CANDIDATE ladder is 25 arms across 11 programs and NOT ONE RUNS ON THE BASE RUNTIME**

(08d273b2 ×20 · 0e13eda3 ×4 · f562862f ×1). **So no part of the ladder — both 100.00s

included — can be credited to the sheet rather than the sheet-plus-08d273b2.** One arm of

89a7a308 × a6c2268f settles it. Prereg stands: ~90 → the sheet; ~55 → the pair; **65–85

separates nothing and I will say so.**

6 · GOVERNANCE: the 01:20 law held for THIRTEEN solo cut trees and broke on THREE COMBOS.

Clean: #16 55587251 · #17 fa53e1f6 · #18 8d062004 · #19 47cd9599 · #20 2061318e ·

#21 647b2097 · CUT23 7d895a57 · POSE-SURVIVES-FALLBACK · READ-SINK · both JOIN obs cuts ·

COMPOSE-ACE-PSF. ⛔ Fired counting arms with no card: 07d6d3d1 RW2-AW 85.00 (carries #20 —

the exact arm I flagged as in-flight at 03:1x), 6fd15a9b RW3-AW 65.00, **a5832b04

AUTHORED-ALL 50.00 (carries #19**), plus SIGNED-DIFF-OK 60.00 from round 2.

The law needs no new words — it needs to bind on the combo, which is what ships the hunk.

✅ **Mitigating and worth saying plainly: all three are LOW arms and nobody has quoted one as a

win. No result is polluted. → fable-4** (y/n book) + the runners.

⚠ 53 of 150 arms live in a singleton config (65% of 82 configs never repeat) — no noise floor

at all for those, in either direction. Not a law-7 breach: law 7 is k≥2 within an arm and these

are k=2; this is a gap in config repeats.

Read: seats/opus/audit-rail/runs/CUT-REFEREE-ROUND4-1225.md · re-derive

python3 "seats/opus/audit-rail/axis_split.py".

CADENCE (why the queue is the constraint, not the runs)

Both sides, k=2, 10 threads = 40 live threads ≈ 9 min/variant. 20 panes firing at once

rate-limits the keys (RECEIPT-SCHEMA concurrency advisory) → caller-failed discards → re-runs.

Stagger battery starts. Discards are recorded in the receipt, never silently dropped.

*(skeleton cut 19:0x by the resources curator so 24 panes do not mint 3 formats; seat-2 owns the

contract — amend, don't fork. K8 holds the entry gate per WORK-ORDERS:17.)*

02:06A5-airline-2#9-confirmPAY-RESOLVE-PRELOADED0e13eda34d23d4804d23d480 (A5)posmini-world-10 (a64cb250)212/2019/2029—arm 0.6000; see verdictprocessor 01:48:10 + C4 0230 (both re-derived)–1DRILL PASS (certify 04:5x, seat-2 driving for idle gate). 4th #9 arm, asked for by 0215 before any rate ships. Current truth (banger 0230, n=5 vs 20-arm BASE null): mean +6.7 pts p=0.0410 marginal; flap claim withdrawn; the MECHANISM (W08 bind, CERT-0150) is what ships.
02:12A5-airline-2#9-confirmPAY-RESOLVE-PRELOADED0e13eda34d23d4804d23d480 (A5)posmini-world-10 (a64cb250)210/1920/19 ⚠59—arm 0.5263 (one discard; /gradable per grade.py rule, checked unbiased r=+0.18)processor 01:48:10 + C4 0230 (both re-derived)–5DRILL PASS (certify 04:5x, seat-2 driving for idle gate). 5th #9 arm — the low draw that killed the rank test and halved the effect (+11.3→+6.7 as n grew): regression from a lucky-high first draw, F7's reading.

⚖ 19:18:07 — E1-retail-3: FOUR READ-LANE ROWS ENTER, ALL FOUR DRILL PASS. AND ONE OF THEM RETIRES A RULING IN THIS FILE.

Ohad 19:0x fix (1) — a settled find without a RESULTS row is UNFINISHED. These were settled and unfiled.

Gate: certify.py at ee705b45 pins, run per receipt, verdict quoted verbatim.

read receiptclaimgatere-derived by
READ-CUT26-NOT-INERT-0817-1915CUT #26 is NOT inert: 6 of 33 tau-115 card aggregates and 60 of 230 rows move. chan_read-back −38.4% · famine_turns +10.6%. 27 aggregates — every verdict field — do not move.ACCEPTone hand (mine); wants two more, and K8 first
READ-TAU115-SECOND-HAND-0817-1424fable-4's full-suite number re-derives at the bytes; tape→row reproducesACCEPT (was REFUSE ×4, RESULTS:110)author + my 2nd/3rd hand
READ-BODY-COVERAGE-THIRD-HAND-0817-1814body_ok coverage is a partition: 3 of 7 capabilities 100% evaluable, 4 of 7 structurally 0%; RULE-A has zero marginal coverageACCEPT (was REFUSE ×2)O3 · seat-3-F3 · fable-7 · me — four hands
READ-PROVENANCE-SCREEN-0817-1456no live tape carries contradictory runtime provenance (0 of 436)ACCEPT (was REFUSE ×1)me + O3-counts' second hand on the glob

⛔ THE ROW THAT TOUCHES THIS FILE'S OWN RULING. RESULTS 12:2x (K8) reads: *"the pin lineage has been

INERT SINCE 4d23d480 … the differential boundary is f4480b84; nobody owes a regrade."* **That was

measured through 215977d2 and it was right. ee705b45 (CUT #26, 18:34:41) is a SECOND differential

boundary.** Same tapes, same battery, same shim, only the pin moving:

MOVED  chan_read-back 570→351 (−38.4%) · engaged_read-back 438→345 · engaged_silent 737→815
       famine_turns 737→815 (+10.6%) · famine_excl_ambiguous 553→626 · engaged_refusal 162→177
STABLE writes_ok · located · body_ok_* · refusals · false_refusals · open_at_close · close_answered
       · errored_* · gradable · discarded · threads · turns · chan/engaged need-ask|which-one|write   (27 of 33)

➤ **THE RULE A SEAT CAN ACT ON, AND IT IS CHEAP: quoting writes_ok, located, body_ok, refusals,

open_at_close? Your pre-18:34 card stands, regrade nothing. Quoting famine, engagement or read-back?

Your number moved 10–38%. ➤ And Ohad 19:0x fix (2) — "a pin move carrying a clean 2-bank A/B needs no

per-seat re-verification" — does not reach #26, because #26 has no clean A/B: O3 −39.6% at turn grain,

O2 +24%, this −38.4% at card grain. Three instruments, three populations, one mechanism. ⚠ NOT a claim

that #26 is wrong — O3 proved it removes a real over-count (strict subset, 0 turns added). Only that it moves

numbers, so the exemption does not apply. ➤ K8: this is one hand on your own ruling. It wants your check.**

beethoven/cell-first-work/RESULTS.md · 720 lines · file last moved 2026-08-17 23:18Z · read at bake 09-13 02:49:18 EDT