RESULTS — one row per audited run (PLAN §4: Ohad reads only this). Rows enter ONLY after an audit pane re-derived the receipt.
THE ENTRY GATE (K8-keeper, WORK-ORDERS:17 — amended 19:5x, not forked)
A row enters only when its receipt scores DRILL PASS.
python3 "seats/opus/K8-keeper/receipt-rail-drill/certify.py" <receipt.json> [--grade <grade.json>]
16 planted receipts, one per clause, + a clean control. A gate must catch **16/16 and accept
the clean one** (a gate that rejects everything scores 16/16 and still fails — L1 law 5).
audit pane names the seat that re-derived at the bytes. owed = not yet certified.
✅ 12:3x — THE GATE NOW HAS THREE LANES. THE FREEZE HAD TWO CAUSES AND I OWN BOTH.
RESULTS.md sat frozen ~9 h behind me. usage-warden 04:51:15 said the gate was unmanned;
O3-counts 04:56:44 said the gate was correct but admitted only one shape. I swept all 607
receipts myself, one at a time so a crash cannot hide inside a batch: **ACCEPT 95 · REFUSE 511 ·
CRASHED 1. Both seats were right.** O3 sampled the first 200 by path, got 0, and flagged that
limit themselves — the full sweep finds 95, which dedupe to 36 distinct arms. So there was
a real backlog nobody drove and a whole class that could not enter.
⛔ The crash was a hole in my own rail — os.path.isdir() on a receipt whose program is a
dict threw a TypeError: no ACCEPT, no REFUSE, 1 of 607 fell straight through. **A gate that
crashes neither accepts nor refuses.** Fixed: any non-string refuses in words.
| lane | for | what replaces the runtime hash |
|---|---|---|
run (default) | fires off the pinned program | nothing — unchanged, exactly what the drill scores |
read | offline analyses — the night's biggest findings | re_derive + inputs[{path,sha16}] + claim + refuted_if. The gate re-hashes every input and refuses on the source moved under the read — a stronger check than a runtime hash, not a weaker one |
pair | pairs whose control is a named arm, not BASE-cell | a named control on disk + a declared axis; the gate verifies every other axis matches at the bytes |
The pair lane is TIGHTER where it counts. c1 only ever asked *"is there a BASE-cell
sibling?"* — it never checked that one axis differs. The new lane does. **Law 1 is verified
here for the first time instead of assumed.** O3's line — *"the gate asks for the comparison the
methodology rejects"* — was the correct diagnosis.
**Opt-in only; a receipt with no lane is a run, exactly as before. reference_gate untouched —
drill still 16/16, escape PASS, DRILL PASS. First receipt through the read lane was my own, and
it was REFUSED** (placeholder input shas) until I fixed it.
**➤ EVERY SEAT WITH AN OFFLINE FINDING: add "lane": "read" + those four fields and hand me the
path.** It no longer needs a runtime hash to enter.
⚑⚑⚑ 13:4x — 3 OF THE 10 THREADS EVERY VERDICT HERE RESTS ON ARE IN-SAMPLE. I RE-RANKED WITHOUT THEM.
A5-airline-2 (14:0x) and processor (14:44) found it; **I reproduced it thread-by-thread at my own
bytes before using it. battery-world-117.yaml is byte-identical** (78a77c286f350ddf) to
personas.generated.yaml inside the v7.2 certification lane — the pinned program's own.
| thread | subject ids | in the certification set? |
|---|---|---|
| W03-exchange | #W1067251 raj.sanchez2046 | ⚑ SEEN |
| W04-exchange-then-return-chain | #W1578930 #W3323013 harper.silva1192 | ⚑ SEEN |
| W07-address-then-default-chain | #W1046662 aarav.wilson8531 | ⚑ SEEN |
| W01 W02 W05 W06 W08 W09 W10 | — | clean |
⛔ WHAT I AM NOT CLAIMING. The in-sample threads pool 38.53% against the clean seven's
60.77%. I am not reporting that as a finding — it is the seen-vs-clean contrast that
processor (14:58) and A5-airline-2 (15:05) already showed has no stable sign, and that A5
retracted in their own seat. W03/W04/W07 are also the chain-heavy tasks. **The contrast is
confounded with task class and I will not launder it through my table.**
WHAT IS MINE, AND IT IS IMMUNE TO THAT CONFOUND: re-score every arm on the **same 7 clean
threads. Every arm faces the identical subset, so task class cancels. Two of my own published
verdicts move:**
| verdict | all 10 threads | clean 7 only | |
|---|---|---|---|
| V-BIND (my 12:2x row) | 35.00 — 5.00 below the null min | 50.00 vs null min 42.86 | ⛔ OUTLIER DELETED — row corrected above |
| 2×2 cell vs BASE null | +3.51 | −3.37 | SIGN FLIPS. Inside the null both ways, so the verdict holds — the point estimate does not |
| PAY-RESOLVE vs BASE | +8.83 | +9.13 | direction survives; still only 1 of 4 arms above the clean null max |
| BASE null itself | 40.00–65.00, pooled 54.46 | 42.86–71.43, pooled 60.51 | the clean ruler is better controlled and NOISIER — ~14 gradable per arm, not ~20 |
**The honest summary: dropping three in-sample threads deletes an outlier I published this
morning and flips the sign of another point estimate — and changes no verdict's direction.** The
table is more robust than the alarm implied and less robust than my V-BIND row implied.
→ Any clause cut against the world-ten should say whether it rules on 10 threads or the clean 7.
⚑⚑⚑ 12:2x — THE MISSING 2×2 CELL EXISTS NOW, AND TWO PRE-REGISTRATIONS ON IT DISAGREE
At 00:3x I wrote that the settling arm — COMPOSE-ACE's runtime on the PINNED program — "STILL
DOES NOT EXIST, unclaimed." O1-citations built it. Three arms. Certified and re-graded by me on
the current pin. It is the cell that separates engine from authoring, and it splits:
| the same cell, read against a different BASE | verdict |
|---|---|
| E1's PREREG-J2c-2x2 — NAMED iff ≥ +6.7 vs pooled r0 = 52.54 (2 arms) → k=3 cell 60.00 = +7.46 | NAMED |
| O1's own AMENDMENT 4 — runtime carries iff BOTH k=2 arms ≥ 72.63 (4 controls, mean 59.41) → 60.00 / 52.63 | REFUTED |
| vs the n=20 tape-deduped null (pooled 54.46, range 40.00–65.00) → three arms pooled 57.97 = +3.51 | UNRESOLVED — and by E1's own scale (≤+3.3 = "not there") it is 0.21 pts above not there |
One cell. Three bars. Two opposite verdicts. The disagreement is 100% the choice of null — the
arms are identical bytes in all three readings. All three cell arms (52.63 · 60.00 · 60.00) sit
inside the null's range; 60.00 is beaten by two plain BASE controls and tied by four.
What that settles, narrowly and honestly: the 66.67 that COMPOSE-ACE scored needed its
program (12cc120a). With the program held to pinned, its runtime lands inside the null.
The runtime alone is not carrying it. That is the strongest statement the bytes support — and
it is not a statement about which pre-registration was right, because they cannot both be.
→ E1-retail-3 + O1-citations: one line each, naming which null your clause rules on. Whichever
you pick is defensible; picking after seeing this table is not.
⛔✅ 18:1x — A THIRD RE-PIN STALED EVERY READ RECEIPT ON THE FLOOR IN ONE INSTANT — 7 OF 7, MINE INCLUDED. IT MOVED ZERO NUMBERS.
grade.py was re-pinned ca54268b45683a07 → 215977d2dee0eda2 (~17:2x). Nobody told the gate and
nobody had to: PIN_GRADER is derived live from the file, so **the whole rail re-pinned itself with
zero edits and stayed 16/16 + escape PASS + DRILL PASS. What it did instead was refuse every
read receipt on the floor**, all on the same clause, all in the same second:
| read receipt | seat | causes |
|---|---|---|
READ-noise-floor-0817-1340 | audit-rail | 1 — the pin |
READ-w08-program-gate-0817-1458 | audit-rail | 1 — the pin |
READ-name-vs-bytes-0817-1459 | audit-rail | 1 — the pin |
READ-2x2-and-pin-1235 | K8 (mine) | 1 — the pin |
READ-clean7-rerank-1345 | K8 (mine) | 1 — the pin |
READ-PROVENANCE-SCREEN-0817-1456 | E1-retail-3 | 1 — sha16: "SELF" placeholder on their own re-derive script |
READ-TAU115-SECOND-HAND-0817-1424 | E1-retail-3 | 4 — the pin · "SELF" · a grade doc named as an input · a comment written inside a path string |
The refusal is correct and I did not loosen it. A read's whole claim is that its re_derive
command reproduces it, and that command no longer runs the grader the receipt names. **But "your
source rotted" and "a shared pin moved under you" need different fixes**, and the old wording could
not tell a seat which. Fixed in words only — the gate now names a re-pin as a re-pin and gives the
one-line repair. Message change, not a check change.
Then I measured what the re-pin actually cost, on my own 36 banked arms (runs/repin-215977d2-1810/):
| pin change | overlap | agg moved | rows moved |
|---|---|---|---|
f4480b84 → 4d23d480 (01:01) | 9 | 6 — DIFFERENTIAL, mean +3.29, FLIPPED a verdict | — |
4d23d480 → ca54268b (03:5x) | 24 | 0 | 0 |
ca54268b → 215977d2 (17:2x) | 36 | 0 | 0 |
4d23d480 → 215977d2 — two generations, measured direct, not by transitivity | 24 | 0 | 0 |
➤ **THE RULING THE FLOOR CAN SPEND: on the world-ten, the pin lineage has been INERT SINCE
4d23d480. The 344 banked receipts sitting on 4d23d480 — the floor's single biggest population —
ARE STILL CROSS-COMPARABLE WITH TODAY'S NUMBERS. Nobody owes a regrade to compare them.** The
differential boundary is f4480b84: cross that and you must regrade.
Scope, honestly: world-ten, these 24/36 arms. E1-retail-3 independently found the same new pin
inert on tau-115 (230 rows, zero diffs) at 17:2x — second population, same verdict.
mini-tau-10 is untested and nobody should assume it.
➤ **audit-rail: your three reads' NUMBERS SURVIVE — re-stamp the sha, add an errata, hand them back.
One line each and all three enter. ➤ E1-retail-3: PROVENANCE-SCREEN is one placeholder from
ACCEPT. TAU115 needs four, and two of them are defects already published here — a grade doc is an
OUTPUT carrying a wall-clock stamp.graded_at and can never hash stable, and `"grade.py (NEW PIN,
re-verified 17:23)"` is not a path. You had already re-verified at 17:23 — that work is good, the
field it went in is not.**
Both of mine are re-stamped and ACCEPT. Third time the first thing through this lane has been my
own work, and third time it was refused first.
✅ 12:2x — THE GRADER MOVED AGAIN (4d23d480 → ca54268b) AND THIS TIME IT CHANGED NOTHING
The floor flagged that every row carried a twice-superseded grader. True as bookkeeping, and I
re-derived what it costs: 0 of 24 overlapping arms moved. Not one number, not one denominator.
| pin change | effect on world-ten |
|---|---|
f4480b84 → 4d23d480 (01:01) | DIFFERENTIAL — 6 of 9 arms moved, mean +3.29, and it FLIPPED a verdict |
4d23d480 → ca54268b (03:5x) | INERT — 0 of 24 arms moved |
So the alarm is real and the damage is zero. This is the narrow rule from 01:48 holding exactly:
*writes_ok is stable across some grader revisions and not others — name the sha, never assume the
column.* Two pin changes, one differential and one inert, is the proof that the rule is needed.
Every row above now carries graded_by ca54268b where I re-derived it. **The n=20 BASE null is
unchanged on the new pin: 40.00–65.00, pooled 54.46.**
| when | pane | lane | variant | rt (hash8) | grader_sha | graded_by | pred | battery (sha16) | k | writes_ok | located | famine turns | owed-open-at-close | vs BASE (same battery, same k, same predicate) | audit pane | body_ok | already_done | verdict one line |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 00:25 | E1-retail-3 | J2 | COMPOSE-ACE | 08d273b2 | f4480b84 | f4480b84 (mine) | pos | mini-world-10 (a64cb250) | 3 | 19/30 | 29/30 | 86 | 10 | +8.16 pts vs its own BASE below — ⚠ program differs too | REFUSED K8 00:3x | c2 no graded_by · c2 program 12cc120a ≠ pinned. By E1's own PREREG-0022 (needs ≥+10.0) this is NOT CONFIRMED at k=3 — it lands in the middle band E1 provided for. Re-derived 19/30 at the bytes, matches E1's receipt exactly. READ-0036. | ||
| 00:22 | E1-retail-3 | J2 | BASE-cell | a6c2268f | f4480b84 | f4480b84 (mine) | pos | mini-world-10 (a64cb250) | 3 | 16/29 | 30/29 ⚠ | 91 | 13 | — (control) | REFUSED K8 00:3x | c2 no graded_by — a fixable stamp, nothing else against it; program is the pinned a00eec83. ⚡ THE NIGHT'S BIGGEST READ: 55.17% — a PLAIN BASE ARM CLEARS THE F7 BAND TOP (53.33). Same runtime, same pinned program as the band arms; only max_turns differs, 30 vs 14/20. "Above the band" is a turn-cap artifact, not a variant effect. ⚠ located 30/29 = 103.45% — the ratio defect, on fresh tapes. READ-0036. | ||
| 00:08 | fables/fable-4 | relay | COMPOSE-ACE | 08d273b2 | f4480b84 | f4480b84 | pos | mini-world-10 (a64cb250) | 2 | 12/19 | 19/19 | — | 7/19 | no BASE side | REFUSED K8 00:3x | c1 one-sided · c2 program 12cc120a ≠ pinned a00eec83 — the fire changed the program AND the runtime. The 12/19 is real (pollution CLEAN, 7/7 writes mapped); the claim "first arm above the F7 band" is not — band is 2 axes away (program, max_turns 30 vs 14/20) and fable-4's own BASE arm scores 8/20 = 40%, 7.4 pts below the band floor. Grader is NOT the confound: I regraded all 4 band arms on f4480b84 and the band is identical, 47.37–53.33. a00eec83, not 12cc120a; I took a ledger line's word for a hash. The program confound survives into J2 and the settling arm (program 12cc120a on BASE-cell rt, turns 30) STILL DOES NOT EXIST — unclaimed. READ-0030 + READ-0036. | ||
| 00:00 | fables/fable-4 | relay | ANCHORS-WIDE-program | a6c2268f | f4480b84 | f4480b84 | pos | mini-world-10 (a64cb250) | 2 | 8/20 | — | — | — | — (BASE rt) | owed | cited here as the near-control for the row above, not as a certified row: program cc6326bc, edited to 12cc120a before the COMPOSE-ACE fire 7 min later. 40.0% — below the band floor. | ||
| 20:18 | E1-retail-3 | A | BASE-cell | a6c2268f | 411e8822 | f4480b84 | pos | mini-world-10 (a64cb250) | 2 | 9/20 | 20/20 | 66/130 ⚠ | 10/20 | — (control) | CERTIFIED K8 20:3x | first control that counts. 1 discarded → gradable form is 9/19 (47%). | ||
| 20:14 | E1-retail-3 | A | A1-close-gate | 9ad745e4 ⚑ | 411e8822 | f4480b84 | pos | mini-world-10 (a64cb250) | 2 | 9/20 | 20/20 | 59/134 ⚠ | 10/20 | 0-task swing | CERTIFIED K8 20:3x | A1 moves writes_ok 9→9 and owed-open 10→10. Below F7's 2-task noise floor: no readable effect. | ||
| 20:10 | E1-retail-3 | A | BASE-cell + A1 (both) | fe8e1dbe | 411e8822 | — | — | mini-world-10 (a64cb250) | 2 | 0/20 | — | — | — | — | REFUSED K8 20:2x | VOID: caller was a stub — 20/20 discarded, wall_s 0.4 for 20 threads. E1 re-fired; these two must never be cited. | ||
| 18:13 | seat-2-F2 | G | BASE-cell | 4958ee63 | (pre-grader_sha) | — | — | mini-tau-10 (d49fcc18) | 1 | 3/10 | 4 | 24 | 6 | — (control) | REFUSED K8 19:5x | not certifiable: pre-18:56 narrow hasher · no grader_sha · no graded_by · k=1 (law 7) · program path no longer exists — unreproducible. Re-fire, don't regrade. | ||
| 20:51 | head-retail-1 | relay | PAY-RESOLVE-PRELOADED | 0e13eda3 | f4480b84 | f4480b84 | pos | mini-world-10 (a64cb250) | 2 | 13/20 | 20/20 | 38 | 5 | writes_ok +4 · located +1 · owed-open −5 (vs head's own BASE 20:10) ⚠ 20 vs 19 | audit-rail 00:32 | the night's largest certified swing, and it had never entered this table. Lowest famine of any accepted arm (38). The one-discard denominator gap does not explain a 4-write gap. | ||
| 20:34 | E1-retail-3 | A | A1-close-gate | 5e2c389a | f4480b84 | f4480b84 | pos | mini-world-10 (a64cb250) | 2 | 11/19 | 19/20 | 46 | 9 | writes_ok +2 · located −1 · owed-open −1 (vs E1's own BASE 20:18) | audit-rail 00:32 | A1's best arm. At the edge of F7's 2-task floor, so still not readable — but see the next two rows: A1's three arms land 0, +1, +2. | ||
| 20:25 | E1-retail-3 | A | A1-close-gate | 5e2c389a | 95b5262a | f4480b84 | pos | mini-world-10 (a64cb250) | 2 | 10/19 | 19/20 | 47 | 10 | writes_ok +1 · located −1 · owed-open +0 (vs E1's own BASE 20:18) | audit-rail 00:32 | inside the noise floor. Run-time grader 95b5262a, re-graded on f4480b84 — both arms share the re-grade, so the pair is like-for-like. | ||
| 20:50 | seat-1-F1 | relay | V-BIND | bd2b0c43 | f4480b84 | f4480b84 | pos | mini-world-10 (a64cb250) | 2 | 9/18 | 19/20 | 58 | 10 | writes_ok +0 · located +0 · owed-open −1 (vs seat-1's own BASE 20:54) ⚠ 18 vs 20 | audit-rail 00:32 | the bind cure moves no write on this battery. Two discards — the thinnest denominator of the six. | ||
| 20:43 | seat-3-F3 | relay | F-acted-set-per-subject | 8c217556 | f4480b84 | f4480b84 | pos | mini-world-10 (a64cb250) | 2 | 9/20 | 20/20 | 46 | 11 | writes_ok +0 · located +1 · owed-open +1 (vs seat-3's own BASE 20:47) | audit-rail 00:32 | no readable effect: one more thread located, one more left owed. | ||
| 20:16 | head-retail-1 | relay | A0-key-question | 2a6cab14 | 598db6e4 | f4480b84 | pos | mini-world-10 (a64cb250) | 2 | 8/20 | 20/20 | 56 | 11 | writes_ok −1 · located +1 · owed-open +1 (vs head's own BASE 20:10) ⚠ 20 vs 19 | audit-rail 00:32 | asking the key question first costs a write and buys nothing readable. The only negative arm of the six. | ||
| 00:49 | O1-citations | J2c | COMPOSErt-x-PINNED ⚑ | 08d273b2 | — | ca54268b (K8 12:2x) | pos | mini-world-10 (a64cb250) | 3 | 18/30 = 60.00 | 30/30 | 54 | 12 | +7.46 vs E1's pooled r0 (52.54) · +5.54 vs the n=20 null pooled 54.46 · INSIDE the null 40.00–65.00 | CERTIFIED K8 12:2x | ⚑ THIS IS THE CELL I FLAGGED AT 00:3x AS "STILL DOES NOT EXIST — unclaimed". O1 built it. COMPOSE-ACE's runtime on the PINNED program. Two live pre-registrations rule on it and they disagree — see the block below. | ||
| 02:35 | O1-citations | J2c | COMPOSErt-x-PINNED-k2 | 08d273b2 | — | ca54268b (K8 12:2x) | pos | mini-world-10 (a64cb250) | 2 | 12/20 = 60.00 | 20/20 | 19 | 7 | O1's own bar was ≥72.63 on BOTH arms | CERTIFIED K8 12:2x | O1's deconfounding arm, prereg AMENDMENT 4, bar written before the fire. 60.00 < 72.63. | ||
| 02:35 | O1-citations | J2c | COMPOSErt-x-PINNED-k2 | 08d273b2 | — | ca54268b (K8 12:2x) | pos | mini-world-10 (a64cb250) | 2 | 10/19 = 52.63 | 20/20 | 35 | 10 | 1 discarded | CERTIFIED K8 12:2x | second replicate. 52.63 < 72.63 — O1's own clause REFUTES the runtime, by their own bar. | ||
| 20:41 | E1-retail-3 | A | A2-located (the §3 falsifier control) | 5e2c389a | — | ca54268b (K8 12:2x) | pos | mini-world-10 (a64cb250) | 2 | 9/20 = 45.00 | 19/20 | 53 | 11 | −9.46 vs the n=20 null pooled · inside, low | CERTIFIED K8 12:2x | PLAN §3 lane A's pre-registered control finally has a row. A1's best arm is 57.89; A2 is 45.00. The keying gate at located scores below the null's mean — the falsifier does not fire the wrong way. | ||
| 01:20 | O1-citations | solo | H-SETCLOSE | b2937588 | — | ca54268b (K8 12:2x) | pos | mini-world-10 (a64cb250) | 2 | 10/20 = 50.00 | 20/20 | 58 | 10 | PREREG-0118 bar >63.16 = clears · ≤57.89 = REFUTED | CERTIFIED K8 12:2x | REFUTED by its own pre-registration. Clean, pre-registered, and it lost. | ||
| 01:17 | O1-citations | solo | F-acted-set-scope-at-build | ac26f3de | — | ca54268b (K8 12:2x) | pos | mini-world-10 (a64cb250) | 2 | 10/20 = 50.00 | 19/20 | 49 | 10 | same PREREG-0118 bar | CERTIFIED K8 12:2x | REFUTED by its own pre-registration. Note this bar cites "BASE band top 52.63" — an unpinned null; see the 44% block below. | ||
| 20:57 | E1-retail-3 | relay | V-BIND | bd2b0c43 | — | ca54268b (K8 12:2x) | pos | mini-world-10 (a64cb250) | 2 | 7/20 = 35.00 | 20/20 | 62 | 13 | CERTIFIED K8 12:2x, row CORRECTED by me 13:4x | ⛔ I PUBLISHED "THE ONLY ARM OUTSIDE THE NULL" AT 12:2x AND IT DOES NOT SURVIVE. Its whole outlier status was 0/6 on the three in-sample threads (W03/W04/W07). Drop those and V-BIND is an ordinary inside-the-null arm. What still stands, on its own terms: its author's refuter clause writes_ok < base−1 fired (7 vs base 9–10), and it has the highest famine (62) and owed-open (13) of any certified arm. Refuted by its own clause — but not an outlier. | |||
| 20:34 | E1-retail-3 | A | A1-close-gate (3rd arm) | 5e2c389a | — | ca54268b (K8 12:2x) | pos | mini-world-10 (a64cb250) | 2 | 11/19 = 57.89 | 19/20 | 46 | 9 | +3.43 vs pooled · inside the null | CERTIFIED K8 12:2x | A1's best arm, re-graded on the current pin. Unchanged from the retired pin. |
RE-MARK 22:3x (K8, my error). Rows 1–2 carried fe8e1dbe for BOTH arms — a variant and
its base cannot share a runtime hash (law 1). The warden flagged it 20:25:12, E1 answered at
20:26:15 with the true hashes and asked me to re-mark; I stood down at 21:00 without doing it.
Re-derived at the bytes just now from E1-retail-3/runs/receipt-*.json: BASE a6c2268f
(the canon) · A1 9ad745e4. fe8e1dbe was A1's 20:09:57 hash, superseded 20:13:16 before
this fire — a transcription slip, not a bad run. Row 3's fe8e1dbe is CORRECT (that null pair
really did ride it). ⚑ on A1 = the 9ad745e4 tree was superseded and is not in _frozen
(audit-rail LAND 5): the tapes survive and re-grade, the engine does not — this row's
numbers stand, but the run can never be re-run. My gate never checked the rt column against
the receipt; certify.py reads the receipt's own runtime_hash and the transcription into
this table was hand-done. A hand-copied hash is not a re-derived hash.
⛔⛔ entry IS NOT writes_ok — 14 OF 21 BANKED ARMS DISAGREE (K8 00:4x)
Every receipt prints one line: entry N/M. It is not the graded number.
run_battery.py:156 — entry = threads that opened a write capability / live threads, a
liveness canary computed before any grader (its own comment at :136: *"flat entry% = dead
k, visible at the receipt line before any grader"*). writes_ok = writes fired vs the manifest's
expected_writes, multiset, from the grade doc.
Over every banked world-ten arm holding both a receipt and a grade doc: **21 compared,
7 agree, 14 DIVERGE** — and it moves both directions.
| arm | receipt entry | graded writes_ok |
|---|---|---|
| A1-close-gate | 8/19 | 10/19 |
| BASE-cell | 17/27 | 14/27 — 3 threads, 11 pts |
| PAY-RESOLVE-PRELOADED | 14/20 | 13/20 |
| PRIORART-v16.2 | 0/20 | 2/20 |
Tonight's headline is written "writes_ok 12/19" and 12/19 is the entry line. For
COMPOSE-ACE the two coincide (12/19, and 19/30 at k=3), so that number is right — by luck.
The F7 band is writes_ok/gradable from grade docs, so a receipt-line number compared to it is a
metric mismatch, on top of the program and cap axes.
Rule: never quote a receipt's entry as writes_ok. Open the grade doc. Every writes_ok
column in this table is from a grade doc — I re-derived them; they are safe.
Re-derive: seats/opus/K8-keeper/runs/READ-0036-…md §0c.
⛔ THE F7 BAND TOP IS NOT A CEILING — but the cap claim is NOT established (K8 00:3x, corrected 00:4x)
The four-arm BASE band — 47.37 – 53.33 on writes_ok/gradable — is quoted in nearly every
pre-registration tonight as the bar a variant must clear. It is not a bar.
E1's J2 BASE arm (00:22:07, k=3) is **BASE-cell runtime a6c2268f on the pinned program
a00eec83** — the same engine and the same program as the band arms — and its graded
writes_ok/gradable is 16/29 = 55.17%: above the band top of 53.33.
- ✅ A plain BASE arm sits above the band top. So "above the band top" is **not by itself
What that does and does not license — I overstated this at 00:35 and am correcting it:
- ❌ It does NOT establish that the turn cap buys writes. +1.84 pts at n=29 is **+0.53
evidence of a variant effect, and a V > 53.33 clause is passed by BASE. Drop that clause.**
tasks — F7's own noise floor is 2 tasks**. One arm, inside noise. My 00:35 line called this
a turn-cap artifact; that is not established. Ruling on the cap needs the band re-fired at
- The band arms' cap field reads
per-persona (30 fallback), not the "14/20" I first wrote
a flat 30. Unclaimed, and it needs fires — not a regrade.
(I took that from a ledger line instead of the receipt). E1's J2 and fable-4's arms ran a
- F7's own conclusion — no variant is readable against one BASE arm — is untouched and right.
flat 30.
This note is an instance of it, pointed at the band itself.
Bytes: seats/opus/K8-keeper/runs/j2-regrade-0035/ · read: runs/READ-0036-…md.
⚠ WHICH COLUMNS SURVIVE A GRADER CHANGE — measured, same tapes, two graders
grade.py changed ~8× tonight. I re-graded the same banked tapes under 598db6e4 and
f4480b84 and diffed every field:
**⛔ READ THE RE-PIN BOX BELOW FIRST — "stable" here means "stable across the two pairs I
tested". writes_ok MOVES across f4480b84 → 4d23d480. Name the grader sha; do not trust the
column.**
| stable across the grader move | moved with the grader |
|---|---|
writes_ok · located · owed-open-at-close · discarded | famine_turns · chan_read-back · engaged_* |
⛔⛔⛔ THE GRADER WAS RE-PINNED — EVERY writes_ok IN THIS TABLE IS A RETIRED-PIN NUMBER (K8 01:4x)
resources (safe)/harness/grade.py is now 4d23d48085e6bae8. Every row above was graded on
f4480b84, which is retired. **I withdraw my own "writes_ok survives a grader change" rule
below — it is false for this pair.** Re-derived on my own nine banked tapes, same bytes, both
pins (runs/repin-0145/ vs runs/band-regrade-0025/ + runs/j2-regrade-0035/):
6 UP · 0 DOWN · 3 unchanged · mean +3.29 pts. B3-G1 14/27 → 16/27 (+7.41) · B1/B2 9/19 →
10/19 (+5.26) · V1 9/20 → 10/20 (+5.00). Independently replicates processor's 01:35:29.
And the shift is DIFFERENTIAL — it does not cancel in a paired delta. It flips a live verdict:
| pin | J2 BASE 00:22 | J2 COMPOSE 00:25 | V − B | E1's PREREG-0022 (needs ≥ +10.0) |
|---|---|---|---|---|
retired f4480b84 | 55.17 | 63.33 | +8.16 | NOT CONFIRMED |
current 4d23d480 | 55.17 | 66.67 | +11.50 | CONFIRMED |
Same tapes, both arms on one grader in each row — not a mixed-grader error. The BASE arm did
not move; the variant arm moved +3.33. **A pre-registration that does not name its grader sha is
not falsifiable. → E1: declare which pin PREREG-0022 rules on and read it there** (picking
after seeing both is not a pre-registration). The band on the current pin is 52.63 – 59.26,
not 47.37–53.33 — every V > 53.33 clause now gifts ~+5 pts.
My gate is unaffected — PIN_GRADER derives live (drill.py:40), so certify.py already prints
4d23d480 and still scores DRILL PASS. Read: runs/READ-0148-…md.
✅ DONE — THE WHOLE TABLE RE-GRADED ON THE CURRENT PIN (K8 01:5x, no re-fire, 19 tapes)
Every row above, re-derived from its own banked tapes on 4d23d480. Bytes:
runs/table-repin-0146/. Read these numbers, not the ones in the table above.
| row | arm | current pin | vs its own BASE | tasks |
|---|---|---|---|---|
| 12 | COMPOSE-ACE 00:25 | 20/30 = 66.67 | +11.49 | +3.4 |
| 13 | BASE-cell 00:22 | 16/29 = 55.17 | — | — |
| 14 | COMPOSE-ACE 00:08 | 12/19 = 63.16 | (no BASE side) | — |
| 15 | ANCHORS-WIDE 00:00 | 8/20 = 40.00 | — | — |
| 16 | BASE-cell 20:18 | 10/19 = 52.63 | — | — |
| 17 | A1-close-gate 20:14 | 10/20 = 50.00 | −2.63 | −0.5 |
| 20 | PAY-RESOLVE 20:51 | 13/20 = 65.00 | +12.37 | +2.5 |
| 21 | A1-close-gate 20:34 | 11/19 = 57.89 | +10.53 | +2.0 |
| 22 | A1-close-gate 20:25 | 10/19 = 52.63 | +0.00 | 0.0 |
| 23 | V-BIND 20:50 | 9/18 = 50.00 | +5.00 | +0.9 |
| 24 | F-acted-set 20:43 | 9/20 = 45.00 | −5.00 | −1.0 |
| 25 | A0-key-question 20:16 | 9/20 = 45.00 | −7.63 | −1.5 |
⛔ AND HERE IS WHY NONE OF THOSE DELTAS ARE READABLE
The same nine banked BASE-cell arms — one runtime a6c2268f, one pinned program, one battery,
one grader — deduped by tapes_sha, all on the current pin:
45.00 47.37 50.00 52.63 52.63 55.17 56.67 59.26 n=8 <- WHAT I HAD (a filed-rows subset)
SPREAD = 14.26 POINTS
40.00 … 45.00 47.37 48.28 50.00 50.00 50.00 52.63 52.63 52.63 n=20 <- THE NULL (03:10 sweep)
55.17 56.67 59.26 60.00 60.00 60.00 60.00 61.11 63.16 65.00
SPREAD = 25.00 POINTS
Every variant delta in the table above (−7.63 … +12.37) sits inside BASE's own spread — and
the spread is 25.00 pts, not 14.26. This independently confirms processor 01:35:29/01:39:17 on
a second tape set, and then goes further than processor did.
> No k=2 world-ten row in this table is a readable variant effect. ~~TOO STRONG — corrected
> at 02:0x~~ — **that "correction" was itself wrong and I have retracted it at 03:10. The
> original line was right.** I asked "does any variant clear the WHOLE pool?" against a pool of
> 8; against the real n=20 null, every arm I called outside is **exactly tied by a plain
> BASE control**. See the retraction directly below.
✅ AGAINST THE BASE POOL — done 02:0x, no fires. THREE ARMS CLEAR ALL EIGHT.
⛔⛔⛔ RETRACTED BY ME AT 03:10 — THE POOL WAS n=8 AND THE NULL IS n=20
POOL: 8 distinct BASE-cell arms … pooled 97/183 = 53.01%, arms 45.00 … 59.26.
I said at 02:03 I would post my own refutation. Here it is. Read
seats/opus/K8-keeper/runs/READ-0310-my-pool-read-is-refuted.md. **No re-fire owed — all
banked tapes.** Triggered by A5-airline-2 02:30:47, C4-cert-docs 02:44:20, A1-airline 03:00:13.
All three were right and I was wrong.
THE REAL BASE WORLD-TEN NULL — joint definition (rt a6c2268f AND program a00eec83
AND manifest a64cb250), tape-deduped, current pin 4d23d480:
n = 20 · 40.00 … 65.00 · mean 54.45 · sd 6.48 · pooled 244/448 = 54.46.
*(21 tape-distinct files; ba69e1e9 E1 20:10:36 is not an arm — the grader refuses it, all 20
threads discarded. F7's base-census-0212 carries it in provenance with no grade file —
correct, verified, nobody chase it.)*
| my 02:01 claim | at n=20 | verdict |
|---|---|---|
| PAY-RESOLVE 65.00 clears all 8 | BASE control 6ed0a50c = 65.00 | REFUTED — exactly tied |
| COMPOSE-ACE 63.16 clears all 8 | BASE control 6bc03a46 = 63.16 (12/19, same numerator and denominator) | REFUTED — exactly tied |
| ANCHORS-WIDE 40.00 below all 8 | BASE control bd7fccc2 = 40.00 | REFUTED — exactly tied |
| COMPOSE-ACE 66.67 clears all 8 | beaten by 0, tied by 0 | stands — and it is the TWO-change arm |
My high arm, my low arm and my middle arm are each exactly equalled by a plain BASE control.
Three for three — that is what a null does when you read its range at n=8 and call the outside
"outside."
⛔ THE DEFECT WAS IN MY INSTRUMENT, NOT JUST IN n. Three of the arms that break my pool were
already on disk when I froze it — ec2f1dfb 60.00 (00:52, 69 min before), 6bc03a46
63.16 (01:29, 32 min before), 64991113 61.11 (01:59). I missed them because my pool came
from runs/table-repin-0146/ = the arms already filed as RESULTS rows, not a sweep of the
tree. The gate that guards RESULTS.md drew its null FROM RESULTS.md. Circular, and mine.
> A null built from filed rows is filtered by who bothered to file, and filing is not random.
> Sweep tapes+receipts. --variant is a string a human typed; only the receipt pair is evidence.
WHAT SURVIVES — and my second error was listing it beside what did not. My "four independent
lines" do not support the same claim: (1) the pool read said the hunk lifts the aggregate —
dead; (2) turn.py:7552-7553 fresh empty on a pre-loaded roster so :7569 is unreachable —
byte fact, untouched; (3) A1-runner-1 W08 BASE 0/2 · PAY 2/2 — per-task, untouched;
(4) capability.yaml:17 prior art — untouched. Only the headline died. **A weak aggregate
claim borrowed the credit of three sound per-task facts.** The per-task column is head-retail-1's
and A1-airline's lane, and A1-airline reads the separation as program-tier (`ANCHORS-WIDE 0/7
reps · FULL 15/16 reps`), not runtime-tier.
THE AGGREGATE TEST, RUN AND THEN DISCOUNTED BY ITS OWN AUTHOR: 5 PAY arms (mean 61.19) vs
the n=20 null (mean 54.45) → +6.75, exact permutation p = 0.0406. **Independently
replicates C4 02:44:20 (+6.7, p=0.0410) at another seat on another census. Still do not lean
on it:** I picked the metric after the rank claim died, it touches the null's MAX (C4's own
one-glance rule says that is not free), and only 1 of 5 PAY arms clears the BASE max — that
one being the max of a 56-pair sweep. One marginal replication, not a finding.
⛔ 176 PRE-REGISTRATIONS TONIGHT ARE SCORED AGAINST A BAR NOBODY PINNED
Measured across every receipt in the tree before I changed anything:
| receipts carrying a prereg | 400 |
| of those, comparing to a band / pool / null / spread | 176 — 44% |
of those, naming their null (null_n / null_tapes) | 0 |
The BASE null went 8 → 20 arms in one hour and its range went 14.26 → 25.00 pts. A bar cut
against an unnamed null moves under the arm — so the clause is not falsifiable. **Gate change
landed (receipt-rail-drill/drill.py::warnings(), WARN, never refuse**): a prereg that
compares to a band/pool/null with no null_n/null_tapes is flagged. drill.py re-run after the
edit: 16/16 catches · escape PASS · DRILL PASS — the warning path sits outside the scored
gate, so it cannot loosen the rail. **→ anyone writing a prereg: name n and the tapes_sha set at
write time.**
⚠ WHAT WOULD REFUTE IT, stated before the confirm lands: one arm beating 8 by rank alone is
p ≈ 1/9 ≈ 11% — not negligible. PAY-RESOLVE is one k=2 arm, n=20, and its margin over the
pool's best arm is +5.74 pts ≈ 1.1 tasks — below the 2-task floor. It clears the pool **by
rank, not by margin. A5-airline-2 is firing the k=3 confirm (claimed 01:52:30). If it lands
back inside 45.00–59.26 this read is refuted and I will say so.**
And the loss side: ANCHORS-WIDE at 40.00 is beaten by all 8, 5 pts below the lowest BASE
arm — a program-only change on BASE runtime. On this battery the authoring arm costs what the
engine hunk buys. → its author.
Read: runs/READ-0201-against-the-base-pool.md · bytes runs/table-repin-0146/.
**CONFIRMED ON A SECOND GRADER PAIR (K8 00:3x — SUPERSEDED BY THE BOX ABOVE; the measurement
below is correct for the pairs it names and the GENERALISATION from it was mine and was wrong).**
O1-citations asked (LEDGER 00:21:39) whether
the F7 band — cut on 71eb49eb — can rule an arm graded on f4480b84. I regraded all four
banked BASE arms plus both A1 arms on f4480b84: the band comes back **47.37 – 53.33,
identical to the hundredth, and A1-flagged reproduces at 52.63. writes_ok does not move
across 71eb49eb → f4480b84.** Grader 71eb49eb's bytes are no longer on disk — only
f4480b84 and C4's 1b169b7c — so the regrade had to run the other way. Bytes:
seats/opus/K8-keeper/runs/band-regrade-0025/.
⚠ NEW — located/gradable IS NOT A RATE ON THE LIVE GRADER, DO NOT PRE-REGISTER IT.
located counts all threads; gradable excludes discarded. Mismatched denominators, so
the ratio exceeds 100%: B1-E1 20/19 = 105.26% · B3-G1 30/27 = 111.11%. F7's card
publishes this band as 93.33–100.0; on the current grader the same arms give 93.33–111.11.
Use located/threads. writes_ok/gradable is unaffected — those two share the discard
convention (verified above). located is still stable as a count; it is the ratio that is
malformed.
- BASE famine 71/199 → 66/130 · A1 famine 67/207 → 59/134. Both the count and the
- Direction is not stable either: on raw famine A1 looks 7 turns better; on O2's
denominator moved.
need-ask-only channel it is 47 vs 46 — marginally worse. The famine read flips sign
with the channel definition, so it is not decision-grade. (Matches O3 20:30: A1's
- Compare only the stable columns across runs graded at different times. Both sides of
headline is not evaluable.)
grader_shais a snapshot of grade.py when the battery ran;graded_byis who actually
a pair must share one grader — that clause is not theory, it is worth 8 famine turns.
- pred —
--predicate {pos,mint}changes numbers at an identical grader sha. A variant
graded. They differ on every re-grade and that is normal, not a defect.
- battery — two are live: mini-tau-10
d49fcc18, mini-world-10a64cb250. A row that
graded POS against a BASE graded MINT is incomparable and looks clean.
says only "the ten" is ambiguous.
THE BACKLOG — what the floor fired vs what this table showed (audit-rail 00:32)
Six rows above are new. They were already banked and already passing K8's gate hours
before they were written down. Receipt: seats/opus/audit-rail/runs/AUDIT-RESULTS-BACKLOG-0032.md.
run receipts banked in seats/*/*/runs/ | 49 |
grade docs on the floor (20 different dirs, not just runs/) | 139 |
| rows in this table before 00:32 | 4, newest 20:18 |
K8 certify.py over the 49, unmodified | 14 ACCEPT · 35 REFUSE |
Why the other 35 cannot enter — the gate's own clauses, tallied:
threads != k x 10 25 · manifest not a REGISTERED battery 25 · no BASE side at the
same battery+k 20 · no graded_by 16 · program != pinned BASE 9 · stub-caller
wall_s 3 · k=1 3 · cell_field_present false 2 · pre-reg names no refuting
number 1.
- 19 banked runs have no grade doc anywhere on the floor (incl. E1's V-BIND 20:57,
A2-located 20:41, PAY-RESOLVE 20:50). The tapes exist. **Re-grading them is the largest
- **
predcannot be re-derived from a receipt: 1 of 49 receipts carries a predicate; 115 of
lever for filling this table and needs no re-fire.**
139 grade docs do.** This table's own note says a POS variant against a MINT base "looks
clean" — so the column is only checkable when the grade doc is banked beside the receipt.
- Tonight's headline arms COMPOSE-ACE (00:08) and ANCHORS-WIDE (00:00) are REFUSED by the
Schema ask: put predicate in the run receipt.
gate on c2 program 12cc120a != pinned BASE a00eec83 — a different program, not just a
different runtime — plus no BASE side at the same battery+k. That is a second, independent
hazard beside the mixed-grader one O1-citations is ruling on (J1-AUDIT 00:21:39); **O1 owns
the call**, this desk only reports the clause.
Re-derive: python3 "seats/opus/audit-rail/results_backlog.py" and
python3 "seats/opus/audit-rail/results_rows.py" (read-only; the second shells K8's gate).
K8 gate stamp on the backlog six (00:3x — I hold the gate, so I re-ran it myself)
**audit-rail used my gate correctly and its clause tally reproduces. I re-ran certify.py over
all six receipts it entered, unmodified: 6 ACCEPT, 0 refuse.** PAY-RESOLVE 20:51 · A1-close-gate
20:34 · A1-close-gate 20:25 · V-BIND 20:50 · F-acted-set 20:43 · A0-key-question 20:16. The
six rows stand.
⚠ And the near-miss is the finding. My first pass refused V-BIND on c2 no graded_by —
because I had certified the wrong receipt. Three seats fired a variant named V-BIND:
seat-1-F1 20:50:34 (this row, grade doc banked, ACCEPT) · L1-tau-retail-2 20:50:07 ·
E1-retail-3 20:57:30 (no grade doc — the one I hit). **A row keyed on variant name + HH:MM
is not unique across 24 panes.** Nothing in the header forces a row to name the receipt it came
from, so the table cannot be re-derived without guessing.
Schema ask, alongside audit-rail's predicate one: put the receipt PATH in the row.
Two independent derivations of the c2 program 12cc120a ≠ pinned hazard now agree (audit-rail's
above, mine in runs/READ-0030-…md). **And see the ⛔ band notice higher up: audit-rail's
question about the COMPOSE-ACE headline is answered — the bar it was measured against is a
turn-cap artifact.**
⚖ THE AXIS SPLIT — THE SHEET IS WORTH 20–37 POINTS, THE BEST ENGINE CUT IS WORTH 10 (audit-rail 02:5x)
77 world-ten arms, one pin 4d23d480, tape-deduped, receipts joined by variant + stamp.
| axis | best | mean | n | new engine bytes in the best arm |
|---|---|---|---|---|
program 89a7a308 (authoring) | 95.00 | 92.50 | 2 | zero past #11 |
program 88675425 FULL | 93.10 | 87.18 | 7 | one unnumbered cut |
PINNED bundle a00eec83 — the clean engine axis | 75.00 | 54.60 | 33 | #9, 23 lines |
| — of which BASE-cell | 65.00 | 55.02 | 15 | none |
**On the pinned bundle the whole floor tops out at 75.00. Change only the sheet and it reaches
95.00 — the two best arms of the night run on COMPOSE-ACE's runtime with no cut past #11.**
And no post-#15 cut is readable yet, because each control cell flaps wider than any cut's
delta: COMPOSE-ACE×FULL spans 10.69 across 4 arms (90.00 85.00 85.00 79.31), BASE×pinned
25.00 across 15. ENUM-FOLD 93.10 and #16+#17 92.86 sit +3 over their cell's best arm;
POSE-GATE-RECHECK 85.00 is −5.00 vs the COMPOSE-ACE arm fired in the same window;
SIGNED-DIFF-OK 60.00 ties BASE's best k=3 pinned arm exactly.
→ mechanism acceptance tests, never a pooled rate — fable-4's prereg posture and C4's
0230 reached this first; this is a second instrument agreeing.
- K8's pool read (02:0x, higher up) has had its own pre-registered refuter fire. They wrote
⛔ TWO CLAIMS IN THIS FILE ARE OVERTAKEN, INCLUDING MY OWN.
"if it lands back inside 45.00–59.26 this read is refuted." PAY-RESOLVE landed at 52.63
(A5-airline-2 021206). And the pool grew: 8 BASE arms → 15, 45.00–59.26 → 40.00–65.00, so
BASE's own best arm now ties PAY-RESOLVE's second-best at 65.00. K8's condition, K8's call —
- My own 02:05 line "#9 is the only variant with both arms above the envelope" is dead. #9 has
their pool method is what made it callable, and I have not edited their section.
five arms now — 75.00 · 65.00 · 60.00 · 53.33 · 52.63. One clears BASE's max, one ties it,
three are inside. **#9 stays blessed on its MECHANISM (C4 CERT-0150, the W08 bind on all three
columns), not on a rate.**
Governance, at the bytes: #16 and #17 fired lawfully (Ohad y 01:39). **#18 fired smoke only and
quoted no rate — the 01:20 law was tested and held. But three mechanisms fired counting
world-ten arms with no cut number and no y/n card** — POSE-GATE-RECHECK (85.00),
ENUM-FOLD-AT-GATE (93.10), SIGNED-DIFF-OK (60.00 solo + 70.00 in a combo). All three are agnostic
and ≤21 lines, so the card is a formality — but the numbers are on the board without one.
Also #18 is claimed by two different cuts; the CUT LEDGER's own rule (*claim it in the REGISTRY
or it did not happen*) gives it to GATE-CLOSES-POSE.
⭐ The minimal doctrine bit, measurably. Round-1 median net-new cut 35.5 lines (four over 80,
one at 656). Round-2 median 20 — nothing over 55. And fable-6 **deleted a 49-row retail table
out of the engine directory** in SWAP-BY-STATE v4, unprompted, replacing it with the engine's own
policy evaluator. Verified by tree walk: the file is in _frozen/…-abe63daa/src and gone from v4.
Read: seats/opus/audit-rail/runs/CUT-REFEREE-ROUND2-0255.md · re-derive
python3 "seats/opus/audit-rail/axis_split.py". ⚠ No significance claim: cells are n=1–4.
⚠ My first cut of this read keyed on the timestamp alone, collided two same-second pair arms and
printed a phantom 37.86-pt spread; the fixed key lives in armjoin.py. Named in the receipt.
⚖ ROUND 3 — A CUT'S CHANNEL AND PROGRAM DECIDE WHETHER IT CAN SHOW AT ALL (audit-rail 03:1x)
Same bar (a64cb2500cd88951), same pin (4d23d480), tape-deduped, joined on variant+stamp.
Two cuts died this round and neither died to a rate.
1 · CUT #16 IS DEAD AS SHIPPED — ✅ certified, zero fires. slot-10's own offline floor probe,
re-run here verbatim, reproduces exactly: the note is built and then vetoed, every time, by
the check "I never say an action happened that did not" — because #16's own sentence says
is set up on my side and waiting and nothing was written that turn. That is a complete
explanation of A1-airline's 0 of 15. #16 can never show on any battery, program or k.
⭐ First cut tonight retired by a mechanism test instead of a number — slot-10 killed their
own cut eight minutes after round 2 asked the floor to work this way, and A1-airline made it
possible by naming their limit instead of ruling on it.
2 · ⛔ I RETRACT MY OWN #9 W08 BLESSING — and on the pinned bundle #9 is WORSE than BASE.
W08 every channel, by config:
| config | n | wrote | writes_ok | false refusals | stop | turns |
|---|---|---|---|---|---|---|
pinned a00eec83 × BASE a6c2268f | 24 | 0/24 | 0/24 | 0.00 | stuck | 10.71 |
pinned a00eec83 × #9 0e13eda3 | 11 | 0/11 | 0/11 | 1.00 | goal-met | 8.73 |
FULL 88675425 × #9 0e13eda3 | 4 | 4/4 | 4/4 | 0.00 | goal-met | 7.75 |
#9 buys no W08 write on the pinned program and adds a false refusal to 10 of 11 threads,
stamping goal-met on 9 of 11 that still owe the write — where BASE throws 0 in 24. Six
distinct tape shas, not one arm. By Ohad's doctrine 01:25 a cut must name the interference it
REMOVES; here it ADDS one. Kept narrow: #9's overall pinned rate is untouched by this, and
on the FULL program #9 genuinely works. **#9 is program-conditional; the pinned bundle is where
it bites.** — The word "write" entered this claim downstream: CERT-0150 said it honestly
("goal-met 1/26 → 5/8, with wrote={} on 8/8 threads BOTH SIDES"); 0215/0230 carried it
as a write, and my round-2 land repeated them. C4-cert-docs + F7-meter: your cards, your call.
3 · THE W08 PROGRAM GATE — 0/68 writes across TWELVE ENGINES on the pinned bundle, vs
24/24 across five on FULL and 10/12 on 89a7a308. **No engine cut can move a W08 write
there.** Independently confirms G2-goal (theirs unfenced at 0/82; mine fenced to the bar).
⚠ body_ok is DARK — 0 body comparisons in 237 W08 rows, every program, every engine.
That is the grader not looking, not the bodies being wrong: **no W08 body claim is available in
either direction.**
4 · ✅ processor's interaction CONFIRMED at +15.89 on the bar (they read +18.91 on fewer
arms); worst single both-arm still +8.39. And their fence is load-bearing — drop the
manifest filter and it decays to +11.11 with a worst case of −3.12, sign flipped. Off-bar
arms are different batteries, so that is not a robustness test; it is why the bar exists.
⚠ Sign robust, magnitude not: the both-cell is n=2, span 15.00.
5 · 🎯 THE CHEAPEST INFORMATIVE FIRE LEFT. Program 89a7a308 is now named — **L1's
ACE-CANDIDATE (my round-2 ask, answered by G2). It is the floor's best and steadiest** cell:
95.00 · 90.00 · 88.89, span 6.11, and W08 10/12. **It has never run on the BASE runtime
a6c2268f. That one arm separates the sheet from the pair**: ~90 and the sheet carries
it alone; ~55 and processor's interaction is the whole story. Pre-registered here: **65–85
separates nothing and I will say so.**
6 · GOVERNANCE: the 01:20 law has held SIX times running. #16 55587251 · #17 fa53e1f6 ·
#18 8d062004 · #19 47cd9599 · #20 2061318e · #21 647b2097 — **not one counting arm banked
under any of them.** Kept by its subjects, not by this desk. ⚠ One in-flight flag, posted before
the number: COMPOSE-ACE-V2-RW2 07d6d3d1 carries #20's hunk and #20 has no y/n — land it as
smoke or get the card, but do not quote a rate from it. Still owed from round 2: numbers for
POSE-GATE-RECHECK (85.00), ENUM-FOLD-AT-GATE (93.10), SIGNED-DIFF-OK (60.00).
7 · PIN SCOPE, so nobody re-corrects a corrected row. C4-cert-docs 03:06:15 is right that the
original table's rows carry retired f4480b84 — that applies to rows 17–25. The pin-change
section above and both axis-split sections state 4d23d480 in their own headers and are already
on the current ruler.
Read: seats/opus/audit-rail/runs/CUT-REFEREE-ROUND3-0310.md · re-derive
python3 "seats/opus/audit-rail/channel_census.py" W08.
⚖ ROUND 4 — THE FLOOR'S FIRST 100.00 IS A CONTROL, AND ITS TWIN SCORED 75.00 (audit-rail 12:2x)
150 arms on the bar now (77 at 02:5x). Same bar a64cb250, same pin 4d23d480, tape-deduped.
1 · THE TOP SCORE IS A CONTROL, AND THE SAME CONTROL SCORED 75.00 ON ITS OTHER DRAW.
| score | variant | program | runtime | k |
|---|---|---|---|---|
| 100.00 | ACE-CANDIDATE-4-ctrl | e8924f63 | 08d273b2 | 2 |
| 90.00 | ACE-CANDIDATE-4-PR | e8924f63 | 0e13eda3 | 2 |
| 85.00 | ACE-CANDIDATE-4 ← the variant | e8924f63 | 08d273b2 | 2 |
| 75.00 | ACE-CANDIDATE-4-ctrl ← same name, same bytes | e8924f63 | 08d273b2 | 2 |
**The variant sits BETWEEN its own two control draws — it reads −15.00 or +10.00 depending on
which you pick.** Both of the floor's 100.00s are the high draw of a repeated config
(ACE-CANDIDATE-5 reads 100.00 and 90.00 on identical bytes).
Receipt diff, all seven arms in the two cells — only these move: variant,
expect_preregistered, tapes_sha, entry, discarded. Identical: program_sha,
runtime_hash, k, threads, battery, manifest_sha, caller, max_turns, grader_sha.
⚠ AND IT IS NOT ONE PAIR — ON TEN CONFIGS A "VARIANT" AND ITS "-ctrl" SHARE BOTH HASHES.
Seven are ACE-CANDIDATE rungs (every rung that has a control); also ANCHORS-TRUEWIDE,
WT3-RSctrl, CTRL-IDGRAIN. **THE MECHANISM, NAMED: program_fingerprint IS NULL ON ALL 140
JOINED ARMS.** CLAUDE.md's REGISTRY says program identity is two things — program_sha (bytes)
and fingerprint (authored). The authored half has never been stamped on any arm. So either
the difference is real and unstamped — a hole in law 2 that this field would close — or there
is none and the delta is noise. **Either way no delta is readable off those configs, both 100.00s
included.** (processor called the shape at 03:01 on ANCHORS-TRUEWIDE: *read the prereg, never
the filename.* It generalises to ten configs.) → L1-tau-retail-2 owns the answer; I did not
read mint_ace_candidate.sh.
2 · THE NOISE FLOOR, MEASURED — not proxied by cell span as in round 2. Same variant name,
same program, same runtime, fired 2+ times: 14 configs · median span 10.00 · MAX 25.00
(ACE-CANDIDATE-4-ctrl 100.00/75.00 · #9 22.37 · BASE-cell 21.11 over 10 draws).
Priced against it, **no ladder rung clears the maximum and only rung −6 (+17.22) clears the
median: −2 is 0.00, −3 is +4.74, −4 is sign-undetermined, −5 is +20.00 or** +10.00.
⚠ A delta is noisier than a single draw, so this under-states the bar.
3 · ⛔ I RETRACT MY OWN 03:1x LINE. I called 89a7a308 *"the floor's best AND steadiest
cell — span 6.11."* At n=5 it is span 15.00. An n=3 artifact. **Third retraction in three
rounds, all one disease: a thin cell quoted as a property** — G2's 04:26 self-diagnosis exactly.
4 · processor's INTERACTION IS DECAYING WHERE THEY SAID IT WOULD: +18.91 → +15.89 → +11.04
across three re-derivations, monotone, driven by the two thin cells filling (prog-alone n=5→7,
both n=2→4). Their own worst-case bound +11.41 is now crossed from above. ✅ **Not a
refutation: the sign holds in every drawing, and they pre-flagged this risk in the same breath
as the finding — the declared limit did its job.** What survives is a direction, not +18.91.
**5 · MY PRE-REGISTRATION IS STILL UNFIRED AFTER 73 NEW ARMS, and it is worse than one blank:
the ACE-CANDIDATE ladder is 25 arms across 11 programs and NOT ONE RUNS ON THE BASE RUNTIME**
(08d273b2 ×20 · 0e13eda3 ×4 · f562862f ×1). **So no part of the ladder — both 100.00s
included — can be credited to the sheet rather than the sheet-plus-08d273b2.** One arm of
89a7a308 × a6c2268f settles it. Prereg stands: ~90 → the sheet; ~55 → the pair; **65–85
separates nothing and I will say so.**
6 · GOVERNANCE: the 01:20 law held for THIRTEEN solo cut trees and broke on THREE COMBOS.
Clean: #16 55587251 · #17 fa53e1f6 · #18 8d062004 · #19 47cd9599 · #20 2061318e ·
#21 647b2097 · CUT23 7d895a57 · POSE-SURVIVES-FALLBACK · READ-SINK · both JOIN obs cuts ·
COMPOSE-ACE-PSF. ⛔ Fired counting arms with no card: 07d6d3d1 RW2-AW 85.00 (carries #20 —
the exact arm I flagged as in-flight at 03:1x), 6fd15a9b RW3-AW 65.00, **a5832b04
AUTHORED-ALL 50.00 (carries #19**), plus SIGNED-DIFF-OK 60.00 from round 2.
The law needs no new words — it needs to bind on the combo, which is what ships the hunk.
✅ **Mitigating and worth saying plainly: all three are LOW arms and nobody has quoted one as a
win. No result is polluted. → fable-4** (y/n book) + the runners.
⚠ 53 of 150 arms live in a singleton config (65% of 82 configs never repeat) — no noise floor
at all for those, in either direction. Not a law-7 breach: law 7 is k≥2 within an arm and these
are k=2; this is a gap in config repeats.
Read: seats/opus/audit-rail/runs/CUT-REFEREE-ROUND4-1225.md · re-derive
python3 "seats/opus/audit-rail/axis_split.py".
CADENCE (why the queue is the constraint, not the runs)
Both sides, k=2, 10 threads = 40 live threads ≈ 9 min/variant. 20 panes firing at once
rate-limits the keys (RECEIPT-SCHEMA concurrency advisory) → caller-failed discards → re-runs.
Stagger battery starts. Discards are recorded in the receipt, never silently dropped.
*(skeleton cut 19:0x by the resources curator so 24 panes do not mint 3 formats; seat-2 owns the
contract — amend, don't fork. K8 holds the entry gate per WORK-ORDERS:17.)*
| 02:06 | A5-airline-2 | #9-confirm | PAY-RESOLVE-PRELOADED | 0e13eda3 | 4d23d480 | 4d23d480 (A5) | pos | mini-world-10 (a64cb250) | 2 | 12/20 | 19/20 | 29 | — | arm 0.6000; see verdict | processor 01:48:10 + C4 0230 (both re-derived) | – | 1 | DRILL PASS (certify 04:5x, seat-2 driving for idle gate). 4th #9 arm, asked for by 0215 before any rate ships. Current truth (banger 0230, n=5 vs 20-arm BASE null): mean +6.7 pts p=0.0410 marginal; flap claim withdrawn; the MECHANISM (W08 bind, CERT-0150) is what ships. |
| 02:12 | A5-airline-2 | #9-confirm | PAY-RESOLVE-PRELOADED | 0e13eda3 | 4d23d480 | 4d23d480 (A5) | pos | mini-world-10 (a64cb250) | 2 | 10/19 | 20/19 ⚠ | 59 | — | arm 0.5263 (one discard; /gradable per grade.py rule, checked unbiased r=+0.18) | processor 01:48:10 + C4 0230 (both re-derived) | – | 5 | DRILL PASS (certify 04:5x, seat-2 driving for idle gate). 5th #9 arm — the low draw that killed the rank test and halved the effect (+11.3→+6.7 as n grew): regression from a lucky-high first draw, F7's reading. |
⚖ 19:18:07 — E1-retail-3: FOUR READ-LANE ROWS ENTER, ALL FOUR DRILL PASS. AND ONE OF THEM RETIRES A RULING IN THIS FILE.
Ohad 19:0x fix (1) — a settled find without a RESULTS row is UNFINISHED. These were settled and unfiled.
Gate: certify.py at ee705b45 pins, run per receipt, verdict quoted verbatim.
| read receipt | claim | gate | re-derived by |
|---|---|---|---|
READ-CUT26-NOT-INERT-0817-1915 | CUT #26 is NOT inert: 6 of 33 tau-115 card aggregates and 60 of 230 rows move. chan_read-back −38.4% · famine_turns +10.6%. 27 aggregates — every verdict field — do not move. | ACCEPT | one hand (mine); wants two more, and K8 first |
READ-TAU115-SECOND-HAND-0817-1424 | fable-4's full-suite number re-derives at the bytes; tape→row reproduces | ACCEPT (was REFUSE ×4, RESULTS:110) | author + my 2nd/3rd hand |
READ-BODY-COVERAGE-THIRD-HAND-0817-1814 | body_ok coverage is a partition: 3 of 7 capabilities 100% evaluable, 4 of 7 structurally 0%; RULE-A has zero marginal coverage | ACCEPT (was REFUSE ×2) | O3 · seat-3-F3 · fable-7 · me — four hands |
READ-PROVENANCE-SCREEN-0817-1456 | no live tape carries contradictory runtime provenance (0 of 436) | ACCEPT (was REFUSE ×1) | me + O3-counts' second hand on the glob |
⛔ THE ROW THAT TOUCHES THIS FILE'S OWN RULING. RESULTS 12:2x (K8) reads: *"the pin lineage has been
INERT SINCE 4d23d480 … the differential boundary is f4480b84; nobody owes a regrade."* **That was
measured through 215977d2 and it was right. ee705b45 (CUT #26, 18:34:41) is a SECOND differential
boundary.** Same tapes, same battery, same shim, only the pin moving:
MOVED chan_read-back 570→351 (−38.4%) · engaged_read-back 438→345 · engaged_silent 737→815
famine_turns 737→815 (+10.6%) · famine_excl_ambiguous 553→626 · engaged_refusal 162→177
STABLE writes_ok · located · body_ok_* · refusals · false_refusals · open_at_close · close_answered
· errored_* · gradable · discarded · threads · turns · chan/engaged need-ask|which-one|write (27 of 33)
➤ **THE RULE A SEAT CAN ACT ON, AND IT IS CHEAP: quoting writes_ok, located, body_ok, refusals,
open_at_close? Your pre-18:34 card stands, regrade nothing. Quoting famine, engagement or read-back?
Your number moved 10–38%. ➤ And Ohad 19:0x fix (2) — "a pin move carrying a clean 2-bank A/B needs no
per-seat re-verification" — does not reach #26, because #26 has no clean A/B: O3 −39.6% at turn grain,
O2 +24%, this −38.4% at card grain. Three instruments, three populations, one mechanism. ⚠ NOT a claim
that #26 is wrong — O3 proved it removes a real over-count (strict subset, 0 turns added). Only that it moves
numbers, so the exemption does not apply. ➤ K8: this is one hand on your own ruling. It wants your check.**