Skip to content

Run 2026-09-14_x7_rejudge_votes5

Preregistered decision rule, v0.3 (PREREGISTRATION-v0.3.md section 5)

  • Validity precondition (B4-strict answered on 12/24 filler notes, rate 50.0% [31, 69], must be below 10%): FAIL
  • H1 (planted recall 10/22 vs decoy false positives 1/12, Fisher one-sided): p = 0.0296 -> PASS
  • H2a (B7 planted card pairs 3 in 300 draws, expected 1.396, hypergeometric): p = 0.1651; arm-label permutation vs B1: gap = +0.0067, p = 0.3203 -> FAIL
  • H2b (non-NONE rate B7 36/300 vs B1 37/300, Fisher one-sided): p = 0.5985 -> FAIL
  • H3-strict (S0 two-note recall 10/22 vs B4-strict one-note recall 2/22 on bridge notes, Fisher one-sided): p = 0.0078 -> PASS (not interpretable: the validity precondition failed)
  • H4 (partner-domain filler, estimation only, not a recombination test): S0 recall 10/22 = 45.5% [27, 65] vs S1 recall 0/22 = 0.0% [0, 15]
  • H5 (recall by writer family, estimation only, no pass/fail):
  • A: 8/10 = 80.0% [49, 94]
  • B: 2/12 = 16.7% [5, 45]

Decision: NULL (signal requires the validity precondition, H1 and H3-strict all passing; H2, H4, H5 are reported and do not enter the rule).

Run summary

Total measured cost: $88.273001

Models: - cards: claude-haiku-4-5-20251001 (alias haiku) - critic: claude-haiku-4-5-20251001 (alias haiku) - critic_haiku: claude-haiku-4-5-20251001 (alias haiku) - critic_sonnet: claude-sonnet-5 (alias sonnet) - generator: claude-sonnet-5 (alias sonnet) - match_gold: claude-haiku-4-5-20251001 (alias haiku)

Exploratory finds (non-planted survivors, manual review only): 47

Oracle recall on planted: 45.5% [27, 65]; permutation p = 0.0001

T1. Corpus

notes cards cards_per_note bridges decoys leakage_failures corpus_sha256
96 521 5.43 22 12 0 d044765e9dff

T2. Sampler enrichment; base rate 0.47% of 134289 cross-note card pairs are planted

arm draws planted_hits distinct_bridges expected_hits p_hypergeom
B1 300 1 1 1.396 0.754
B7 300 3 3 1.396 0.165

T3. Oracle set: generator and critic without the sampler; mean cosine to gold on planted = 0.566

S0 group units NONE rate non-NONE rate survivor rate (post critic + dupgate)
planted 22 22.7% [10, 43] 77.3% [57, 90] 72.7% [52, 87]
decoy 12 75.0% [47, 91] 25.0% [9, 53] 8.3% [1, 35]
random 36 94.4% [82, 98] 5.6% [2, 18] 2.8% [0, 14]
recall on planted (gold match) 22 45.5% [27, 65]

T4. Per arm: NONE rate, critic kill rate, survivors, recovered planted bridges, measured cost

arm units NONE critic kill dup survivors bridges reachable bridges recovered recall cost usd usd per recovered
B1 300 87.7% [83, 91] 48.6% [33, 64] 0 19 1 0 0.0% [0, 79] 27.0952 n/a
B4 68 63.2% [51, 74] 12.0% [4, 30] 0 22 22 2 9.1% [3, 28] 7.2638 3.6319
B7 300 88.0% [84, 91] 36.1% [22, 52] 0 23 3 0 0.0% [0, 56] 27.0758 n/a
S0 70 68.6% [57, 78] 18.2% [7, 39] 0 18 22 10 45.5% [27, 65] 6.998 0.6998
S1 44 88.6% [76, 95] 20.0% [4, 62] 0 4 22 0 0.0% [0, 15] 3.9967 n/a

T5. Label-permutation null over the S0 oracle set

statistic observed expected under null p shuffles
survivors among planted S0 units 16 5.684 0.0001 10000

T6. Recombination vs single-note reflection on the same bridge notes

arm bridges reachable recovered recall fisher p (S0 > B4)
S0 two notes 22 10 45.5% [27, 65] 0.00785
B4 one note 22 2 9.1% [3, 28]

T7. Exploratory control (not preregistered): does the gold mechanism appear when the partner note is replaced by a mechanism-free filler from the same domain?

arm bridges reachable recovered recall fisher p (S0 > S1)
S0 bridge note + true partner 22 10 45.5% [27, 65] 0.000261
S1 bridge note + partner-domain filler 22 0 0.0% [0, 15]
B4 bridge note alone 22 2 9.1% [3, 28]

T8. Recall after critic, per critic (same generations, before the duplicate gate)

critic model id correct recoveries surviving killed decoy answers decoy surviving kill rate on random
haiku claude-haiku-4-5-20251001 10 10 0 3 1 50.0% [9, 91]
sonnet claude-sonnet-5 10 9 1 3 1 50.0% [9, 91]

T9. Filler abstention under the strict single-note prompt: the single-note arm split by note kind. The filler row is the v0.3 validity precondition; bridge and decoy rows for comparison

note kind units answered NONE errors answer rate
bridge 44 13 31 0 29.5% [18, 44]
filler 24 12 12 0 50.0% [31, 69]