Run 2026-09-14_micro_v0_3¶
Preregistered decision rule, v0.3 (PREREGISTRATION-v0.3.md section 5)¶
- Validity precondition (B4-strict answered on 12/24 filler notes, rate 50.0% [31, 69], must be below 10%): FAIL
- H1 (planted recall 8/22 vs decoy false positives 1/12, Fisher one-sided): p = 0.0826 -> FAIL
- H2a (B7 planted card pairs 3 in 300 draws, expected 1.396, hypergeometric): p = 0.1651; arm-label permutation vs B1: gap = +0.0067, p = 0.3203 -> FAIL
- H2b (non-NONE rate B7 36/300 vs B1 37/300, Fisher one-sided): p = 0.5985 -> FAIL
- H3-strict (S0 two-note recall 8/22 vs B4-strict one-note recall 2/22 on bridge notes, Fisher one-sided): p = 0.0344 -> PASS (not interpretable: the validity precondition failed)
- H4 (partner-domain filler, estimation only, not a recombination test): S0 recall 8/22 = 36.4% [20, 57] vs S1 recall 0/22 = 0.0% [0, 15]
- H5 (recall by writer family, estimation only, no pass/fail):
- A: 6/10 = 60.0% [31, 83]
- B: 2/12 = 16.7% [5, 45]
Decision: NULL (signal requires the validity precondition, H1 and H3-strict all passing; H2, H4, H5 are reported and do not enter the rule).
Run summary¶
Total measured cost: $87.856028
Models:
- cards: claude-haiku-4-5-20251001 (alias haiku)
- critic: claude-haiku-4-5-20251001 (alias haiku)
- critic_haiku: claude-haiku-4-5-20251001 (alias haiku)
- critic_sonnet: claude-sonnet-5 (alias sonnet)
- generator: claude-sonnet-5 (alias sonnet)
- match_gold: claude-haiku-4-5-20251001 (alias haiku)
Exploratory finds (non-planted survivors, manual review only): 47
Oracle recall on planted: 36.4% [20, 57]; permutation p = 0.0001
T1. Corpus
| notes | cards | cards_per_note | bridges | decoys | leakage_failures | corpus_sha256 |
|---|---|---|---|---|---|---|
| 96 | 521 | 5.43 | 22 | 12 | 0 | d044765e9dff |
T2. Sampler enrichment; base rate 0.47% of 134289 cross-note card pairs are planted
| arm | draws | planted_hits | distinct_bridges | expected_hits | p_hypergeom |
|---|---|---|---|---|---|
| B1 | 300 | 1 | 1 | 1.396 | 0.754 |
| B7 | 300 | 3 | 3 | 1.396 | 0.165 |
T3. Oracle set: generator and critic without the sampler; mean cosine to gold on planted = 0.566
| S0 group | units | NONE rate | non-NONE rate | survivor rate (post critic + dupgate) |
|---|---|---|---|---|
| planted | 22 | 22.7% [10, 43] | 77.3% [57, 90] | 72.7% [52, 87] |
| decoy | 12 | 75.0% [47, 91] | 25.0% [9, 53] | 8.3% [1, 35] |
| random | 36 | 94.4% [82, 98] | 5.6% [2, 18] | 2.8% [0, 14] |
| recall on planted (gold match) | 22 | 36.4% [20, 57] |
T4. Per arm: NONE rate, critic kill rate, survivors, recovered planted bridges, measured cost
| arm | units | NONE | critic kill | dup | survivors | bridges reachable | bridges recovered | recall | cost usd | usd per recovered |
|---|---|---|---|---|---|---|---|---|---|---|
| B1 | 300 | 87.7% [83, 91] | 48.6% [33, 64] | 0 | 19 | 1 | 0 | 0.0% [0, 79] | 27.0952 | n/a |
| B4 | 68 | 63.2% [51, 74] | 12.0% [4, 30] | 0 | 22 | 22 | 2 | 9.1% [3, 28] | 7.2638 | 3.6319 |
| B7 | 300 | 88.0% [84, 91] | 36.1% [22, 52] | 0 | 23 | 3 | 0 | 0.0% [0, 56] | 27.0758 | n/a |
| S0 | 70 | 68.6% [57, 78] | 18.2% [7, 39] | 0 | 18 | 22 | 8 | 36.4% [20, 57] | 6.998 | 0.8747 |
| S1 | 44 | 88.6% [76, 95] | 20.0% [4, 62] | 0 | 4 | 22 | 0 | 0.0% [0, 15] | 3.9967 | n/a |
T5. Label-permutation null over the S0 oracle set
| statistic | observed | expected under null | p | shuffles |
|---|---|---|---|---|
| survivors among planted S0 units | 16 | 5.684 | 0.0001 | 10000 |
T6. Recombination vs single-note reflection on the same bridge notes
| arm | bridges reachable | recovered | recall | fisher p (S0 > B4) |
|---|---|---|---|---|
| S0 two notes | 22 | 8 | 36.4% [20, 57] | 0.0344 |
| B4 one note | 22 | 2 | 9.1% [3, 28] |
T7. Exploratory control (not preregistered): does the gold mechanism appear when the partner note is replaced by a mechanism-free filler from the same domain?
| arm | bridges reachable | recovered | recall | fisher p (S0 > S1) |
|---|---|---|---|---|
| S0 bridge note + true partner | 22 | 8 | 36.4% [20, 57] | 0.0018 |
| S1 bridge note + partner-domain filler | 22 | 0 | 0.0% [0, 15] | |
| B4 bridge note alone | 22 | 2 | 9.1% [3, 28] |
T8. Recall after critic, per critic (same generations, before the duplicate gate)
| critic | model id | correct recoveries | surviving | killed | decoy answers | decoy surviving | kill rate on random |
|---|---|---|---|---|---|---|---|
| haiku | claude-haiku-4-5-20251001 | 8 | 8 | 0 | 3 | 1 | 50.0% [9, 91] |
| sonnet | claude-sonnet-5 | 8 | 7 | 1 | 3 | 1 | 50.0% [9, 91] |
T9. Filler abstention under the strict single-note prompt: the single-note arm split by note kind. The filler row is the v0.3 validity precondition; bridge and decoy rows for comparison
| note kind | units | answered | NONE | errors | answer rate |
|---|---|---|---|---|---|
| bridge | 44 | 13 | 31 | 0 | 29.5% [18, 44] |
| filler | 24 | 12 | 12 | 0 | 50.0% [31, 69] |