Skip to content

Run 2026-09-14_x6_strict_prompt_v02

Preregistered decision rule (PREREGISTRATION.md section 5)

  • H1 (planted recall 0/0 vs decoy false positives 0/0, Fisher one-sided): p = 1.0000 -> FAIL
  • H2 (sampler enrichment, hypergeometric): -> FAIL
  • H3: not computable (missing S0 or B4)

Decision: NULL (signal requires H1 and H2 both passing).

Run summary

Total measured cost: $9.930771

Models: - cards: claude-haiku-4-5-20251001 (alias haiku) - critic: claude-haiku-4-5-20251001 (alias haiku) - critic_haiku: claude-haiku-4-5-20251001 (alias haiku) - generator: claude-sonnet-5 (alias sonnet) - match_gold: claude-haiku-4-5-20251001 (alias haiku)

Exploratory finds (non-planted survivors, manual review only): 0

Oracle recall on planted: 0.0% [0, 0]; permutation p = 1.0000

T1. Corpus

notes cards cards_per_note bridges decoys leakage_failures corpus_sha256
96 491 5.11 16 12 13 f8a6e9e33577

T2. Sampler enrichment; base rate 0.38% of 119205 cross-note card pairs are planted

arm draws planted_hits distinct_bridges expected_hits p_hypergeom

T3. Oracle set: generator and critic without the sampler; mean cosine to gold on planted = 0.000

S0 group units NONE rate non-NONE rate survivor rate (post critic + dupgate)
recall on planted (gold match) 0 0.0% [0, 0]

T4. Per arm: NONE rate, critic kill rate, survivors, recovered planted bridges, measured cost

arm units NONE critic kill dup survivors bridges reachable bridges recovered recall cost usd usd per recovered
B4 55 47.3% [35, 60] 27.6% [15, 46] 0 21 16 6 37.5% [18, 61] 6.4149 1.0692

T5. Label-permutation null over the S0 oracle set

statistic observed expected under null p shuffles
survivors among planted S0 units 0 0.0 1.0000 2000

T8. Recall after critic, per critic (same generations, before the duplicate gate)

critic model id correct recoveries surviving killed decoy answers decoy surviving kill rate on random
haiku claude-haiku-4-5-20251001 0 0 0 0 0 0.0% [0, 0]

T9. Filler abstention under the strict single-note prompt: the single-note arm split by note kind. The filler row is the v0.3 validity precondition; bridge and decoy rows for comparison

note kind units answered NONE errors answer rate
bridge 32 14 18 0 43.8% [28, 61]
filler 23 15 8 0 65.2% [45, 81]