Adversarial review of the v0.1 sealed run (2026-09-13)¶
Reviewer: Claude Fable 5.1 acting as a hostile-but-fair area chair, with raw-output access to experiments/runs/public/2026-09-13_micro/. Every number below was recomputed or read from generations.jsonl, critic.jsonl, match_gold.jsonl, gold.json, the snapshot notes and the tables. Unit ids are quoted so anyone can check.
1. Verdict¶
There is something here, but it is smaller and different from what the brief hoped. What survives: on this corpus the generator is selective in the right direction. It spoke on 11 of 12 planted pairs, abstained on 5 of 6 decoys and 34 of 42 random cross-domain pairs, and the grounded judge accepted 8 of 12 planted answers as the gold mechanism, with the 4 rejections being real mechanism mismatches, not judge noise. H1 passes by the preregistered test. What dies: H2 (distance-forced sampling) is null and was never testable at a 0.65 percent base rate; worse, the post-hoc analysis shows planted bridges are embedding-near, so the far-band arms could not have contained them by construction. H3 is not significant (8/12 vs 4/12, Fisher p = 0.11), and the four single-note recoveries expose the real problem: several bridge notes state their half of the mechanism nearly outright, so "recovery" is often reading one note well, not recombining two. The decision rule ("signal" only if H1 and H2 both pass) therefore returns NO SIGNAL, and that must be the headline. The honest positive is narrower: a NONE-permitted generator plus a grounded match judge can separate mechanism-sharing pairs from vocabulary-sharing pairs on a small synthetic corpus, and the critic as configured destroys half of the correct recoveries. Probability that the selection thesis holds on a real corpus, given this run: unchanged from prior, maybe slightly up; this run mostly calibrated the instrument and found two flaws in it (critic, corpus obliqueness).
2. Hardest attacks¶
H1: selection beats decoys (PASSES, wounded)¶
Attack (a): the 8 "matches" are paraphrases of one note, not recombination. Evidence: read each against gold and both notes. - Real two-note recombination, defensible: S0-0000 (br01) combines A's missing payment-age check with B's payment-recency signal; S0-0002 (br03) abstracts "moving reference vs fixed boundary" from a dose-reminder note and a gateway note, which neither states; S0-0003 (br04) freeze-plus-cutoff from an A/B note and a temporal-split note; S0-0008 (br09) short-window velocity vs weekly sum; S0-0001 (br02) anchoring across judge and grant panel. - Mostly reading note B: S0-0011 (br12). Note br12-b says verbatim "The screening happens after purchase, not before... Better to catch it upfront". The generated connection adds only that the grant side has the same shape. S0-0005 (br06): note br06-a says "Can't yet separate improvement source. Supplement, ring measurement, or combination." The mechanism is on the page. S0-0006 (br07): "thundering herd" is inferred from br07-a's power-restore storm and applied to br07-b's checkout timeout; reasonable, but br07-a effectively names the mechanism. - Verdict: 5 of 8 matches are genuine cross-note recombination; 3 are one-note reading plus a domain transfer. The corpus, not the pipeline, is at fault: the note writer was handed the ingredients and, for several bridges, wrote the conclusion. Wounds H1's interpretation, does not kill the test as run.
Attack (b): the match judge is lenient. Evidence: the 4 non-matches (S0-0004 br05 "relational checks, not sequential identifiers"; S0-0007 br08 "period tagging, not continuous telemetry"; S0-0009 br10 "discrete lapse, not expectation lag") are all correct rejections of a related-but-different mechanism; the judge distinguished mechanism from topic. Among matches, S0-0011 and S0-0005 are the softest (mechanism stated in one note), but the judge's job is gold equivalence, not novelty, and they are equivalent. One lenient B4 match: B4-0002 (br02-a alone) was accepted because it restates anchoring from note A, which is the mechanism, so the judge is consistent. Misses.
Attack (c): the critic kills correct recoveries. Evidence: critic killed S0-0005 (restates_claim), S0-0006 (generic), S0-0011 (generic), all three judge-confirmed matches, plus S0-0004, S0-0007, S0-0009 (restates_claim; those were non-matches anyway). Of 8 correct recoveries, 4 survived the critic. The T5 permutation null (5 survivors among planted vs 2.2 expected, p = 0.032) is computed on post-critic survivors, so the preregistered H1 statistic passes despite the critic, but the critic halves recall for no gain in precision on decoys (the one decoy non-NONE, S0-0016, was killed by the critic; the random-pair survivors, 6 of 42, were not). The "generic" reason on br07 and br12 is wrong: a thundering herd and a screen-before-commit are specific mechanisms. As configured (Haiku, single call, "an empty result beats a boring one"), the critic is a liability on this task. This is the most actionable finding in the run.
Attack (e): decoys are too easy. Evidence: decoys share one word (cart, ring, token, protocol, call, benchmark) and nothing else. That is vocabulary matching, but the weakest form; a model that reads past a homonym passes trivially. The one decoy non-NONE (S0-0016, dc05) is a plausible grant-fit observation, not a hallucinated mechanism, and the critic killed it. Decoys at this difficulty cannot support the claim "selection rejects plausible-but-wrong connections". Wounds H1: the pass is against a low bar. v0.2 needs decoys that share a failure shape but not a mechanism (e.g. two notes both about "spikes" with unrelated causes).
H2: distance-forced sampling enriches planted pairs (NULL, and untestable as designed)¶
Attack (f): power. Base rate 0.65 percent of 52,887 card pairs; 100 draws expect 0.65 planted. B1 drew 1, B3 drew 1, B6 drew 0; hypergeometric p = 0.48. Detecting even a 3x enrichment (2 vs 0.65) at alpha .05 needs several hundred draws per arm. The preregistration's own power section said this. H2 was a hypothesis the design could not test. Kills H2 as evidence either way. The NONE gradient (Q1 12/25 non-NONE, Q2-3 5/25, Q4 0/25, top 5 percent 2/25) is the informative result: abstention rises monotonically with distance. Two readings: the generator is honest (far pairs in a 60-note corpus mostly have no connection, and it says so) or lazy (it uses surface overlap as its cue). The 8 non-NONE random S0 pairs and the exploratory finds in Q1 (below) favour honest-but-cue-driven: it speaks when notes share structure. Cannot separate honest from lazy without a far-band pair that is known to connect; the corpus contains none, by construction.
H3: two notes beat one note (NOT SIGNIFICANT, and the corpus is the reason)¶
Attack (d): the four B4 recoveries leak. Evidence: B4-0002 (br02-a): the note states "the judge context includes prior scores. Strip that out and the issue disappears"; the mechanism is on the page. B4-0010 (br06-a): "Can't yet separate improvement source. Supplement, ring measurement, or combination." Stated. B4-0017 (br09-b): the note lists "skipped Tuesday, took 2x Wednesday... Both weeks ended at 7 total, within protocol spec... Tracker records only the weekly sum, doesn't flag the skip-then-double pattern". The gold is one sentence away. B4-0020 (br11-a): "Even pinned snapshots still drift... Scores from three months ago vs. today on the same reference set show measurable difference"; the reference-set re-scoring is stated as current practice. All four are corpus leakage: the note writer wrote the bridge's conclusion on one side. The 6-gram check caught none of this because the wording differs. Indicts corpus construction, not the pipeline. Also note B4 has no abstention (0 NONE in 24) despite the prompt permitting it, and its critic kill rate (25 percent) is lower than S0's (45 percent): the single-note prompt produces confident, specific, mostly-correct implications, which is what makes the incubation control dangerous for the thesis. H3's 8 vs 4 is consistent with "two notes add something" but at n = 12 it is also consistent with nothing.
Post-hoc distance finding (real, but half artifact)¶
Attack (g): all 12 planted bridges have their closest card pair in Q1 (median min distance 0.608 vs 0.737 for random cross-domain pairs) because the note writer was told the same ingredients for both halves; two paraphrases of one spec are near by construction. That explains part of it. What the artifact does not explain: mechanism-sharing notes in real corpora also share ingredients (a fraud note and a refund note that both mention "payment method added in the last 7 days" are near because the world made them so). The finding is a real property of mechanism-sharing pairs and an artifact of how strongly this writer copied ingredients. It cannot be disentangled on this corpus. It does establish one thing firmly: the brief's far-band sampler is aimed at the wrong band for this kind of bridge, and the cross-domain-near arm (B7) is the right next test.
3. The 36 exploratory finds (non-planted survivors)¶
Single-rater classification, reading connection, mechanism and implication against both notes. Counts: (i) plausible genuine unplanted connection, specific to both notes and checkable: 15. (ii) generic bridge-like synthesis ("both are about X"): 8. (iii) restatement or arithmetic on one note: 6. (iv) confabulation, a fact not in either note doing the work: 7.
Examples of (i): - B1-0007 (br01-b x fl11): "The 14-day appeal window is shorter than the months chargebacks take to surface fraud, so prepaid-only restriction substitutes for the chargeback evidence that hasn't arrived yet." Two notes, one real inference, checkable. - B3-0010 (br03-a x fl13): injection timing drifts ~20 min/day, so a liver panel drawn at a fixed clock time floats relative to the last dose and confounds AST/ALT trends; fix: draw at a fixed post-dose interval. Neither note says this. - B1-0000 (dc01-a x fl08): a month-long discount test's last two days are censored by the 36 to 48 hour review lag. Small, correct, useful. - B3-0082 (br06-b x fl10): the failed-authorisation burst is an upstream signal of card testing, independent of the charges it precedes. Examples of (ii): S0-0059, B1-0083, B3-0044 ("both are about trend validity"). Examples of (iv): B3-0024 (iron deficiency raises resting heart rate: outside knowledge, not in the notes), B1-0069 (assumes power-loss recovery involves repositioning the unit), B6-0015 (refunds caused by token truncation: two unrelated notes forced together).
The surprising positive is real but modest: roughly 40 percent of unplanted survivors are specific, grounded, checkable cross-note inferences the corpus author did not plant. On a private corpus these are exactly the "huh" items; here they are also the first evidence that the generator does more than retrieve planted structure. They need a second rater before any number is published.
4. What the paper may say (abstract, real numbers)¶
"We plant 12 cross-domain connections in a 60-note synthetic corpus written in the style of one builder's private notes, and measure whether a generate-then-select pipeline recovers them. On the oracle set, the generator produced a non-NONE answer on 11 of 12 planted pairs and the grounded judge accepted 8 of 12 as the planted mechanism (Wilson 95 percent interval 39 to 86 percent); it abstained on 5 of 6 decoy pairs and 34 of 42 random cross-domain pairs, and no decoy answer survived the critic (label-permutation p = 0.032 on post-critic survivors). Distance-forced sampling did not enrich planted pairs: banded sampling drew 1 planted pair in 100 draws and anchor-plus-remote drew 0 in 50, against a random-pairing expectation of 0.65 (hypergeometric p = 0.48); this hypothesis was underpowered by design, and a post-hoc analysis shows every planted bridge's closest card pair lies in the nearest distance band. Single-note reflection recovered 4 of 12 mechanisms, against 8 of 12 for two notes (Fisher p = 0.11, not significant); reading the four single-note recoveries shows the corpus states those mechanisms on one side. By the preregistered decision rule (H1 and H2 both required) the run reports no signal. The binary critic killed 4 of the 8 correct recoveries. Synthetic ground truth measures recovery of planted structure, not real-world novelty or usefulness; we release the corpus, answer key, prompts, raw outputs and a reproduction script."
Not permitted, still: "novel", "discovers", "first", any lead time, any cost per idea beyond "measured list-price cost of the run was $33.62".
5. What to run next, ranked by information per dollar¶
- Critic ablation on the existing outputs (cost: about $2, Haiku only). Re-run the critic with (a) a Sonnet-class model, (b) the "empty beats boring" line removed, (c) a two-vote rule, and report recall-after-critic vs decoy false positives for each. The critic is the cheapest lever and the clearest flaw.
- S1 partner-domain control (already queued, about $3). Right run: it directly measures how much of S0's recovery is one note plus domain cue. Given the leakage found in section 2(d), expect S1 to recover several bridges; that number is the honest ceiling on "recombination".
- B7 cross-domain-near (queued, about $9). Right run for the sampler question, but interpret with care: on this corpus it will enrich planted pairs partly because of the paraphrase artifact. Report enrichment and recall, and say so.
- Second rater on the 36 exploratory finds (human, 30 minutes). Turns the one interesting positive into a reportable number.
- Corpus v0.2 before any "supports selection" sentence is repeated: (a) obliqueness: forbid the note writer from stating the mechanism's conclusion, and add a check that a single-note reflection with the gold in context does not rate the note as already containing it; (b) decoys that share a failure shape (two "spike" notes, two "drift" notes) with different mechanisms; (c) a second note-writer family (GPT or Gemini via OpenRouter, or a local model) for at least half the bridges; (d) 24 bridges, so H3 has power; (e) drop H2 in its current form; replace with a testable sampler hypothesis at note level (B7 vs B1 on 300 draws each).
- X2 generator-size runs (about $35 each): informative but lower priority than fixing the corpus; a bigger generator on a leaky corpus tells you about leakage.
6. Is the code pivot warranted?¶
Warranted in direction, premature in sequence. The numbers say the pipeline's weak parts are the critic and the ground truth, and both are exactly what execution fixes: a failing test is a critic with no opinion and ground truth with no leakage. The prior-art sweep says the measurement (distance-forced pairing, abstention rate, random control, permutation null, yield-vs-distance) is unclaimed on code, while the verifier is off the shelf. Two things should happen first, in this order, because they are cheap and they change the v0.2 design: run S1 and the critic ablation on the existing outputs (one evening), and rebuild the corpus with oblique bridges and shape-matched decoys (one build, about $2) so the synthetic instrument is trustworthy as the benchmark companion. Then move the headline experiment to real repositories with execution as the critic, carrying over three things this run established: near-in-embedding, far-in-source pairing as the primary sampler (not far-band), NONE rate as a reported quantity, and a second rater or a differential oracle on every survivor. The daydreaming framing stays; the verifier changes.