Preregistration: daydreamd v0.2 planted-bridge recovery experiment (DRAFT, NOT SEALED)¶
Status (sealed 2026-09-13, tag v0.2.0-prereg): DRAFT. Written 2026-09-13 after the v0.1 sealed run and its adversarial review. This file becomes frozen when its SHA-256 and the SHA-256 of the three v0.2 answer-key files are committed together and stamped with OpenTimestamps (see "Sealing"). Until then it may be edited freely. After sealing, any edit above "DEVIATIONS" is a deviation and must be logged there.
Template: OSF Preregistration, open-ended form. Supersedes nothing: the v0.1 preregistration (PREREGISTRATION.md) stays as the record of the v0.1 run.
1. Title¶
Does a generate-then-select pipeline recover oblique cross-domain connections planted in a private-style note corpus, does it reject shape-matched decoys, and does two-note recombination beat one note plus a domain cue?
2. Authors¶
Julian Caraulani (caraulani@gmail.com). Design assistance from Claude Fable 5.1 (Anthropic), disclosed in paper/CHECKLIST.md.
3. Study type¶
Computational experiment with synthetic ground truth. No human participants. Owner-blind rating over a real private corpus remains Track C, deferred.
4. Background and motivation¶
The v0.1 sealed run (experiments/runs/public/2026-09-13_micro, preregistration commit 056b30e) ended NULL by its decision rule: H1 passed (8 of 12 planted bridges recovered, 0 of 6 decoy false positives after the critic, Fisher p = 0.011), H2 failed (distance-forced sampling drew 1 planted pair in 100 against an expectation of 0.65, untestable at that base rate), H3 was not significant (two notes 8 of 12 versus one note 4 of 12, p = 0.11). The adversarial review (research/06-adversarial-review-v0.1-results-2026-09-13.md) found three problems in the corpus rather than the pipeline: four single-note recoveries were leakage by paraphrase that the 6-gram check cannot detect; three of the eight two-note matches were one-note reading plus a domain cue; the six decoys shared one word each and were too easy. A post-hoc analysis found every planted bridge's closest card pair in the nearest embedding-distance band, so far-band sampling was aimed at the wrong band.
v0.2 rebuilds the corpus with oblique ingredients, a paraphrase-leak judge, shape-matched decoys, 24 bridges for power, two note-writer model families, and replaces the far-band sampler hypothesis with a near-in-embedding, far-in-domain one. It adds a direct control for "one note plus a domain cue" (arm S1) and pre-registers two critics side by side.
5. Hypotheses (pre-committed)¶
- H1 (selection). On the oracle set S0, recall on planted pairs exceeds the false-positive rate on decoy pairs. Test: Fisher exact, one-sided, alpha = .05, on 24 planted versus 12 decoys. Recall = generator non-NONE AND recovery judge MATCH (pre-critic). False positive = a decoy pair whose non-NONE output survives the v0.1 Haiku critic and the duplicate gate.
- H2 (near, cross-domain sampling). (a) Arm B7 (near band, different domains) includes more planted card pairs than arm B1 (random) at equal draws. Test: hypergeometric survival for B7's count plus a 10,000-shuffle permutation of planted labels across the candidate pool, alpha = .05, one-sided. (b) The generator's non-NONE rate on B7 exceeds that on B1. Test: Fisher exact, one-sided, alpha = .05.
- H3 (recombination beats single-note reflection). Recall of planted mechanisms in arm B4 (single note) is below recall in S0 on the same 24 bridges. Test: Fisher exact, one-sided, alpha = .05.
- H4 (recombination beats one note plus a domain cue). Recall in S0 exceeds recall in arm S1, where each bridge note is paired with a mechanism-free filler from the partner's domain. Test: Fisher exact, one-sided, alpha = .05, on 24 bridges each. This is the test the v0.1 review asked for.
- H5 (writer family, estimation only). Recall on bridges whose notes were written by the Claude-family writer versus the non-Anthropic writer. Reported as a difference with a 95 percent Wilson-based interval and a two-sided Fisher p. No pass or fail; the purpose is to bound the shared-priors effect.
Decision rule. "Signal" is declared if and only if H1 and H4 both pass. H2, H3 and H5 are reported regardless and do not enter the rule. Any other outcome is reported as a null result, in full, with the same tables.
6. Corpus¶
Synthetic private-style corpus v0.2, committed and public. Ninety-six markdown notes in the voice of one fictional solo builder across six domains, sixteen per domain, 150 to 300 words each (a note under 150 words is regenerated by the builder; over 300 is reported, not regenerated), memory-note style. Specs: data/synth/v0.2/bridges.yaml, decoys.yaml, fillers.yaml. Built notes, manifest, and gold.json will be committed to data/synth/v0.2/.
Planted bridges (24). Every one of the 15 domain pairs appears at least once; the 9 pairs between {ecom, grants, iot} and {mleval, fraud, health} appear twice; every domain carries exactly 8 bridge notes. Ingredients are observations only (counts, dates, readings, events). Each bridge carries forbidden_phrases_a and forbidden_phrases_b (mechanism nouns and verbs the writer must not use on that side) and a one_side_test sentence. Gold connection and implication are authored by the experimenter with Claude as a disclosed writing assistant, fixed under the seal, and never produced by the note-writer call or the generator.
Decoy pairs (12). Nine shape-matched decoys: each imitates the failure shape or surface pattern of a named bridge with unrelated causes and a written why_not. Three homonym decoys as in v0.1. A correct pipeline outputs NONE on all twelve.
Filler notes (24). Four per domain, no planted role.
Note writers (two families). Bridges with odd ids, and decoys and fillers at odd positions, are written by claude-haiku-4-5-20251001 (or the current Haiku snapshot, recorded from the run log). Bridges with even ids, and decoys and fillers at even positions, are written by a non-Anthropic open-weight model run locally through the ollama backend: WRITER_B_MODEL_ID = ollama/qwen2.5:7b-instruct (Qwen 2.5 7B Instruct, Alibaba; the exact model digest Ollama reports is recorded per note in the manifest). Prompt prompts/synth_note.md version 2 for both. Family A at the provider's default temperature; family B at temperature 0.7 (the Ollama backend default); no seed control; both recorded. Family B is a small open model on purpose: if a 7B model's notes are recovered at the same rate as Haiku's, shared frontier-lab priors cannot explain recovery.
Leakage check, two parts, both required for every bridge note. (1) No 6-gram of any gold connection or gold implication appears in the note. (2) The leak judge (prompts/leak_judge.md version 1, Haiku-class) returns CLEAN when shown the note alone with its gold connection and implication. A note failing either check is regenerated, up to five attempts; a bridge whose note cannot pass in five attempts is excluded before sealing the manifest and listed in the datasheet. Both check outputs are committed with the corpus.
Freezing. The corpus is frozen when its manifest SHA-256 is recorded in data/synth/v0.2/MANIFEST.sha256 and committed. Generation runs read only the frozen corpus.
7. Design¶
7.1 Concept cards, embeddings, duplicate gate¶
As in v0.1: cards with haiku (snapshot recorded), three to six per note; sentence-transformers/all-MiniLM-L6-v2 locally; cosine distance; pairs within one note excluded; duplicate gate at cosine 0.85.
7.2 Arms¶
| Arm | What is sampled | N | Purpose |
|---|---|---|---|
| S0 | Oracle set: 24 planted pairs + 12 decoy pairs + 36 random non-planted cross-domain note pairs, labels hidden from the pipeline | 72 | H1, H3, H4 reference, H5 |
| S1 | Each bridge note paired with a randomly chosen filler note from its partner's domain (mechanism-free) | 48 | H4: one note plus a domain cue |
| B1 | Random cross-note card pairs, uniform | 300 | H2 control |
| B7 | Card pairs in the nearest distance quartile (Q1) whose notes are in different domains | 300 | H2 |
| B4 | Single-note reflection over each of the 48 bridge notes | 48 | H3 |
The far-band arms B3 and B6 from v0.1 are not run in v0.2: the v0.1 post-hoc analysis showed they cannot contain planted pairs on this kind of corpus, and the v0.1 run already reports their behaviour (0 of 25 non-NONE in Q4).
7.3 Generator, critics, recovery judge¶
Generator: alias sonnet, snapshot recorded, prompts/generate_pair.md for pair arms and prompts/generate_single.md for B4, unchanged from v0.1 (SHAs recorded). NONE permitted.
Critics, two, run in parallel on every non-NONE output, both pre-registered here: (C1) the v0.1 Haiku critic, prompts/critic.md version 1, alias haiku; (C2) the same prompt with alias sonnet. Recall-after-critic and decoy false positives are reported for each. H1's false-positive definition uses C1 only, for continuity with v0.1.
Recovery judge: prompts/match_gold.md unchanged, alias haiku, MATCH or NO_MATCH with the gold in context, never sees the arm label; cosine to gold reported alongside, never used as the criterion.
8. Variables and measures¶
- Recall on planted pairs per arm (S0, S1, B4) with Wilson 95 percent intervals.
- Decoy false positives (S0), separately for shape-matched and homonym decoys, after each critic and before the duplicate gate (table T8), and after the primary critic plus duplicate gate (table T3).
- NONE rate per arm and, within S0, per label group; recall on planted pairs is also split by the writer family of the first note of each unit (H5); card pairs can span families, so no per-family NONE rate is claimed.
- Sampler enrichment: planted card pairs drawn in B7 and B1 against the hypergeometric expectation; permutation p.
- Recall by writer family (H5), with the difference and its interval.
- Critic effect: number of correct recoveries killed by C1 and by C2, with reasons.
- Cost per arm and per recovered bridge, from logs.
- Exploratory finds: non-planted survivors, reported as a count and released as text, never counted as hits.
9. Sampling plan and power¶
Counts are fixed by the corpus design and are not increased after seeing results.
- H1: 24 planted versus 12 decoys. If recall is 12 of 24 and decoy false positives are 1 of 12, Fisher one-sided p is about 0.01. If recall is 16 of 24 and false positives 1 of 12, p is below 0.001. A recall of 8 of 24 against 1 of 12 gives p about 0.09 and is reported as a fail.
- H2(a): with about 530 cards the candidate pool is about 138,000 cross-note card pairs, of which about 700 are planted (base rate about 0.5 percent). B1 with 300 draws expects about 1.6 planted pairs. B7 must draw about 6 or more to reach p below .05 by hypergeometric; the permutation test is reported alongside because the B7 pool is restricted. H2(b): with 300 units per arm, a non-NONE rate of 25 percent versus 12 percent is detected at alpha .05 with power above 0.95.
- H3: 24 bridges per arm. Recall 16 of 24 versus 8 of 24 gives p about 0.02 (power about 0.7 for that effect); 12 of 24 versus 8 of 24 gives p about 0.19 and is reported as a fail.
- H4: 24 bridges per arm (S1 has two units per bridge; a bridge counts as recovered in S1 if either unit matches). Recall 16 of 24 versus 6 of 24 gives p about 0.004.
- H5: 12 bridges per writer family; estimation only, intervals will be wide.
10. Analysis plan¶
make reproduceregenerates every table from committed raw outputs without API access.- H1: Fisher exact, one-sided, on the 2 x 2 table (planted recovered or not, decoy false positive or not).
- H2(a): hypergeometric survival for B7's planted count against the full candidate pool, with the hypergeometric p governing pass or fail; in addition, a 10,000-shuffle permutation that swaps the arm label (B1 or B7) across the pooled B1 plus B7 units and recomputes the planted-rate gap is reported alongside, never used for the decision.
- H3 and H4: Fisher exact, one-sided, on bridges recovered.
- H5: difference in recall between writer families with interval; two-sided Fisher p.
- Label-permutation null over S0: shuffle planted, decoy and random labels across the 72 units 10,000 times; the statistic is the number of post-critic survivors among planted S0 units (as in v0.1, table T5); report the permutation p.
- Wilson 95 percent intervals on every proportion. Cost from logs.
11. Exclusion criteria¶
- Generator outputs that fail JSON parsing after one repair attempt count as NONE and are logged with raw text.
- API errors are retried three times; a unit that still fails is excluded and listed by id.
- A bridge whose note fails the leak check in five attempts is listed in
manifest.leak_judge.failures; the experimenter removes it fromgold.jsonand from all counts, records the removal in the datasheet, and only then seals the manifest. The builder itself never edits the answer key. - No unit is excluded for its content.
12. What we will report regardless of outcome¶
Tables for every arm, every raw output, every prompt, every model snapshot id for both writer families and the leak judge, the corpus build cost including leak-judge calls, the corpus, the answer key, both leakage-check outputs, the run logs, the cost. A null on any hypothesis is written into the abstract. The v0.1 result stays in the paper as the first run.
13. Threats to validity, stated before the run¶
- Shared priors. The bridges were written with a Claude-family assistant and the generator is Claude-family. v0.2 bounds the note-writer half of this with two writer families (H5) but does not remove the bridge-authoring half; a bridge set written by an independent human remains future work.
- Obliqueness is judged by a model. The leak judge is a Haiku-class model; it can miss implication by world knowledge. Its outputs are released so the judgement can be audited.
- Shape-matched decoys are still authored. A decoy's
why_notis the experimenter's claim that the synthesis is wrong; a reviewer may disagree on individual pairs. All twelve are released with their reasons. - Synthetic ground truth measures recovery of planted structure, not real-world novelty or usefulness.
- Single corpus, single generator. No claim about other corpora or other models is made.
Track C (future). Owner-blind scoring over a real private corpus with a sealed answer key, as described in the v0.1 preregistration, Section 13. Not part of v0.2.
14. Pre-committed abstract template¶
"We plant 24 oblique cross-domain connections in a 96-note synthetic corpus written in the style of one builder's private notes by two model families, add 12 shape-matched decoys, and measure whether a generate-then-select pipeline recovers the connections while rejecting the decoys. On the oracle set the recovery judge accepted [x] of 24 planted pairs and [y] of 12 decoys survived the critic (Fisher p = [p1]). Pairing each bridge note with a mechanism-free note from the partner domain recovered [z] of 24 (Fisher p = [p4] against the true pairs); single-note reflection recovered [c] of 24 (p = [p3]). Near-band cross-domain sampling drew [a] planted card pairs in 300 draws against a random expectation of [e] (p = [p2]), and its non-NONE rate was [r7] percent against [r1] percent for random pairs. Recall on bridges written by the two writer families was [f1] and [f2]. By the preregistered decision rule (H1 and H4 both required) the run reports [signal | no signal]. The Haiku critic killed [k1] correct recoveries and the Sonnet critic [k2]. Synthetic ground truth measures recovery of planted structure, not real-world novelty or usefulness; we release the corpus, answer key, both leakage checks, prompts, raw outputs and a reproduction script."
Claims not permitted in v0.2 text: "novel ideas", "discovers", "anticipates", any lead time, any cost headline not read from logs, "first", "benchmark", generalization beyond this corpus, any sleep-biology claim.
15. Sealing¶
WRITER_B_MODEL_IDis filled in Section 6. Write this file and the answer key (data/synth/v0.2/bridges.yaml,decoys.yaml,fillers.yaml). Do not run any model call for the experiment before step 5. Development smoke tests on a throwaway corpus are not the experiment; they are logged underexperiments/runs/public/*_smoke/and never enter a table.- Compute
shasum -a 256of this file and of the three answer-key files, list the three answer-key hashes below, and commit the four files together. - Stamp:
ots stamp PREREGISTRATION-v0.2.md data/synth/v0.2/bridges.yaml data/synth/v0.2/decoys.yaml data/synth/v0.2/fillers.yaml; commit the.otsproofs. Tag the seal commitv0.2.0-prereg. - Record the commit hash here after commit:
PREREG_COMMIT = TBD. - Build the notes from the sealed specs with both writer families (
make synth SPEC=data/synth/v0.2 WRITERS=experiments/micro-v0.2/writers.yaml, after filling family B's exact model id inwriters.yaml; the builder refuses aTBDmodel), run both leakage checks, freeze the corpus manifest and recordCORPUS_MANIFEST_SHA256 = TBD. - Only then run the v0.2 micro configuration (
experiments/micro-v0.2/config.yaml, to be added with the sizes in Section 7.2) andmake reproduce.
Answer-key hashes at seal time (SHA-256): TBD, filled at step 2.
Sealed-file hashes at seal time (SHA-256):
dcb4e4b639a308e90885ca5a82714b704d70411b10e5d3ea207142885b393145 data/synth/v0.2/bridges.yaml
fb7abfecc2a6a5250b60db9cde5e418dd4619a38fd055b20b82af59b036ef2da data/synth/v0.2/decoys.yaml
c8a613d27987b5f10d06b388ac9e522f1e34d7bbccd2d5eb106eeb9f3e575d25 data/synth/v0.2/fillers.yaml
41d4bf1df5a305e0a07f42fa637ceb75f19739ac31a46cadacda8be18c096295 prompts/leak_judge.md
9555386b1075e4c53fe64a80a563310480f88723c40ef861bab83daeac9684a3 experiments/micro-v0.2/config.yaml
7a896ab471a954ac3ad1e43c789ce2c61b4d25c971b2872c0338d76d5ce6f903 experiments/micro-v0.2/writers.yaml
The hash of this file itself is the git blob recorded by the sealed commit and by its .ots proof.
DEVIATIONS¶
Format: date, section, what changed, why, who decided.
- 2026-09-13, Section 6 (writer family B) and
experiments/micro-v0.2/writers.yaml. The sealed tagollama/qwen2.5:7b-instructcould not be pulled: the Ollama registry's blob host (a Cloudflare R2 endpoint) timed out from this network on two attempts. The same model and quantization, Qwen 2.5 7B Instruct Q4_K_M, was pulled as a single-file GGUF from Hugging Face (hf.co/bartowski/Qwen2.5-7B-Instruct-GGUF:Q4_K_M, Ollama digest recorded per note in the manifest). The writers file was edited to that tag before the build; no note had been written. Decided by the experimenter. - 2026-09-13, Section 6 (note length) and the builder. The first full build (96 notes, both families, leak judge on) produced 46 notes under the 150-word floor, 31 of them under 100 words, almost all from family B: Qwen 2.5 7B under the identical prompt writes 60 to 140 words, so a 150 floor rejects one family systematically and would confound H5 with note length. The floor is lowered to 100 words for both families. The corpus is rebuilt in resume mode: notes that already pass the 6-gram check, the em-dash check and the 100-word floor are kept as written; bridge notes among them that the leak judge never saw (it runs only on notes that pass the other checks) are judged now on their existing text; everything else is regenerated with five fresh attempts. The first build's manifest values (cost, timestamps, corpus hash) are carried in the new manifest under
resumed_fromandcost_usd_prior. Decided by the experimenter, before any experiment call on this corpus. - 2026-09-14, Section 6 (five attempts) and the builder's cleaner. After the resume rebuild, 16 family-B notes were still under 100 words and 12 of the 13 family-B bridge notes that the leak judge had never seen were judged LEAK when judged on their existing text: the small model writes tersely, so the two ingredients sit next to each other and the judge reads the implication off the page, and it appends meta-commentary ("Note: the note adheres to...") that the cleaner did not strip. Haiku's notes from the same specs passed. Excluding every LEAK bridge would remove 9 of the 12 family-B bridges and make H5 unmeasurable. Decision: the cleaner now strips trailing meta-commentary (a builder fix, no prompt change), and the failing notes only are regenerated in resume mode with up to ten attempts instead of five, judged by the unchanged leak judge; notes that still fail are handled by the exclusion rule in Section 6. Prompt and judge are unchanged. Decided by the experimenter, before any experiment call on this corpus.
- 2026-09-14, Section 6 (exclusion) and corpus freeze. After the ten-attempt pass, 10 family-B notes were still under 100 words; the leak judge only runs on notes that pass the other checks, so these bridge notes had no verdict. They were judged once each on their existing text after the build (recorded as
judged_after_buildin the manifest) rather than left unjudged. Final LEAK verdicts: br03-a, br15-b (family A) and br08-a, br08-b, br10-a, br12-b, br16-a, br16-b, br18-b, br24-a (family B). The exclusion rule removes bridges br03, br08, br10, br12, br15, br16, br18 and br24 fromgold.json: 16 planted bridges remain, 10 from family A and 6 from family B. This changes the power stated in Section 9 (H1 now 16 planted vs 12 decoys; H5 compares 10 vs 6 bridges and is estimation only, as preregistered). Four notes had a trailing one-line summary starting "Note:" removed after the build (dc04-a, fl02, fl22, fl24); shas recomputed. Short family-B notes are kept and listed. Decided by the experimenter, before any experiment call on this corpus. - 2026-09-14, outcome. Sealed run
experiments/runs/public/2026-09-14_micro_v0_2(generator claude-sonnet-5; cards, match and one critic claude-haiku-4-5-20251001; second critic claude-sonnet-5; measured list-price cost $85.04). H1 PASS (6/16 planted recovered, 0/12 decoy answers survive, p = 0.021). H2a FAIL (B7 and B1 each drew 2 planted card pairs in 300, expected 1.09; p = 0.30), H2b FAIL (non-NONE 58/300 vs 47/300, p = 0.14). H3 FAIL, inverted (B4 single-note recall 8/16 vs S0 6/16, p = 0.86). H4 PASS (S0 6/16 vs S1 0/16, p = 0.009). H5 estimation: family A 3/10, family B 3/6. Decision by the rule in Section 5 (H1 and H4): SIGNAL. The experimenter's reading, recorded here so the rule's outcome and the interpretation are both on the record: the rule omitted H3, and H3 inverted; the run demonstrates selection (H1) and that a mechanism-free partner does not elicit the mechanism under the pair prompt (H4), but it does not demonstrate that two notes are needed, because one note under the single-note prompt recovered more mechanisms than two notes under the pair prompt. "Recombination" is not claimed from this run.