Skip to content

daydreamd: A Selection Engine for Machine-Generated Connections over a Private Corpus, Evaluated by Planted-Bridge Recovery

Julian Caraulani caraulani@gmail.com, ORCID 0009-0006-3889-3563

September 2026. Draft v0.1 with the sealed run filled in. Every number in Section 6 is read from results/public/2026-09-13_micro/, regenerated by make reproduce from committed run logs. The experiment is preregistered in ../PREREGISTRATION.md; the decision by its rule is stated in the abstract.


Abstract

Three preregistered runs test whether a generate-then-select pipeline recovers cross-domain connections planted in a synthetic corpus written in the style of one builder's private notes. v0.1 (60 notes, 12 bridges): given the right two notes, the generator answered on 11 of 12 planted pairs and the grounded judge accepted 8 of 12 (Wilson 39 to 86 percent); it abstained on 5 of 6 decoys and 34 of 42 random pairs, and no decoy answer survived the critic (Fisher p = 0.011). Distance-forced sampling drew 1 planted pair in 100 against 0.65 expected (p = 0.48), and single-note reflection recovered 4 of 12 against 8 of 12 for two notes (p = 0.11); by the preregistered rule the run reports no signal. v0.2 (96 notes, 16 bridges written obliquely on both sides by one of two model families, shape-matched decoys, two critics): the generator answered on 9 of 16 planted pairs and the judge accepted 6 (Wilson 18 to 61 percent); it abstained on 10 of 12 decoys and 33 of 36 random pairs, and no decoy answer survived the primary critic (p = 0.021). Pairing a bridge note with a mechanism-free filler from the partner domain produced 0 recoveries in 16 (p = 0.009), the preregistered H4, and the run reports SIGNAL by the preregistered rule. That rule's recombination test was not one: single-note reflection recovered 8 of 16 mechanisms, including five the two-note oracle had marked NONE (H3 inverted, p = 0.86), and a follow-up over every note in the corpus found the single-note prompt answers on 23 of 23 fillers and 21 of 21 decoy notes. The planted mechanisms are recoverable from one note, and recombination is not demonstrated. Near-in-embedding, far-in-domain sampling drew 2 planted card pairs in 300 against 1.09 expected (p = 0.30). Recovery did not differ by writer family (3 of 10 Haiku-written, 3 of 6 Qwen-written), with the caveat that the exclusion rule removed 6 of 12 Qwen bridges. Both critics removed correct recoveries (3 of 6 and 1 of 6). v0.3 (96 notes, 24 fact-type bridges of which 22 were certified by the generator itself, under a strict NONE-permitted single-note prompt with a two-vote judge, to need both sides): two-note recombination recovered 8 of 22 planted mechanisms (36 percent, Wilson 20 to 57) and single-note reflection under the strict prompt recovered 2 of 22 (9 percent), against 8 of 16 for one note in v0.2; selection missed its bar (8 of 22 against 1 of 12 decoy answers surviving both critics, p = 0.083); the preregistered validity precondition failed because the strict prompt answered on 12 of 24 filler notes, and on inspection all twelve answers are inferences from specifics the fillers contain, so the precondition measured filler richness rather than abstention. The run reports NULL by its preregistered rule; three judge decisions a human reader would reverse (units S0-0002, S0-0004, S0-0015) would move recall to 11 of 22, and we report the preregistered number. Recovery differed by note-writer family (6 of 10 Haiku-written, 2 of 12 Qwen-written); the review attributes the gap to note length and quality (bridge-note medians 166.5 against 108.5 words, cosine to gold 0.65 against 0.49), and shared priors are not separable from length in this design. Measured list-price costs: $33.62, $85.04 and $87.86 for the runs, $15.05 for the follow-up, $1.54, about $7.0 and about $18.00 for the corpus builds. Synthetic ground truth measures recovery of planted structure, not real-world novelty or usefulness; the corpora, answer keys, prompts, raw outputs and a reproduction script are released.


1. Introduction

Language models generate candidate ideas cheaply and in bulk. What they lack is a way to tell which candidates are worth anything. A physicist at Anthropic, using Claude for quantum chromodynamics work, put it this way: "LLMs are profoundly creative. They simply lack a sense of which paths might be fruitful before walking them" (Schwartz, 2026). Ríos-García et al. (2026) measured the same thing at scale: across more than 25,000 agent runs in eight domains, 41.4 percent of outcome variance is attributable to the base model and 1.5 percent to the agent scaffold, evidence is ignored in 68 percent of traces, and the authors conclude that "outcome-based evaluation cannot detect these failures, and scaffold engineering alone cannot repair them." Si, Yang and Hashimoto (2024) found that 4,000 generated research ideas collapsed to about 200 unique ones. Generation is not the bottleneck. Selection is.

This paper describes daydreamd, a pipeline that samples pairs of distant concepts from a corpus of one person's notes, asks a language model whether a genuine connection exists (with explicit permission to answer that none does), and passes survivors through a binary critic and a retrieval-based duplicate gate. The system is designed to run as a background daemon over a private corpus, but the daemon is packaging. The contribution we can defend in this version is the evaluation protocol: a synthetic corpus with hand-planted connections and decoys, a preregistered analysis, a permutation null, and a full release of prompts, raw outputs and logs.

Zahn, Evans and Eagleman (2026) recently argued that memory consolidation exists to drive cross-domain recombination, and validated a replay-based system against 50,000 OpenAlex cross-field pairs and a held-out 2026 window. Their work establishes "dreaming as recombination" as a theoretical frame with credentialed authors and a large public-literature validation. We differ on four points. First, their pipeline has no critic and no permission to decline; every replay produces output. Second, they have no distance control at selection; distance is measured post hoc. Third, their holdout is not leakage-safe by construction, which their limitations section concedes, whereas our v0.1 ground truth is planted by hand and checked for n-gram leakage. Fourth, their corpus is public literature, whereas our target is a private corpus whose owner is the only competent judge. We treat their result as the strongest available evidence that recombination over memory is worth doing, and our protocol as the missing measurement of whether a given pipeline actually does it.

The rest of the paper is organized as follows. Section 2 places the work among consolidation systems, ideation pipelines over public literature, temporal-validation methods, the literature on judge unreliability, and creativity science. Section 3 describes the system. Section 4 states the selection thesis. Section 5 gives the evaluation protocol. Section 6 holds the result tables from the sealed run and the post-hoc audits. Section 7 lists limitations, including the ones we expect reviewers to raise first.

Four design commitments, each traceable to a prior finding:

  1. The generator may say NONE. A prompt that demands a connection forces confabulation. Structured output with a required testable_implication field is itself a filter, since vague output cannot produce one. The NONE rate doubles as an honesty measure per model and per distance band.
  2. Novelty is a retrieval check, not an opinion. The literature has discredited LLM-rated novelty (Section 2.4). In daydreamd, "already in the corpus" is a cosine threshold against real documents, and "matches the planted answer" is a grounded entailment judgment with the answer in context, never a novelty score.
  3. A stimulus control and a statistical null, kept separate. Random pairing (arm B1) is the stimulus control: the same pipeline, the same fluency, a different sampling policy. Label permutation is the statistical null. The founding brief conflated these; the adversarial review in research/02-sanity-review-2026-09-13.md separated them.
  4. Recovery before usefulness. No claim about usefulness is made until the pipeline has demonstrated that it can recover connections known to be present and decline connections known to be absent.

2.1 The day-dreaming loop and its prior implementations

Gwern (2025) proposed a background loop in which a generator explores non-obvious links between concept pairs drawn from a knowledge store, a critic filters for novelty and usefulness, and survivors are written back to compound. The essay has not been modified since 2025-07-14 and links no implementation beyond a toy. Goedecke (2025) built idea-mill within a day of the essay, described it as "pretty half-assed", hand-wrote facts in YAML because extraction was unreliable, and still reported "a few genuinely novel ideas". zby (2025) piloted temporal novelty testing with pre-cutoff models and stopped at two stated walls: "I didn't implement the search algorithm, I manually selected combinations" and "a domain-agnostic novelty verifier is the fundamental research bottleneck". Vault Daydream (glebis, 2025) runs Sonnet generators and Haiku critics over an Obsidian vault with a 7.0 threshold and no evaluation. None of these measure whether the loop finds anything that is there. daydreamd differs by evaluating recovery against planted ground truth before making any other claim, and by replacing the hand-selected combinations with a sampler whose enrichment is measured.

2.2 Consolidation systems

Sleep-time compute (Lin et al., 2025) runs a background agent over an idle context to precompute and reorganize. Generative Agents (Park et al., 2023) periodically synthesizes stored memories into reflections. Auto-Dreamer (Ye et al., 2026) learns a fast per-session acquisition and a slow cross-session consolidation, reporting a 7-point gain on ScienceWorld at 12 times less memory. Sleep-Consolidated Memory (Shinde, 2026) models NREM and REM stages with forgetting. Anthropic's Dreams API, OpenClaw's Dreaming stage, and OpenDream all perform deduplication, contradiction resolution and cleanup. Every one of these consolidates; none generates connections and none filters for novelty. daydreamd builds no consolidation and treats it as commoditized. It borrows one contract shape from the Dreams API: an asynchronous job with immutable input, a separate output store, and a human review gate.

2.3 Ideation pipelines over public literature

SciMON (Wang et al., 2023), ResearchAgent (Baek et al., 2024), MOOSE-Chem (Yang et al., 2025), the AI co-scientist (Gottweis et al., 2025), BioDisco (Ke et al., 2025), FARS (Tang et al., 2026), Alien Space (Artiles et al., 2026) and SciMuse all run generator-then-critic loops over published literature. SciMuse personalizes, but from a researcher's published record. Alien Space's mechanism is the relevant warning: when prompted for novel ideas, language models "recombine high-density regions of the literature". Co-scientist's proximity agent (a similarity graph for deduplication) and meta-review agent (recurring review patterns fed forward into the next round) are the two components daydreamd adopts directly. The difference is the substrate. Every system above reads what everyone can read. A private corpus is the one setting where the judge is the world expert on the input, and where the ideas cannot have been recombined from high-density public literature because they are not in it.

2.4 Temporal validation and the unreliability of LLM novelty judgment

MOOSE-Chem froze a model with a pre-2024 cutoff and tested against 51 post-January-2024 chemistry papers. HindSight (Jiang, 2026) used Llama-3.3-70B with a June-2023 cutoff and a six-month safety margin, and found that LLM-judged novelty correlates negatively with anticipating real future research (ρ = −0.29, p < 0.01). CKM (Tao et al., 2026) measured a mean of 404 days (median 399, range 66 to 757) between hypothesis and matching paper. Wainrib et al. (2026) ran 800 independently replicated experiments and showed that permuting the hit and miss labels erases a 53.4 percent gain, which is the null-design daydreamd adopts. On the judge side, FARS reports automated review scores of 5.00 against 3.23 from 88 human reviewers; the Ideation-Execution Gap (Si et al., 2025) found that ideas rated novel before execution scored below human ideas after execution; RQ-Bench (Sinhahajari et al., 2026) documents a "novelty mirage" in which judges rate model-written research questions as highly novel while experts prefer author-anchored ones; Ideation Arena (Chen et al., 2026) reports that LLM judges align with 105 expert researchers' pairwise preferences only 72.56 percent of the time, and that scaffolds vary wildly, some underperforming their base model. RINoBench (Schopf and Färber, 2026) evaluates nine automated novelty metrics against 1,381 expert-judged ideas and finds that reasoning aligns with human rationales but "does not reliably translate into accurate novelty judgments". AgentIdeaBench (Mo et al., 2026) scores originality against retrieved prior art across 33 models and 40 subfields. Nusrat and Nusrat (2025) audited three Kosmos hypotheses against random-gene nulls and found one of three indistinguishable from noise, caught only by the null model.

daydreamd inherits the old-cutoff and temporal-holdout methodology as hygiene and does not claim it. Its v0.1 differs from all of the above in three ways: the ground truth is planted by hand rather than inferred from later publication; the judge is asked a grounded entailment question with the answer in context, never a novelty question; and a permutation null is preregistered. Ideation Arena, RINoBench and AgentIdeaBench occupy the "novelty benchmark for public-literature ideas" slot; daydreamd does not claim that slot.

2.5 Creativity science

Boden (1990) classifies this work as combinational creativity. Campbell (1960) and Simonton frame it as blind variation and selective retention. Mednick (1962) named remote association, though his flat-hierarchy mechanism was refuted by Benedek and Neubauer (2013) and is not relied on here. Uzzi et al. (2013), over 17.9 million papers, found that the highest-impact work combines a conventional core with an injection of atypical combinations, a two-dimensional prescription that motivates arm B6. Liu et al. (2026), with 140 humans, 2,800 ideas and seven models, found that semantic distance predicts originality linearly (humans β = .10, models β = .18) with no quadratic term, and Shen, Druckmann and Zou (2026) obtained their best results by maximizing distance (novel-solution rate from 1.6 percent to over 50 percent at fixed model and temperature). Orwig et al. (2025) found a plateau, not a peak. These results kill the intermediate-distance "fertile zone" for raw novelty; what survives, and what nobody has plotted, is usable yield as novelty times feasibility against distance. That curve is not part of v0.1.

Three further results constrain the design. Peeperkorn et al. (2024) showed temperature is a risk parameter, not a creativity parameter, so sampling policy is never conflated with temperature. FunSearch (Romera-Paredes et al., 2024) got results with a fast weak model and about a million samples and reported insensitivity to model choice, but AlphaEvolve (Novikov et al., 2025) reverses this in the same lab: it "performs increasingly better as the underlying LLM improves", trading roughly a thousandfold in samples for capability. The defensible statement is that a generate-evaluate-select loop is necessary for verified novelty at any capability level, and better models make the search cheaper, not unnecessary. Argument Collapse (Kim et al., 2026) reports 65.3 percent unique arguments from humans against 3.4 percent from models, and Shared Imagination (Zhou et al., 2024) shows models hallucinate alike, so multi-model ensembling is not a diversity fix. Chen, Zhao and Cohan (2026) warn that models already over-produce "bridge-like opportunities and synthesis methods", exactly the output shape of a combination daemon.

On the biological analogy: Wagner et al. (2004) largely failed to replicate (Brodt et al., 2018; Schönauer et al., 2018), so no sleep-biology claim is made. What survives is Hoel's (2021) architectural argument that a system injecting noise during live operation would fail at its job, so a dedicated offline period is needed, and the PAD model (Deperrois et al., 2022) in which ablating the mixing of multiple memories costs 4.4 points on CIFAR-10 and 18 points on SVHN, evidence that recombination rather than mere replay does measurable work. Model collapse (Shumailov et al., 2024) destroys distributional tails first, and remote associations are the tails; any writeback loop must segregate generated material, anchor each cycle in fresh real input, and gate on human verification. DreamCoder (Ellis et al., 2021) is the existence proof that a fantasy-plus-replay loop can be done safely. Hoel's overfitted-brain hypothesis has, as far as we can find, never been connected to language models; that framing is available and is used lightly here.

2.6 The personal-informatics null

Strömel et al. (CHI 2024, N = 273) is the best-powered test of language-model narratives over personal data. Generated narratives moved engagement, attention and reward, and produced a flat null on the Insight subscale: F(2, 270) = 0.43, p = .64. PhysioLLM's generated-insight arm did not beat a plain data summary, and an audit of 14,922 generated explanations over personal sensor data (Zhu et al., 2026) found that models "routinely attribute anomalous days to causes without sufficient support". This null does not test our hypothesis for four reasons: the substrate was seven days of step counts, which contain no latent conceptual structure; the operation was summarization of one dataset, not recombination of two distant concepts; the participants had no expertise in their own step counts; and there was no filter, no rejection, no permission to say NONE. What transfers is the warning that enjoyment is not insight and self-report cannot tell them apart. That is why the human-scored track (Section 7) is a blind test with a sealed key, and why v0.1 uses no self-report at all.

2.7 Claim discipline

Tao et al. maintain a ledger for AI contributions to Erdős problems that classifies each as 1(a) AI-independent with no comparable literature, 1(b) AI solution with literature found afterwards, 1(c) AI building on known literature, 1(d) AI plus human, or 2(a) to 2(d) literature search, formalization, rewriting, computation. The cautionary case is October 2025, when a model was publicized as solving about ten open Erdős problems and had in fact located existing literature solutions. Every future "hit" from daydreamd over a real corpus will be classified against this ledger before the word novel is used. In v0.1, no hit is called novel; hits are called recovered.


3. System

3.1 Pipeline

ingest → concept cards → local embeddings → pair sampler (B1 | B3 | B6)
      → generator (NONE permitted, structured JSON)
      → binary critic → duplicate gate (retrieval, cosine 0.85)
      → morning.md → owner verdicts → writeback (unverified until endorsed)

The engine never knows where the corpus came from. Adapters supply documents: a Claude Code memory directory, an Obsidian vault, a codebase, a Zotero library, or the synthetic corpus used in this paper.

3.2 Concept cards

One cheap-model pass per document distills atomic claims with source pointers into a tight schema (id, source_note, claim of at most 30 words, entities, why_it_matters, confidence). This is the step Goedecke identified as broken in idea-mill. Cards are cached and incremental. In v0.1, card text is committed so anyone can rate fidelity; human fidelity rating is deferred to the first real-corpus run.

3.3 Embeddings

all-MiniLM-L6-v2 runs on device. The privacy rule "your notes never leave your machine" has to hold for embedding, not just storage.

3.4 Samplers

  • B1, random. Uniform over cross-note pairs.
  • B3, banded. Quantile bands of the cross-note cosine-distance distribution: Q1, Q2 to Q3, Q4, top 5 percent. Equal draws per band. The far tail is included because that is where the distinctive prediction lives.
  • B6, anchor plus remote. One note from the densest embedding cluster paired with one from Q4, following Uzzi's conventional-core-plus-atypical prescription.

3.5 Generator

Prompt contract, committed verbatim in prompts/generator.md: two claims from one person's notes; most pairs are unrelated; if no genuine, non-obvious connection exists, output exactly NONE; otherwise JSON with connection (at most 40 words), mechanism, testable_implication (one concrete thing the owner could check this week), and needs (which claim supplies what). No restating. No facts absent from the inputs.

3.6 Critic and duplicate gate

The critic is binary and cheap: kill if the output restates a claim, if the implication is not checkable, or if the connection is generic. The duplicate gate embeds the surviving connection and removes it if cosine similarity to any card or note chunk exceeds 0.85. Neither component rates novelty.

3.7 Morning review and writeback (not exercised in v0.1)

Survivors land in morning.md with provenance. The owner approves or rejects each with a reason. Rejections feed a per-user learnings file that tunes the critic (the co-scientist meta-review pattern). Endorsed dreams gain sampling weight and become concept cards. Every dream carries lineage; a dream built on a later-rejected dream is flagged. Generated material is never silently folded into ground truth.

3.8 Ship shape

Local-first command-line tool, bring-your-own key or a headless Claude Code session, pluggable generator (any API or local model), built-in scheduler, no telemetry. Configuration: corpus paths, schedule, dream count, sampler, models, strictness.


4. The Selection Thesis

daydreamd is not a better generator. It is a selection engine for machine-generated connections, deployed over a corpus whose owner is the ground-truth expert. Three observations support treating selection as the frontier and generation as commoditized.

First, every serious builder of a day-dreaming loop stopped at verification (Section 2.1). Second, the field has measured that LLM-rated novelty anti-correlates with real value (Section 2.4), which removes the obvious cheap judge. Third, a private corpus is the one setting where a competent human judge is available for free, because the owner already knows what is in their own notes and what would be worth acting on.

The thesis makes a prediction that v0.1 tests in the narrowest form: a pipeline built on these principles should recover connections that are known to be present, decline connections that are known to be absent, and surface planted pairs more often when the sampler forces distance than when it does not. If it fails the planted test, no later claim is worth making.


5. Evaluation Protocol

Full detail is in ../PREREGISTRATION.md, which is hashed and timestamped before any model call. This section summarizes it.

5.1 Synthetic corpus with planted ground truth

Sixty notes in the voice of one fictional solo builder across six domains (ten each), 150 to 300 words, in memory-note style. Twelve planted bridges: pairs of notes in different domains carrying, obliquely, the ingredients of one hidden mechanism, each with a human-written gold_connection (at most 40 words) and gold_implication. Six decoy pairs with shared vocabulary and no mechanism. Thirty-six filler notes. A leakage check ensures no 6-gram of any gold connection appears in any note. Corpus, manifest hashes and gold.json are committed and public. There are 1,770 cross-note pairs, of which 12 are planted, a base rate of 0.68 percent.

5.2 Arms

S0 (oracle set: 12 planted, 6 decoy, 42 random non-planted cross-domain pairs, labels hidden) tests generator and critic independent of any sampler. B1 random (100 pairs), B3 banded (25 per band, four bands), B6 anchor plus remote (50 pairs). B4 single-note reflection over the 24 bridge notes, the incubation control.

5.3 Measures

Sampler enrichment of planted pairs against the hypergeometric base rate; recall on planted pairs (non-NONE and a grounded entailment MATCH against gold, with cosine reported but not used as the criterion); specificity as NONE rate on decoys and random pairs; false-positive rate after critic on decoys; B4 recall; NONE rate by band; cost per arm and per recovered bridge from logs; exploratory finds on non-planted pairs, reported separately and never counted as hits.

5.4 Hypotheses and decision rule

H1: S0 recall on planted pairs exceeds the false-positive rate on decoys (Fisher exact, one-sided, α = .05). H2: B3 and/or B6 surface more planted pairs than B1 (hypergeometric and 10,000-shuffle permutation, α = .05). H3: B4 recall is below pair-arm recall (Fisher). Signal is declared only if H1 and H2 both pass. Anything else is reported as null, in full, with the same tables.

5.5 Threats

The notes are drafted by a language model and the generator is a language model; they may share priors. Mitigations: bridges and gold text authored by the experimenter with a language model as a disclosed writing assistant and fixed under the preregistration seal before any note existed, the leakage check, decoy pairs, and a planned second corpus and bridge set written by a different model family or by an independent human. Synthetic ground truth measures recovery of planted structure, not real-world novelty or usefulness. The recovery judge is a language model asked a grounded entailment question with the answer in context, the narrow use the literature supports; its prompt and every judgment are released.


6. Experimental Results

Every number below is read from results/public/2026-09-13_micro/, regenerated by make reproduce from the committed run experiments/runs/public/2026-09-13_micro/ without API access. Model snapshot IDs, prompt hashes and the access date are read from that run's metadata.yaml.

T1. Corpus. Filled by make reproduce (T1.md).

Notes Domains Bridge notes Filler notes Planted pairs Decoy pairs Cross-note card pairs Cards Leakage failures
60 6 24 36 12 6 52,887 328 (5.47 per note) 0

T2. Sampler enrichment of planted pairs. Filled by make reproduce (T2.md). Base rate: 0.65 percent of the 52,887 cross-note card pairs are planted.

Arm Pairs drawn Planted included Distinct bridges Hypergeometric expectation P(at least observed)
B1 random 100 1 1 0.654 0.482
B3 banded 100 1 1 0.654 0.482
B6 anchor+remote 50 0 0 0.327 1.000

T3. Oracle set S0: recall and specificity. Filled by make reproduce (T3.md). Mean cosine between generated connection and gold on planted pairs: 0.428.

Pair type N NONE rate (Wilson 95%) Non-NONE rate Survived critic and dup gate MATCH to gold
Planted 12 8.3% [1, 35] 91.7% [65, 99] 41.7% [19, 68] 8 of 12, 66.7% [39, 86]
Decoy 6 83.3% [44, 97] 16.7% [3, 56] 0.0% [0, 39] n/a
Random non-planted 42 81.0% [67, 90] 19.0% [10, 33] 14.3% [7, 28] n/a

Fisher exact (8 of 12 planted recovered vs 0 of 6 decoy false positives after the critic), one-sided: p = 0.0113.

T4. Per-arm NONE rate, kill rate, survivors, recovery and measured cost. Filled by make reproduce (T4.md). Cost is USD at list price read from logged token usage; the marginal cost on a subscription was zero.

Arm Units NONE % (Wilson 95%) Critic kill % of non-NONE Dup-gate kills Survivors Bridges reachable Bridges recovered Recall USD USD per recovered bridge
S0 60 66.7 [54, 77] 45.0 [26, 66] 0 11 12 8 66.7% [39, 86] 5.95 0.74
B1 100 84.0 [76, 90] 12.5 [3, 36] 0 14 1 0 0.0% [0, 79] 8.77 n/a
B3 100 81.0 [72, 87] 26.3 [12, 49] 0 14 1 0 0.0% [0, 79] 8.85 n/a
B6 50 96.0 [87, 99] 0.0 [0, 66] 0 2 0 0 0.0% [0, 0] 4.08 n/a
B4 24 0.0 [0, 14] 25.0 [12, 45] 0 18 12 4 33.3% [14, 61] 2.86 0.71

NONE rate by distance band, arm B3 (non-NONE out of 25 per band): Q1 12, Q2 to Q3 5, Q4 0, top 5 percent 2. Exploratory finds (non-planted survivors) by arm: B1 14, B3 14, S0 6, B6 2, B4 0; 36 in total.

T5. Permutation null over S0. Filled by make reproduce (T5.md). The statistic is the count of post-critic survivors among the 12 planted S0 units, with the planted, decoy and random labels shuffled across the 60 S0 units.

Statistic Observed Mean under null Permutation p (10,000 shuffles)
Survivors among planted S0 units 5 2.197 0.0320

T6. Single-note reflection (B4) versus the two-note oracle arm. Filled by make reproduce (T6.md).

Arm Bridges tested Mechanisms recovered Rate (Wilson 95%) Fisher p (S0 greater than B4)
S0 two notes 12 8 66.7% [39, 86] 0.11
B4 single note 12 4 33.3% [14, 61] n/a

Preregistered decision rule (DECISION.md, verbatim).

  • H1 (planted recall 8/12 vs decoy false positives 0/6, Fisher one-sided): p = 0.0113 -> PASS
  • H2 (sampler enrichment, hypergeometric): B3 p = 0.482, B6 p = 1.000 -> FAIL
  • H3 (S0 two-note recall vs B4 one-note recall, Fisher one-sided): p = 0.1102 -> FAIL

Decision: NULL (signal requires H1 and H2 both passing).

Exploratory finds. Count: 36. Released as results/public/2026-09-13_micro/exploratory_finds.jsonl. Not counted as hits.

Model and prompt provenance. Run id 2026-09-13_micro, executed 2026-09-13. Generator snapshot: claude-sonnet-5. Card, critic and recovery-judge snapshot: claude-haiku-4-5-20251001. Note-writer snapshot for the corpus: claude-haiku-4-5-20251001. Embedding model: sentence-transformers/all-MiniLM-L6-v2 (sentence-transformers 6.0.1). Prompt SHA-256: cards f2a7ac1a, critic 49ec635b, generate_pair feddedd8, generate_single 1547b975, match_gold 09d4576c (full hashes in the run's metadata.yaml). Measured list-price cost: $33.62 for the run, $1.54 for the corpus build. Seal commit 056b30e37a3e9f72e7977b190a0f7f13a22903fb, tag v0.1.0-prereg. Corpus sha256 f17c4f0864305f54ce92f3a3422b3ccf65acfbcd84d072ec87e334e2b8b4c571.

6.1 Post-hoc and exploratory (not preregistered)

Everything in this subsection was decided after the seal. The post-hoc distance analysis and the exploratory arms were written and committed while the sealed run was executing, after reading only its pre-generation artifacts (sampled units, cards, embeddings); the audits below were done after reading its outputs. None of it enters the decision rule.

Where the planted pairs sit in distance space. For each of the 12 planted bridges, the closest card pair between its two notes falls in the nearest distance band (Q1): 12 of 12. Median minimum card distance between bridge notes is 0.608, against 0.737 for random cross-domain note pairs. Band edges over the 52,887 cross-note card pairs: q25 0.827, q75 0.963, q95 1.042. Two notes that carry the same hidden mechanism share its ingredients, so they are domain-far but embedding-near. Far-band sampling therefore aims at the band that contains none of them, which is one reason H2 could not pass. Part of this is an artifact of a note writer that was told the ingredients; part of it is plausibly a property of mechanism-sharing pairs in general. Which part is which is a v0.2 question.

Abstention rises with distance. In arm B3, non-NONE answers per band of 25: Q1 12, Q2 to Q3 5, Q4 0, top 5 percent 2. Anchor-plus-remote (B6): 2 of 50. The generator is honest at distance on this corpus, or lazy; the two are not separable here.

Critic audit. The binary critic killed 3 of the 8 correct recoveries (units S0-0005, S0-0006, S0-0011; reasons restates_claim, generic, generic) and correctly killed the one decoy answer that got past the generator (S0-0016, generic). As configured, the critic buys specificity at the price of recall; T5 is computed on post-critic survivors (5 of 12), which is why it is weaker than the pre-critic recall (8 of 12).

Single-note leakage audit. All four B4 recoveries (B4-0002, B4-0010, B4-0017, B4-0020) correspond to notes that state the mechanism on one side in paraphrase. The 6-gram leakage check does not catch paraphrase. This indicts the corpus construction, not the pipeline, and it is the main reason H3 is not significant.

Exploratory finds. Of the 36 non-planted survivors, a single rater classified 15 as plausible unplanted inferences, 8 as generic bridge-like syntheses, 6 as restatements and 7 as confabulations. Single rater; not a reportable number until a second rater scores the same items.

X1, partner-domain control (run 2026-09-13_x1_partner_control, same corpus, prompts and models, new seed; measured list-price cost $9.62). Each of the 24 bridge notes was paired with a filler note from its partner's domain that carries the domain's vocabulary but none of the mechanism's ingredients (arm S1). The oracle set was re-run alongside it. Results (tables T3, T6, T7 of that run): S0 planted recall 9 of 12 (75.0 percent, Wilson 47 to 91), decoys answered 0 of 6; S1 answered on 5 of 24 units and recovered 1 of 12 mechanisms (8.3 percent, Wilson 1 to 35), Fisher one-sided S0 greater than S1, p = 0.0014; B4 single note recovered 3 of 12, Fisher S0 greater than B4, p = 0.0196 in this run (p = 0.11 in the sealed run). Reading: on this corpus the generator does not produce the planted mechanism from one note plus domain cues; it needs the second note. This answers the strongest objection to H1 in the adversarial review (section 2(a) of research/06). It does not repair the paraphrase leakage in the four B4 recoveries, which remains a corpus flaw. S1 is not preregistered; it becomes hypothesis H4 in v0.2.

X4, critic ablation (run 2026-09-13_x4_critic_sonnet; the sealed run's 81 non-NONE generations re-judged by a Sonnet-class critic, claude-sonnet-5, prompt unchanged). The larger critic killed 3 of the 8 correct recoveries (S0-0002, S0-0006, S0-0011), against the Haiku critic's 3 (S0-0005, S0-0006, S0-0011), and it kept the single decoy answer (S0-0016) that the Haiku critic had killed. Its overall kill rate was 18 of 81 against the Haiku critic's 45 percent on the oracle set. Reading: the recall loss is not a model-size effect. A binary rubric of "restates a claim, generic, uncheckable" removes correct cross-note recoveries at either size, and the larger model is more lenient on the decoy. In v0.2 two critics run in parallel and are compared in table T8; in the execution-verified track the critic is replaced by a test.

X3, cross-domain-near sampler (run 2026-09-13_x3_cross_domain_near; same corpus, prompts and models; measured list-price cost $29.51; a first attempt was lost to a spent subscription usage window and discarded, and the pipeline now aborts a stage above 20 percent errors). Arm B7 samples card pairs from the nearest distance band (Q1) whose notes belong to different domains, the policy the post-hoc analysis suggested. Against B1 (random) and B3 (banded) at 100 draws each: B7 drew 2 planted card pairs (2 distinct bridges) against a hypergeometric expectation of 0.61 (p = 0.126); B1 drew 1, B3 drew 0. The generator answered on 31 of 100 B7 pairs, 16 of 100 B1 pairs and 9 of 100 B3 pairs; within B3 the far band and the top 5 percent drew NONE on all 50 units, replicating the sealed run's gradient. B7 recovered 1 of its 2 planted pairs. Reading: the direction matches the post-hoc finding and the effect is not significant at this size, which is the power problem the review identified (section 2(f) of research/06). v0.2 tests B7 against B1 at 300 draws each (H2 in PREREGISTRATION-v0.2.md).

6.2 v0.2: oblique bridges, two writer families, two critics (preregistered)

Every number below is read from results/public/2026-09-14_micro_v0_2/, regenerated by make reproduce from the committed run experiments/runs/public/2026-09-14_micro_v0_2/ without API access. The protocol is PREREGISTRATION-v0.2.md, sealed at commit ad88f4314b2875330b194d6b435d6b7a118f261c (tag v0.2.0-prereg) before any note of the corpus existed.

Corpus. 96 notes, six domains, 16 per domain: 24 bridges as specified, 12 decoy pairs (9 shape-matched, 3 homonym) and 24 fillers. Half the notes were written by claude-haiku-4-5-20251001 and half by Qwen 2.5 7B Instruct Q4_K_M run locally (hf.co/bartowski/Qwen2.5-7B-Instruct-GGUF:Q4_K_M), assigned by id parity. Every bridge note passed a 6-gram check and was judged by a paraphrase-leak judge with the gold in context; eight bridges (br03, br08, br10, br12, br15, br16, br18, br24; two Haiku-written, six Qwen-written) had a note that ended LEAK after the allowed attempts and were removed from the answer key under the preregistered exclusion rule, leaving 16 planted bridges, 10 written by family A and 6 by family B. Family B writes about half as much under the identical prompt (median 107 words against 176), and 10 of its notes stayed under the 100-word floor. Measured build cost about $7.0 at list price for the Haiku and judge calls; family B was local. Datasheet: data/synth/v0.2/datasheet.md.

Deviations. Five are logged in PREREGISTRATION-v0.2.md: the family-B model was pulled from Hugging Face under a different tag because the Ollama registry's blob host was unreachable; the length floor was lowered from 150 to 100 words for both families after the small model failed the floor systematically; failing notes were given ten attempts instead of five; unjudged short bridge notes were judged after the build and the exclusion applied; and the outcome entry records the rule's result next to the experimenter's reading. None changed a hypothesis, a test or a threshold.

T1. Corpus. Filled by make reproduce (T1.md).

Notes Cards Planted bridges Decoy pairs Cross-note card pairs Notes failing a builder check after the allowed attempts
96 483 (5.03 per note) 16 12 115,345 13

T2. Sampler enrichment of planted card pairs. Filled by make reproduce (T2.md). Base rate: 0.36 percent of the 115,345 cross-note card pairs are planted.

Arm Pairs drawn Planted included Distinct bridges Hypergeometric expectation P(at least observed)
B1 random 300 2 2 1.09 0.297
B7 near, cross-domain 300 2 2 1.09 0.297

T3. Oracle set S0: recall and specificity. Filled by make reproduce (T3.md). Mean cosine between generated connection and gold on planted pairs: 0.547.

Pair type N NONE rate (Wilson 95%) Non-NONE rate Survived critic and dup gate MATCH to gold
Planted 16 43.8% [23, 67] 56.2% [33, 77] 31.2% [14, 56] 6 of 16, 37.5% [18, 61]
Decoy 12 83.3% [55, 95] 16.7% [5, 45] 0.0% [0, 24] n/a
Random non-planted 36 91.7% [78, 97] 5.6% [2, 18] 0.0% [0, 10] n/a

Fisher exact (6 of 16 planted recovered vs 0 of 12 decoy false positives after the primary critic), one-sided: p = 0.0213.

T4. Per-arm NONE rate, kill rate, survivors, recovery and measured cost. Filled by make reproduce (T4.md). 728 units; two generator calls errored (units whose partner note had no cards, see 7.1).

Arm Units NONE % (Wilson 95%) Critic kill % of non-NONE Dup-gate kills Survivors Bridges reachable Bridges recovered Recall USD USD per recovered bridge
S0 64 78.1 [67, 86] 61.5 [36, 82] 0 5 16 6 37.5% [18, 61] 6.26 1.04
S1 32 90.6 [76, 97] 50.0 [9, 91] 0 1 16 0 0.0% [0, 19] 2.73 n/a
B1 300 84.3 [80, 88] 21.3 [12, 35] 0 37 2 0 0.0% [0, 66] 26.93 n/a
B7 300 80.7 [76, 85] 36.2 [25, 49] 0 37 2 0 0.0% [0, 66] 27.60 n/a
B4 32 3.1 [1, 16] 22.6 [11, 40] 0 24 16 8 50.0% [28, 72] 3.92 0.49

Exploratory finds (non-planted survivors): 75 in total, released as results/public/2026-09-14_micro_v0_2/exploratory_finds.jsonl, not counted as hits.

T5. Permutation null over S0. Filled by make reproduce (T5.md).

Statistic Observed Mean under null Permutation p (10,000 shuffles)
Survivors among planted S0 units 5 1.256 0.0002

T6. Single-note reflection (B4) versus the two-note oracle arm. Filled by make reproduce (T6.md).

Arm Bridges tested Mechanisms recovered Rate (Wilson 95%) Fisher p (S0 greater than B4)
S0 two notes 16 6 37.5% [18, 61] 0.857
B4 single note 16 8 50.0% [28, 72] n/a

T7. Partner-domain filler control (preregistered as H4). Filled by make reproduce (T7.md).

Arm Bridges tested Mechanisms recovered Rate (Wilson 95%) Fisher p (S0 greater than S1)
S0 bridge note plus true partner 16 6 37.5% [18, 61] 0.00884
S1 bridge note plus partner-domain filler 16 0 0.0% [0, 19] n/a
B4 bridge note alone 16 8 50.0% [28, 72] n/a

T8. Recall after critic, per critic, on the same generations, before the duplicate gate. Filled by make reproduce (T8.md).

Critic Model id Correct recoveries Surviving Killed Decoy answers Decoy surviving Kill rate on random
haiku (primary) claude-haiku-4-5-20251001 6 3 3 2 0 100.0% [34, 100]
sonnet claude-sonnet-5 6 5 1 2 1 100.0% [34, 100]

Preregistered decision rule (DECISION.md, verbatim).

  • H1 (planted recall 6/16 vs decoy false positives 0/12, Fisher one-sided): p = 0.0213 -> PASS
  • H2a (B7 planted card pairs 2 in 300 draws, expected 1.09, hypergeometric): p = 0.2973; arm-label permutation vs B1: gap = +0.0000, p = 0.6929 -> FAIL
  • H2b (non-NONE rate B7 58/300 vs B1 47/300, Fisher one-sided): p = 0.1413 -> FAIL
  • H3 (S0 two-note recall vs B4 one-note recall, Fisher one-sided): p = 0.8574 -> FAIL
  • H4 (S0 recall 6/16 vs S1 partner-domain-filler recall 0/16, Fisher one-sided): p = 0.0088 -> PASS
  • H5 (recall by writer family, estimation only, no pass/fail): A: 3/10 = 30.0% [11, 60]; B: 3/6 = 50.0% [19, 81]

Decision: SIGNAL (signal requires H1 and H4 both passing; H2, H3, H5 are reported and do not enter the rule).

Reading. The rule says SIGNAL, and it passed for a reason it did not test. The adversarial review of the raw outputs (research/07-adversarial-review-v0.2-results-2026-09-14.md) supports six statements.

(a) The eight single-note recoveries are principles read off one note's ingredient set. Units B4-0000, B4-0007, B4-0008, B4-0009, B4-0015, B4-0016, B4-0019 and B4-0024 recover card testing, relative-schedule drift, a device step (twice), regression to the mean, receipt lag against expense cadence, an unmonitored absence, and a check interval longer than the event. In each, the note lists every ingredient for its side and never states the mechanism, so the leak judge answered CLEAN and the one_side_test in the specification, written for a naive reader, holds. The generator is not a naive reader. The match judge is not lenient: it rejected S0-0000 (br01), whose generated mechanism is the gold mechanism, so oracle recall is 7 of 16 on a human reading and 6 of 16 by the preregistered judge.

(b) H4 is not a recombination test. S1 asks the model whether a bridge note and an unrelated filler connect; 29 of 32 answers were NONE, and the two answers (S1-0000, S1-0016) were honest attempts to connect the wrong pair. NONE is the correct answer to the question S1 asks, and it says nothing about whether the second note was needed. Five bridges the two-note oracle marked NONE (br05, br06, br11, br13, br20; units S0-0003, S0-0004, S0-0007, S0-0008, S0-0012) were recovered by the single-note arm: given both notes and permission to abstain, the model abstained; given one note under the reflection prompt, it produced the gold. The pair prompt's "most pairs are unrelated" framing suppresses answers the model has.

(c) Two bridges behaved as designed, and the pipeline missed both. Of the seven planted NONEs (five Haiku-written, two Qwen-written), br09 (sequential identifiers exposing one operator) and br17 (expectation-lag dropout at the third cycle) are oblique on both sides and recoverable only from both notes. Neither was recovered by any arm. The difficulty the design was meant to create exists in exactly two of sixteen bridges, and the generator did not clear it.

(d) The sampler effect is real, small and undetectable here. All 16 planted pairs again have their closest card pair in the nearest distance band. The cross-domain part of that band holds 20,163 card pairs, 200 of them planted (0.99 percent against a base of 0.36 percent): enrichment 2.73x, with an expectation of 3.0 planted pairs in 300 draws against 1.1 at the base rate; observed 2. Ranking all 3,526 cross-domain note pairs by minimum card distance places the planted pairs at ranks from 6 to 1748; precision at 300 is 5 of 300 (3.7x the base). Detecting a lift of this size at this base rate needs on the order of 1,500 draws per arm. The measured quantity to report is the pool enrichment; the arm-level test is dropped.

(e) No critic configuration keeps recall without passing a decoy. The Haiku critic killed S0-0011, S0-0014 and S0-0015 (three correct recoveries) and both decoy answers; the Sonnet critic killed S0-0014 and kept decoy S0-0024. "Kill if either kills" keeps 3 of 6 correct and 0 decoys; "kill if both kill" keeps 5 of 6 and 1 decoy. The rubric's "generic" and "restates" categories remove correct answers by construction on a corpus whose gold mechanisms are principles.

(f) Shared priors are not supported as the explanation, and the test is confounded. Recall on Qwen-written bridges (3 of 6; single-note 3 of 6) was at least as high as on Haiku-written ones (3 of 10; single-note 5 of 10). The exclusion removed 6 of 12 Qwen bridges, leaving the ones a 7B model could write without stating the mechanism, and Qwen notes are shorter, which concentrates cues.

X5, single-note prompt over every note (exploratory, not preregistered; run 2026-09-14_x5_single_note_all; measured list-price cost $15.05). The single-note prompt already permits NONE. To test whether it measures knowledge or compliance, it was run over every note in the v0.2 corpus with the same generator; four notes the card extractor had refused (synthetic health notes read as personal medical information) had no cards and were skipped, leaving 92 units.

Note kind Units Answered NONE
Filler 23 23 0
Decoy note 21 21 0
Excluded-bridge note 16 16 0
Planted-bridge note 32 31 1

Planted mechanisms recovered from one note: 6 of 16 (br01, br05, br06, br13, br14, br22), 37.5 percent [18, 61] (T4 of that run). The prompt answers on everything; its NONE permission is inert. The single-note arm therefore measures the model's willingness to state an implication from any note, and the S1-versus-B4 contrast in the sealed run is a prompt-contract artifact, not evidence about recombination.

Model and prompt provenance, v0.2. Run id 2026-09-14_micro_v0_2, executed 2026-09-14. Generator snapshot claude-sonnet-5; card, primary critic and recovery judge claude-haiku-4-5-20251001; second critic claude-sonnet-5. Note writers: claude-haiku-4-5-20251001 (family A) and ollama/hf.co/bartowski/Qwen2.5-7B-Instruct-GGUF:Q4_K_M@a693a5ed1336 (family B). Prompt hashes in the run's metadata.yaml. Measured list-price cost $85.04 for the run. Seal commit ad88f4314b2875330b194d6b435d6b7a118f261c, tag v0.2.0-prereg. Corpus sha256 f8a6e9e335772ff44fd86ff0245933ce73d9c5354ef5164c476cb93cdd1ef44a. OpenTimestamps proofs for both preregistrations and answer keys attest in Bitcoin blocks 966837 (v0.1) and 966878 (v0.2); docs/verify-seal.md gives the recipe.

6.3 v0.3: bridges that need both sides (preregistered)

Corpus and gate. data/synth/v0.3/: 96 notes, 24 specified bridges of six fact types (identifier 6, date 5, number 4, vendor 4, location 3, sequence 2), 12 shape-matched decoys, 24 fillers; family A notes by claude-haiku-4-5-20251001, family B by a local Qwen 2.5 14B Instruct. Every bridge note passed the 6-gram check and the paraphrase-leak judge. The one-side gate of ADR-014 then ran on every bridge side: the generator reads one note under prompts/generate_single_strict.md, and any answer is judged against the gold by two votes with a tie failing closed. Four passes were run. Flags: br02 (pass 1), br06 and br15 (pass 2), br11 (pass 3), br09 and br11 (pass 4); all flags except br11's second were ties. Side-A specifications of br02, br06, br15 and br11 were rewritten once each, as the rule allows; br11 was recovered again with 2 of 2 votes and dropped; br09, first flagged on the final pass, was dropped under a four-pass cap logged as deviation 4. Accepted answer key: 22 bridges, 10 family A and 12 family B. Measured list-price build cost about $18.00 (family B ran locally). Datasheet: data/synth/v0.3/datasheet.md.

Design differences from v0.2. Arms S0 (22 planted, 12 decoys, 36 random), S1 (partner-domain filler, 44 units), B4-strict (the strict single-note prompt over 44 bridge notes and 24 fillers; the filler answer rate is the preregistered validity precondition, bar 10 percent), B1 random 300 card pairs, B7 near-in-embedding, far-in-domain 300 card pairs; two critics; the match judge with two votes and a tie counting as NO_MATCH. Signal iff the precondition, H1 and H3-strict all pass. PREREGISTRATION-v0.3.md, seal commit b700d92, tag v0.3.0-prereg.

Results (run 2026-09-14_micro_v0_3; every value from results/public/2026-09-14_micro_v0_3/; each table is regenerated by make reproduce).

Table Result
T1 corpus 96 notes, 521 cards (5.43 per note), 22 planted bridges, 12 decoys, 0 leakage failures, 134,289 cross-note card pairs
T2 sampler enrichment base rate 0.47 percent; B1 drew 1 planted card pair in 300, B7 drew 3 (expected 1.396), hypergeometric p = 0.165; arm-label permutation gap +0.0067, p = 0.32
T3 oracle set planted: NONE 22.7 percent, answered 77.3 percent, survived 72.7 percent; decoy: NONE 75.0 percent, answered 25.0 percent, survived 8.3 percent (1 of 12); random: NONE 94.4 percent, answered 5.6 percent, survived 2.8 percent; recall on planted 8 of 22, 36.4 percent [20, 57]; mean cosine to gold 0.566
T4 per arm S0 70 units, NONE 68.6 percent, 8 recovered; S1 44 units, NONE 88.6 percent, 0 recovered; B4 68 units, NONE 63.2 percent, 2 of 22 recovered; B1 300, NONE 87.7 percent; B7 300, NONE 88.0 percent; 47 exploratory survivors; measured cost $87.86
T5 permutation null 16 planted survivors against 5.684 expected, p = 0.0001
T6 two notes vs one note S0 8 of 22 (36.4 percent) vs B4-strict 2 of 22 (9.1 percent [3, 28]), Fisher p = 0.034
T7 partner-domain filler S1 0 of 22, Fisher S0 greater than S1 p = 0.0018
T8 critics Haiku: 8 correct recoveries, 8 surviving, 0 killed; 3 decoy answers, 1 surviving. Sonnet: 8, 7, 1 killed; 3 decoy answers, 1 surviving. Both let the same decoy answer through
T9 filler abstention bridge notes: 44 units, 13 answered (29.5 percent [18, 44]); fillers: 24 units, 12 answered (50.0 percent [31, 69])

DECISION.md, verbatim:

Validity precondition (B4-strict answered on 12/24 filler notes, rate 50.0% [31, 69], must be below 10%): FAIL. H1 (planted recall 8/22 vs decoy false positives 1/12, Fisher one-sided): p = 0.0826, FAIL. H2a (B7 planted card pairs 3 in 300 draws, expected 1.396, hypergeometric): p = 0.1651; arm-label permutation vs B1: gap = +0.0067, p = 0.3203, FAIL. H2b (non-NONE rate B7 36/300 vs B1 37/300, Fisher one-sided): p = 0.5985, FAIL. H3-strict (S0 two-note recall 8/22 vs B4-strict one-note recall 2/22 on bridge notes, Fisher one-sided): p = 0.0344, PASS (not interpretable: the validity precondition failed). H4 (partner-domain filler, estimation only, not a recombination test): S0 recall 8/22 = 36.4% [20, 57] vs S1 recall 0/22 = 0.0% [0, 15]. H5 (recall by writer family, estimation only, no pass/fail): A: 6/10 = 60.0% [31, 83]; B: 2/12 = 16.7% [5, 45]. Decision: NULL (signal requires the validity precondition, H1 and H3-strict all passing; H2, H4, H5 are reported and do not enter the rule).

Deviations. Five entries in PREREGISTRATION-v0.3.md: the first build pass used five attempts per note before the resume pass gave every failing note its remaining attempts up to ten, and seven short bridge notes were judged by the leak judge after the build; the four rewrite-once spec edits, with the rewritten answer key's hash recorded next to the sealed one; the gate's stochasticity across passes; the four-pass cap that dropped br09 without a rewrite; and the outcome entry, which also records one reporting fix: the stats code had counted only answered notes as "bridges reachable" for the single-note arm, showing 2 of 12 (16.7 percent) where the right denominator is every planted bridge the arm showed the generator, 2 of 22 (9.1 percent). The fix changed T4, T6 and T7 and the H3-strict line (p = 0.034 after, p = 0.21 before); the decision is NULL under both because the precondition failed.

Reading, from the adversarial review (research/08). (a) The precondition indicts itself. The twelve filler answers (B4-0044, -0045, -0046, -0049, -0053, -0054, -0058, -0060, -0063, -0064, -0066, -0067) are, on inspection, ten genuine local inferences from specifics the fillers contain (an arithmetic gap, two contradictory deadlines, an unverified fix), one generic implication and one restatement, and no named principle. The prompt did what it asks; the fillers were dense working notes with numbers and dates, and half of them carry a checkable inference. A filler answer rate therefore measures filler richness, not abstention, and no prompt would pass the 10 percent bar on this corpus without refusing genuine specifics. (b) H1's misses. Of the 14 planted units not counted as recovered, 5 were NONE (br04, br08, br10 from family B; br17, br23 from family A) and 9 were answered and judged NO_MATCH. Three of the nine are judge false negatives a human accepts (S0-0002 br03, whose output is the gold with a tie vote; S0-0015 br18, the same mechanism with 0 of 2; S0-0004 br05, a tie); three are partial and the judge is defensible; three are real misses, and two of those (br02, br06) are self-inflicted: the rewrite-once edits to their side A removed the ingredient the gold depends on. A human-read recall is about 11 of 22 (Fisher about 0.02 against 1 of 12); the preregistered number stands at 8 of 22. The surviving decoy, S0-0027 (dc06), is advice rather than a connection between the two notes; the critic rubric has no category for that. (c) H5, the writer-family effect. Family A 6 of 10, family B 2 of 12. Family B bridge notes are 35 percent shorter (medians 108.5 against 166.5 words), their cards are terser, and among answered family-B pairs the outputs land farther from gold (mean cosine 0.49 against 0.65). The note-quality reading survives; the two dropped bridges were both family A, which inflates the A rate without explaining the B rate; shared priors cannot be separated from length in this design, and a same-length control is the experiment. (d) H3-strict: the gate worked. Two-note recall did not fall across versions (6 of 16 in v0.2, 8 of 22 here) while one-note recall did (8 of 16 to 2 of 22), so the gap between two notes and one note moved from minus 12 to plus 27 points. The confound is the prompt: v0.2 used the inert single prompt, v0.3 the strict one; the same strict prompt on the v0.2 corpus is the clean comparison (X6). Both single-note recoveries (B4-0003, B4-0027) are half-mechanisms that the gate's judge had rejected on the same side hours earlier. (e) The gate is stochastic. Six flags in 192 side-gates across four passes (3.1 percent per side and pass), five of them ties; the flagged set changed almost completely from pass to pass, and four of the five tie-flags cleared on the next pass without any change to the flagged note. Fail-closed on a tie is right for a certification, but it should trigger a re-vote, not a rewrite; and a rewrite must be checked for gold derivability, which the builder does not yet do. (f) Critics. The Haiku critic killed 0 of 8 correct recoveries and the Sonnet critic 1, against 3 of 8 and 3 of 6 in v0.1 and v0.2. Nothing changed in the critic; fact-type outputs cite identifiers, dates and numbers, so the "generic" and "restates" rules do not fire on them. Both critics passed the dc06 advice; the rubric needs a "cites both notes" rule.

X6, the strict prompt over the v0.2 corpus (run 2026-09-14_x6_strict_prompt_v02, measured list-price cost $9.93). The same strict single-note prompt used in v0.3, applied to the v0.2 corpus's 32 bridge notes and 23 filler notes with a two-vote match judge. One-note recovery: 6 of 16 mechanisms (v0.2's inert prompt: 8 of 16; v0.2's two-note oracle: 6 of 16). Fillers answered: 15 of 23 (65.2 percent, Wilson 45 to 81); bridge notes answered: 14 of 32. Reading: with the corpus held fixed, the strict prompt lowers one-note recovery only slightly and does not abstain on fillers; the fall to 2 of 22 in v0.3 is the fact-type corpus and the gate, not the prompt.

X7, five-vote re-judging of the v0.3 answers (run 2026-09-14_x7_rejudge_votes5; generations, critics and duplicate gate reused verbatim from the sealed run; only the match stage re-run with five votes, majority rule). Of 36 answered planted units, two verdicts change, S0-0002 (br03) and S0-0004 (br05), both NO_MATCH to MATCH, the two units the review read as judge false negatives. Under five votes the oracle recall is 10 of 22 (45.5 percent, Wilson 27 to 65), H1 would give p = 0.030 and H3-strict p = 0.008. This is exploratory: the sealed run used two votes, the validity precondition still fails, and the decision stays NULL. It quantifies the judge's contribution to the miss: two of the fourteen non-recoveries are judge disagreements.

Model and prompt provenance, v0.3. Run id 2026-09-14_micro_v0_3, executed 2026-09-14. Generator snapshot claude-sonnet-5; cards, primary critic and match judge claude-haiku-4-5-20251001; second critic claude-sonnet-5. Note writers: claude-haiku-4-5-20251001 (family A) and ollama/hf.co/bartowski/Qwen2.5-14B-Instruct-GGUF:Q4_K_M@1d373ff5643f (family B). Gate generator claude-sonnet-5, gate judge claude-haiku-4-5-20251001. Prompt hashes in the run's metadata.yaml. Measured list-price cost $87.86 for the run, about $18.00 for the corpus build. Seal commit b700d927924be0f324cac6c257256d49c99265fb, tag v0.3.0-prereg, OpenTimestamps proofs anchored in Bitcoin block 966979. Corpus sha256 d044765e9dff (full value in data/synth/v0.3/manifest.json).


7. Limitations and Future Work

7.1 Limitations of v0.1, v0.2 and v0.3

Synthetic ground truth. Recovery of planted structure is not evidence of real-world novelty or usefulness. A pipeline can pass this test and still produce nothing an owner would act on.

Shared priors between corpus author, bridge author and generator. All three involved a language model of the same family. Sealing the bridges before any note existed, the leakage check and the decoys reduce but do not remove the risk that recovery reflects shared associations rather than reading the notes. A second bridge set written by an independent human is the planned control.

One corpus, one generator, one judge. No generalization is claimed. The judge is a language model, used only for grounded entailment with the answer in context.

No human rating. The blind owner-scoring protocol is specified and deferred.

Scaffold variance. Ríos-García et al. attribute 1.5 percent of variance to scaffolds. The arms in this paper are all scaffold ablations. A base-model axis crossed with samplers, with a variance decomposition, is the obvious next experiment; if the sampler effect is small relative to the model effect, that will be reported as the finding.

Leakage by paraphrase. The 6-gram check guarantees no six-word run of a gold sentence appears in a note; it does not stop the note writer from stating the mechanism in other words. Four of the twelve bridges leak this way (Section 6.1), which inflates single-note recovery and weakens H3. v0.2 adds a paraphrase-leak judge: single-note reflection with the gold in context must not rate the note as already containing it.

The critic is a liability as configured. It killed 3 of 8 correct recoveries and 1 of 1 decoy answer. A critic that trades three true recoveries for one false positive is not a filter worth having at these rates; alternatives (a stronger critic model, a two-vote rule, dropping the "generic" reason) are cheap to test on the committed outputs and are the first v0.2 ablation.

H2 was underpowered by construction. With 12 planted note pairs among 52,887 cross-note card pairs (0.65 percent), 100 draws expect 0.65 planted hits under random sampling. No sampler could show enrichment at that base rate unless the effect were enormous, and the post-hoc shows the effect points the other way. H2 in this form is not a fair test of distance-forcing and is replaced in v0.2.

Cost. No cost number is stated that was not read from a log.

Principle-type bridges are recoverable from one note (v0.2). Eight of the sixteen v0.2 mechanisms are general principles, and each bridge note carries the full ingredient set for its side, so the generator reads the principle off one note (Section 6.2, reading (a)). The obliqueness gate (per-side forbidden phrases, a one-side test written for a naive reader, a paraphrase-leak judge) screened for stated mechanisms and passed implied ones. Only br09 and br17 were oblique to the generator itself, and both were missed.

The single-note prompt's NONE permission is inert. Over every note in the v0.2 corpus it answered on 23 of 23 fillers, 21 of 21 decoy notes and 31 of 32 planted-bridge notes (X5). Any comparison between a NONE-permitting pair prompt and this prompt compares prompt contracts, not knowledge. H3 and H4 as preregistered inherit this flaw.

The writer-family exclusion is family-dependent. The leak rule removed 6 of 12 Qwen-written bridges and 2 of 12 Haiku-written ones, and Qwen notes are half as long. H5 is reported as an estimate and cannot separate priors from length and selection.

Card-extractor refusals. Four synthetic health notes were refused by the card extractor as personal medical information, leaving them with no concept cards; in the sealed v0.2 run this produced units with empty claim lists and two errored generator calls. The pipeline now records refusals, retries once and excludes zero-card notes from note-level arms; the sealed run is reported as executed.

The v0.3 validity precondition was mis-specified. It asked the strict single-note prompt to answer on fewer than 10 percent of fillers, but the fillers were written as dense working notes and half of them carry a genuine local inference; the twelve answers it gave are what the prompt asks for. The bar measured filler richness, not abstention. The property actually needed is that one note alone does not yield the gold, which the gate already certifies per bridge.

Judge noise is now the dominant error in both directions. Three planted outputs a human accepts were judged NO_MATCH (S0-0002, S0-0004, S0-0015), and the two single-note "recoveries" (B4-0003, B4-0027) were half-mechanisms the gate's judge had rejected on the same side hours earlier. With two votes and ties failing closed, the same judge on the same prompt returns opposite verdicts on identical content. Every v0.3 number is bounded by this more than by the generator.

The rewrite-once rule has a gold-derivability hole. Rewriting side A of br02 and br06 to defeat a gate flag removed the ingredient their gold depends on, so those bridges became unrecoverable by construction and show up as misses. A rewrite must preserve, and the builder must assert, that every gold ingredient still appears on some side.

The writer-family effect is confounded with length. Family B bridge notes are 35 percent shorter and their cards terser; the 60 against 17 percent recovery gap tracks that, and shared priors cannot be separated from it without a same-length control.

Plan for v0.4. Keep the fact-type corpus design and the gate. Replace the filler-rate precondition with two checks that target the actual failure: zero filler answers may match any gold, and single-note recovery on bridge notes must stay below one third of two-note recall. If a filler answer rate is kept at all, write the fillers without numbers, dates or conflicts and set the bar from a pilot. Use three or five match-judge votes, with a tie triggering a re-vote rather than a rewrite. Add a gold-derivability check to every rewrite. Add a "cites both notes" rule to the critic so advice that connects nothing is killed. Run a same-length control for the writer-family effect (family B notes regenerated with a higher floor, or family A notes truncated). Report the sampler as pool enrichment or at the size that can detect a 3x lift. In parallel, the next headline experiment moves to real code repositories with execution as the critic, where a failing test replaces the LLM's opinion and a pinned commit replaces the synthetic answer key; the prior-art sweep for that direction is research/05-prior-art-execution-verified-code-2026-09-13.md.

7.2 Track C: owner-blind scoring over a real private corpus

Pool survivors, shuffle, strip arm labels, present in batches of 20. The owner marks each item KEEP (would act on or write down) and KNOWN (already had this thought) as binary decisions. Thirty items are repeated for intra-rater kappa. The key file's hash is committed before scoring and the file is revealed after. Fisher exact and permutation as in v0.1. The engagement-versus-insight warning from Strömel et al. is the reason this is blind and binary, never a rating of enjoyment.

7.3 Track A: retrospective temporal validation

An old-cutoff open-weight generator with a six-month margin (HindSight's discipline) over a time-frozen public corpus, with hit detection by retrieval over post-cutoff literature and a modern-model entailment matcher, reported as lift over a matched null with a lead-time distribution and a future-neighbourhood rate. Inherited methodology, cited, not claimed.

7.4 Track B: prospective public registry

Dreams over live corpora, published with git timestamps and OpenTimestamps anchoring, with two outcome classes: independently discovered later, and adopted from the registry.

7.5 Open measurements nobody has made

Usable yield (owner KEEP rate, with "not in corpus" enforced by retrieval) as a function of embedding distance, including the far tail. Tail diversity of dreams across nights with and without writeback, the model-collapse curve. Concept-card fidelity rates. A false-positive rate for a corpus-scale ideation system.


8. Conclusion

Generation is cheap and selection is the wall, and every previous attempt at a day-dreaming loop stopped at that wall without measuring it. Three preregistered runs measured it on synthetic corpora with planted answers, and the three outcomes are NULL, SIGNAL by rule with recombination not claimed, and NULL. What moved across versions is the thing the corpus design was changed to move: single-note recovery fell from 50 percent of planted mechanisms in v0.2 to 9 percent in v0.3 while two-note recovery held (38 to 36 percent), so the bridges of v0.3 are the first that a capable model cannot read from one side. What did not move is the sampler: distance-forced, near-in-embedding and random pairing drew planted pairs at chance three times. And what now bounds every number is the instrument rather than the loop: a strict abstention prompt that answers on half of dense filler notes, and a match judge that returns opposite verdicts on identical content. By the v0.1 rule: no signal. By the v0.2 rule: signal, for a reason the rule did not test. By the v0.3 rule: no signal, because the rule's own precondition measured the wrong thing. We report all three verbatim, next to the readings that qualify them. The corpora, answer keys, prompts, raw outputs, tables, deviations and reviews are released so that the next run, with a judge that votes more than twice and a precondition that tests abstention rather than filler richness, can be judged against the same discipline; the next headline experiment moves that discipline to real code with execution as the critic.

The system and the protocol are open source under MIT (code) and CC-BY-4.0 (paper and data).


References

All references were checked against their primary source (arXiv abstract page, publisher page, or live URL) on 2026-09-13.

Baek, J., et al. (2024). ResearchAgent: Iterative research idea generation over scientific literature with large language models. arXiv:2404.07738.

Benedek, M. and Neubauer, A. C. (2013). Revisiting Mednick's model on creativity-related differences in associative hierarchies: Evidence for a common path to uncommon thought. Journal of Creative Behavior, 47(4), 273–289. doi:10.1002/jocb.35.

Ke, Y., George, K., Pandya, K., et al. (2025). BioDisco: Multi-agent hypothesis generation with dual-mode evidence, iterative feedback and temporal evaluation. arXiv:2508.01285.

Boden, M. A. (1990). The Creative Mind: Myths and Mechanisms. Weidenfeld and Nicolson.

Brodt, S., Pöhlchen, D., Täumer, E., Gais, S. and Schönauer, M. (2018). Incubation, not sleep, aids problem-solving. Sleep, 41(10), zsy155. doi:10.1093/sleep/zsy155.

Campbell, D. T. (1960). Blind variation and selective retention in creative thought as in other knowledge processes. Psychological Review, 67(6), 380–400.

Chen, Z., Zhao, K., Fu, J., et al. (2026). Ideation Arena: Evaluating LLM generated research ideas with battle-style human expert assessment. arXiv:2608.29696.

Chen, Z., Zhao, Y. and Cohan, A. (2026). Measuring the gap between human and LLM research ideas. arXiv:2607.01233.

Tao, J., Wang, Y., Liu, X., et al. (2026). Continuous Knowledge Metabolism: Generating scientific hypotheses from evolving literature. arXiv:2604.12243. (Cited as CKM.)

Deperrois, N., Petrovici, M. A., Senn, W. and Jordan, J. (2022). Learning cortical representations through perturbed and adversarial dreaming. eLife, 11:e76384.

Ellis, K., et al. (2021). DreamCoder: Bootstrapping inductive program synthesis with wake-sleep library learning. PLDI 2021.

Tang, Q., Sun, T., Hu, X., et al. (2026). FARS: A fully automated research system deployed at scale. arXiv:2606.31651.

Artiles, A. H., Weiss, M., Brinkmann, L., et al. (2026). The alien space of science: Sampling coherent but cognitively unavailable research directions. arXiv:2603.01092.

Goedecke, S. (2025). Practical notes on getting LLMs to generate new ideas. seangoedecke.com/idea-mill; code at github.com/sgoedecke/idea-mill.

Gottweis, J., Weng, W.-H., Daryin, A., et al. (2025). Accelerating scientific discovery with Co-Scientist (v1 title: Towards an AI co-scientist). arXiv:2502.18864.

Gwern (2025). LLM daydreaming. gwern.net/ai-daydreaming. Last modified 2025-07-14.

Jiang, B. (2026). HindSight: Evaluating LLM-generated research ideas via future impact. arXiv:2603.15164.

Hoel, E. (2021). The overfitted brain: Dreams evolved to assist generalization. Patterns, 2(5), 100244.

Si, C., Hashimoto, T. and Yang, D. (2025). The ideation-execution gap: Execution outcomes of LLM-generated versus human research ideas. arXiv:2506.20803.

Lin, K., et al. (2025). Sleep-time compute: Beyond inference scaling at test-time. arXiv:2504.13171.

Liu, Q. E., Dubova, M., Conklin, H., et al. (2026). Assessing the effect of cross-domain mapping on creativity in humans and large language models. arXiv:2603.19087.

Luo, X., et al. (2025). Large language models surpass human experts in predicting neuroscience results. Nature Human Behaviour, 9(2), 305–315.

Mednick, S. (1962). The associative basis of the creative process. Psychological Review, 69(3), 220–232.

Mo, Y., Zheng, T., Gao, Y., et al. (2026). AgentIdeaBench: Benchmarking scientific ideation in the agent era. arXiv:2609.07611.

Nusrat, H. and Nusrat, O. (2025). When AI does science: Evaluating the autonomous AI scientist KOSMOS in radiation biology. arXiv:2511.13825.

Orwig, W., Luchini, S. A., Beaty, R. and Schacter, D. L. (2025). A 'sweet spot' for creative ideation: Non-linear associations between semantic distance and creativity. Cognitive Computational Neuroscience (CCN) 2025, abstract.

Park, J. S., et al. (2023). Generative agents: Interactive simulacra of human behavior. arXiv:2304.03442.

Peeperkorn, M., et al. (2024). Is temperature the creativity parameter of large language models? arXiv:2405.00492.

Penadés, J. R., et al. (2025). AI mirrors experimental science to uncover a mechanism of gene transfer crucial to bacterial evolution. Cell, 188. doi:10.1016/j.cell.2025.08.032. Published online 2025-09-09.

Ríos-García, M., Alampara, N., Gupta, C., et al. (2026). AI scientists produce results without reasoning scientifically. arXiv:2604.18805.

Romera-Paredes, B., et al. (2024). Mathematical discoveries from program search with large language models. Nature, 625, 468–475.

Novikov, A., Vũ, N., Eisenberger, M., et al. (2025). AlphaEvolve: A coding agent for scientific and algorithmic discovery. arXiv:2506.13131.

Schönauer, M., Brodt, S., Pöhlchen, D., Breßmer, A., Danek, A. H. and Gais, S. (2018). Sleep does not promote solving classical insight problems and magic tricks. Frontiers in Human Neuroscience, 12:72. doi:10.3389/fnhum.2018.00072.

Schopf, T. and Färber, M. (2026). Is this idea novel? An automated benchmark for judgment of research ideas (RINoBench). arXiv:2603.10303.

Schwartz, M. D. (2026). Vibe physics: The AI grad student. Anthropic Research, anthropic.com/research/vibe-physics.

Shen, A., Druckmann, S. and Zou, J. (2026). Unlocking LLM creativity in science through analogical reasoning. arXiv:2605.11258.

Shinde, S. S. (2026). SCM: Sleep-consolidated memory with algorithmic forgetting for large language models. arXiv:2604.20943.

Shumailov, I., et al. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755–759.

Si, C., Yang, D. and Hashimoto, T. (2024). Can LLMs generate novel research ideas? arXiv:2409.04109. ICLR 2025.

Sinhahajari, S., Majumder, N. and Poria, S. (2026). On the limits of LLM-as-judge for scientific novelty assessment. arXiv:2606.12071.

Strömel, K. R., Henry, S., Johansson, T., Niess, J. and Woźniak, P. W. (2024). Narrating fitness: Leveraging large language models for reflective fitness tracker data interpretation. CHI 2024. doi:10.1145/3613904.3642032.

Zhu, S., Zhang, H., Chi, J. D., et al. (2026). Causal stories from sensor traces: Auditing epistemic overreach in LLM-generated personal sensing explanations. arXiv:2605.08590.

Tao, T., et al. Erdős problems AI contributions ledger. github.com/teorth/erdosproblems wiki.

Uzzi, B., Mukherjee, S., Stringer, M. and Jones, B. (2013). Atypical combinations and scientific impact. Science, 342(6157), 468–472.

Wagner, U., et al. (2004). Sleep inspires insight. Nature, 427, 352–355.

Wainrib, G., Bodinier, B., Dakhli, H., et al. (2026). Can AI scientist agents learn from lab-in-the-loop feedback? Evidence from iterative perturbation discovery. arXiv:2603.26177.

Wang, Q., et al. (2023). SciMON: Scientific inspiration machines optimized for novelty. arXiv:2305.14259.

Yang, Z., Liu, W., Gao, B., et al. (2025). MOOSE-Chem: Large language models for rediscovering unseen chemistry scientific hypotheses. ICLR 2025. arXiv:2410.07076.

Ye, C., Liu, Y., Wang, Y., et al. (2026). Auto-Dreamer: Learning offline memory consolidation for language agents. arXiv:2605.20616.

Zahn, O., Evans, J. and Eagleman, D. (2026). Discovery by dreaming: Cross-domain recombination in artificial memory. arXiv:2607.16256.

Łukasiak, Z. (zby) (2025). Reinventing daydreaming machines. zzbbyy.substack.com/p/reinventing-daydreaming-machines, 2025-10-13; code at github.com/zby/DayDreamingDayDreaming.

Zhou, Y., et al. (2024). Shared imagination: LLMs hallucinate alike. arXiv:2407.16604.

Kim, Y., Chang, Y., Pham, C. M., et al. (2026). Argument collapse: LLMs flatten long-form public debate. arXiv:2606.01736.

Behrouz, A., Hashemi, F., Javanmard, A. and Mirrokni, V. (2026). Language models need sleep: Learning to self-modify and consolidate memories. arXiv:2606.03979. (Google Research; the 'sleep stage' with an RL-generated synthetic curriculum.)

Ding, T., Nannapaneni, A., Liu, B., et al. (2026). Always-on agents: A survey of persistent memory, state, and governance in LLM agents. arXiv:2606.30306.

Cheung, V. (2026). Dreaming is not a bug: A Jung-inspired dream layer for multi-agent LLM companions. arXiv:2601.06115.