Skip to content

BRIEF — daydreamd

Founding document v3. Written 2026-07-24/25 (Fable 5 @ xhigh: planning pass + 5 parallel prior-art sweeps + publication strategy). Research lineage lives in ../DREAMING.md. This file is the source of truth for future sessions — start at ⭐ SETTLED POSITION.**

HOW TO PICK THIS UP COLD: read ⭐ SETTLED POSITION (the answer) → PUBLICATION STRATEGY (where it goes) → Build plan (the weekend blind-test gate) → Open decisions (what's still Julian's call). Everything between is the audit trail, including claims that were made and then killed by the sweeps — kept deliberately so nobody re-litigates them.

STATUS 2026-07-25: research + positioning COMPLETE. Nothing built yet. Awaiting Julian's go on the weekend MVP. No code, no repo, no name chosen.

Identity rule: STANDALONE research/product project. NOTHING to do with the social-media "Loops" content machine — no shared branding, no "loop" naming. Julian's channels are launch channels only. Working name: daydreamd (final name = Julian's call).


One-liner

A local-first daemon that dreams over your accumulated knowledge while you sleep — deliberately colliding far-apart concepts from your own corpus, filtering the collisions through a strict critic, and serving the survivors with your morning coffee.

Why this, why now

  • Gwern proposed the day-dreaming loop (2025-07-14). Sean Goedecke shipped a semi-manual prototype ONE day later, called it "pretty half-assed" — and it still produced "a few genuinely novel ideas." A full year on, nobody has built the rigorous version. Validated, dormant, citable gap.
  • Agent-memory ecosystem (Letta, Mem0, A-Mem) all consolidate memory; none generate from it. PKM/Obsidian crowd feels their notes are dead weight. Two starving audiences.
  • The deeper prize: nobody can measure idea-novelty credibly. Whoever ships the benchmark gets cited by everyone who measures against it (highest-citation act in ML, per BRAND research).

⭐ SETTLED POSITION (after 5 prior-art sweeps, 2026-07-24) — READ THIS FIRST

Everything below this section is the audit trail of how we got here, including claims that were made and then killed. This is the answer.

THE CONTRIBUTION, in one sentence: join the interestingness triad (novelty × usefulness × surprise) to a temporal ground-truth split, over a private corpus, as a persistent daemon. Two independent 2025-26 teams built the triad (InterFeat: novelty+utility+plausibility; SerenQA 2511.12472: relevance+novelty+surprise) — neither validated it against later independent discovery. Four teams built time-splits (HindSight, CKM, Jansen/AI2, MOOSE family) — none applied it to a triad or to a personal corpus. Nobody has joined them. That join is the paper. HindSight's ρ=−0.29 is the evidence the join matters.

THREE CLAIMS THAT SURVIVED ALL FIVE SWEEPS: 1. Private corpus — verified empty from BOTH directions: zero LBD papers touch personal corpora; arXiv search "personal informatics" AND "large language model" returns 0 hits. 2. Frozen old-cutoff generator run BACKWARDS — MOOSE-Chem freezes the model and moves the test set forward; we freeze the model and let the world move. The inversion is unpublished. 3. Triad × temporal validation (above).

TWO CLAIMS TO DROP — they were true a year ago and are not now: - ~~"no usefulness critic exists"~~ → InterFeat (2026) built one. - ~~"time-splitting is unexplored in the LLM era"~~ → HindSight, CKM, Jansen/AI2, MOOSE family, all 2024-26.

⚠️ THE SOBERING FACT — the personal-corpus gap is empty for a REASON. The best-powered test of LLM narratives over personal data (Strömel et al., CHI '24, N=273, real donated fitness data, validated instruments) moved engagement, attention and reward — and got a FLAT NULL on the Insight subscale: F(2,270)=0.43, p=.64. PhysioLLM's generated-insight arm did not beat a plain data summary. And an audit of 14,922 generated explanations over personal sensor data (2605.08590) found LLMs "routinely attribute anomalous days to causes without sufficient support," with richer context NOT reducing overreach. → Power the study properly, and run PhysioLLM's placebo arm (a condition with NO personal data) — almost nobody does.

⚠️ BUT — why that null does NOT test our hypothesis (have this answer ready; it is the reviewer's and the skeptic's first shot): 1. Wrong substrate. Their corpus was 7 days of STEP-COUNT data — numbers about walking. Ours is years of accumulated written thinking. A week of pedometer data contains nothing to discover; there is no latent conceptual structure to connect. 2. Wrong operation. They SUMMARIZED one dataset ("you were most active Tuesday"). We RECOMBINE two distant concepts. Different mechanism entirely — and the PAD ablation shows the recombination step is precisely where the measurable value lives (removing memory-mixing costs 4.4/18 accuracy points). 3. Wrong expertise. Their participants had no expertise in their own step counts — there is nothing to be expert about. Julian IS the world authority on his own accumulated thinking, which is the entire reason the judging problem is tractable for us and impossible for everyone else. 4. No filter at all. They generated narrative and displayed it. No critic, no rejection, no NONE-permission. What DOES transfer (the real warning, keep it): their participants found the output engaging, enjoyable, well-written and rewarding — every measure moved EXCEPT insight. Enjoyment is not insight, and self-report cannot tell them apart. Our morning.md could easily be a pleasure to read and completely useless. That is exactly why the weekend gate is a BLIND test against permuted controls, not "did you like it?".

THE MOST PORTABLE EVAL DESIGN IN THE LITERATURE — steal it: BrainBench (Luo et al., Nature Human Behaviour 9(2):305–315, 2025): present two versions of an abstract — the real result vs an altered one — and ask which is real. LLMs beat human neuroscience experts; LLM confidence is calibrated. Authors state it "is not neuroscience specific and is transferable." Personal- corpus adaptation: hold out the last N months of the corpus, generate from the earlier slice, and score against the held-out slice. That exact experiment has never been run by anyone.

BEST VALIDATION DESIGN ANYWHERE (aspire to it): Co-Scientist × Penadés (Cell 188:6654, 2025) — the lab posed a question they had already answered experimentally but never published anywhere, and the top-ranked hypothesis matched their confirmed mechanism. Uncontaminable by construction: the answer existed in no corpus on Earth. Our analogue: ask a collaborator (or Julian) to pose a question whose answer is in their head but not written down anywhere.

THE SCAFFOLD WARNING — the deepest cut in all five sweeps: Ríos-García et al. (2604.18805), >25,000 agent runs across 8 domains: variance decomposition attributes 41.4% to the base model and 1.5% to the agent scaffold. Evidence is ignored in 68% of traces; belief revision after refutation occurs in 26% — persisting even when agents are handed near-complete correct reasoning trajectories. Verbatim: "Outcome-based evaluation cannot detect these failures, and scaffold engineering alone cannot repair them." → Our value must come from the CORPUS (private, unexploited) and the EVALUATION (grounded, null-tested), never from clever orchestration.

METRIC SHAPE: report lead-time DISTRIBUTION, not a point estimate — CKM (2604.12243) got a mean of 404 days (median 399, range 66–757) between hypothesis and matching paper. Also report future-neighbourhood rate (pArticleMap: exact gold recovery only 10.8%, but 61.0% landed in the right forward-looking region) — credit for being directionally right when exact matching is too strict.

The four load-bearing design decisions (2026-07-24 planning pass)

Decisions 1 and 2 were CORRECTED the same day by the prior-art sweep — see POSITIONING below. 1 is inherited methodology, NOT our contribution. Do not claim either as novel. 1. Old-cutoff open-weight generator for the retro-eval — INHERITED, cite it. Generator = a model whose cutoff predates the corpus window; frontier model still does hit-MATCHING only (entailment over two in-context texts, no memory). Already claimed: MOOSE-Chem (ICLR 2025, 51 post-Jan-2024 chemistry papers, pre-2024-cutoff models) and HindSight (2603.15164, Mar 2026, Llama-3.3-70B + June-2023 cutoff + 6-month safety margin — adopt that margin, it's more careful than our version). Co-Scientist's cf-PICI result is stronger still (ground truth was UNPUBLISHED). Present as hygiene we inherit; claiming it invites an easy reviewer kill. Still requires the pluggable-generator interface (wanted anyway). 2. NULL DESIGN — UPGRADED to random-permutation (better than my matched-decoy version). Steal from Wainrib et al. (2603.26177, 800 independently replicated experiments): they show feedback access gives +53.4% discoveries (p=0.003), then permute the hit/miss labels and the gain vanishes — proving the effect came from feedback STRUCTURE, not prompt-induced recall. Our version: permute the dream→note assignments. If the hit rate survives permutation, we never had a signal. Tests the mechanism, not just the outcome; cheap; no reviewer can argue. Keep matched-decoy pairs as a secondary base-rate check (still nobody reports a false-positive rate — and Nusrat & Nusrat, 2511.13825, audited 3 Kosmos hypotheses against random-gene nulls and found 1 of 3 was indistinguishable from noise, caught ONLY by the null model). Related prior art to cite: Alien Space's availability model (2603.01092), BioDisco's Bradley–Terry, OptimusKG's negative-edge control (83.4% of injected FALSE edges correctly got no support). 3. Figure 1 = yield-vs-distance — CORRECTED: the y-axis is novelty × FEASIBILITY, not novelty. ⚠️ The raw-novelty fertile-zone claim is FALSIFIED by the two most direct measurements: Liu/Dubova/Griffiths (2603.19087, Mar 2026 — 140 humans, 2,800 ideas, 7 LLMs, our exact manipulation) finds distance→originality LINEAR and positive (humans β=.10 p<.01; LLMs β=.18 p<.001), no quadratic term; Shen/Druckmann/Zou (2605.11258) gets its best results by MAXIMIZING distance (6.8–7.5/10 vs 2.3–3.4 baseline; >50% novel solutions vs 1.6%). Orwig 2025's "sweet spot" is a plateau, not a peak (distance stops helping; it does not hurt). What survives and is genuinely unclaimed: novelty rises with distance (monotone, robust), feasibility FALLS with distance, so usable-idea yield = novelty × feasibility and should be single-peaked. NOBODY has ever plotted that product against distance. Cheap to run, and it confirms or kills the hypothesis outright. Also untested: the far tail (>7/10 distance), where the distinctive prediction lives. Do NOT cite Mednick's flat-hierarchy mechanism (refuted, Benedek & Neubauer 2013) or Uzzi as "intermediate" (his finding is BIMODAL: conventional core + atypical tail — a different design prescription). 4. Generator gets permission to say NONE. "Find a connection" forces confabulation. Prompt contract: "Most pairs are unrelated. If no genuine connection exists, output NONE." Structured output per dream: {connection, mechanism, testable_implication} — requiring a falsifiable implication is itself a novelty filter (vibes can't produce one). NONE-rate doubles as a free honesty metric per model and per distance band.

Architecture (locked)

Pipeline: ingest → concept cards → embed (local) → banded distance sampler → generator (NONE-permitted, structured) → cheap critic (~85% kill) → deep judge (strict rubric) → morning.md → user verdicts → writeback.

  • Concept cards, not raw chunks. One-time cheap-model pass per document distilling atomic claims + source pointers (tight schema). This is exactly the step Sean identified as broken (hand-wrote facts in YAML because extraction was unreliable). Cached + incremental.
  • Local embeddings (small sentence-transformer on-device). "Your notes never leave your machine" must hold for embedding too, not just storage.
  • Banded sampler: sample pairs across embedding-distance bands (not max-distance). After the fertile-zone experiment, default to the measured peak band.
  • Judge rubric — five axes: novel vs THE CORPUS (retrieval check: if a note already says it, kill) · novel vs THE WORLD (recall probe) · mechanism validity · usefulness TO THIS CORPUS'S OWNER · falsifiability. Precision over recall, hard threshold. An empty morning beats a boring one — a lax judge kills trust in a week. Judge slot = config value (default Opus-class; Fable A/B-able for pennies; must be swappable anyway for ZDR enterprises where Fable 400s).
  • Morning review = the flywheel (prototype already validated by the LinkedIn board pattern): approve/reject-with-reason per dream → per-user LEARNINGS file tunes the judge; endorsed dreams gain sampling weight and become concept cards (compounding); rejections carry reasons. Every dream carries provenance lineage; a dream built on a later-rejected dream gets flagged. Verdicts double as ground truth for the product-usefulness eval — product and paper feed each other.
  • Writeback risk (hallucination compounding): dreams stay marked unverified until endorsed; lineage makes contamination traceable.

The daemon (product spec)

One core, four adapters — the engine never knows where the corpus came from. 1. Claude Code memory (~/.claude/projects/*/memory + transcripts) — dogfood + agent-memory wedge. 2. Obsidian (markdown folder) — biggest audience; output = daily note "what your vault dreamed last night" with [[wikilinks]] to the connected sources. Screenshot-bait by design. Community plugin (TS wrapper over core) = phase-2 adoption channel. 3. Codebase — code-aware chunking (functions, docs, TODOs, commit messages). 4. Zotero — Better-BibTeX / zotero.sqlite; the researcher audience — the one that cites.

Ship shape: local-first CLI (uvx daydreamd), BYO Anthropic key (Julian's own runs: drive claude headless on Max = zero marginal cost), pluggable generator (any API or local model — required by eval decision #1, and the community will want Ollama), built-in launchd/cron scheduler, no telemetry by default. Config: corpus paths, schedule, dream count, distance band, judge model, strictness. Outputs: morning.md + a reviewable verdict surface (start with plain markdown checkboxes; board-style UI later).

Cost model (the answer to Gwern's daydreaming tax): 200-dream night ≈ 1M tokens ≈ $2 (Haiku-class generation + cheap critic; Opus-class judge on the ~15% survivors ≈ $0.60). Batch API halves it; caching cuts further → real-world $1–5/night. "A full night of machine dreaming over your vault: $1.40" = launch-post slide + README line + paper result (cost-per-surviving-idea vs naive all-expensive dreaming, ~20–30x).

Evaluation (the paper)

Track A — retrospective, structurally leakage-free: old-cutoff open-weight generator dreams over a time-frozen corpus (recommended: an arXiv cs slice, dense enough that connections get made within 2–3 yrs; Julian to confirm domain; second domain later for generality). Hit-detection = retrieval over post-cutoff literature + modern-LLM entailment matcher. Report strict (mechanism matches) and loose (pair co-occurs as a topic) hit rates, ALWAYS as lift over the matched null. Pre-register thresholds.

Track B — prospective public registry: dream now over live corpora; publish OPEN with git timestamps + OpenTimestamps anchoring. Two outcome classes, both wins: independently discovered and adopted from the registry (someone reads a dream and builds it = impact, not contamination). Registry doubles as an idea faucet and a recurring content series with a built-in season finale.

Track C — human usefulness: morning-review verdicts (Julian's own + opt-in users) → real-world usefulness rate; blind A/B ratings for arm comparisons.

Experiment matrix (same corpus, critic, budget): B0 raw "be creative" sampling · B1 random pairing (sampler ablation) · B2 reflection-style summarization (Generative Agents baseline) · B3 banded distance-forcing (ours). Metrics: lift over matched null · novelty-vs-corpus · novelty-vs-world · blind human usefulness · NONE-rate · cost per surviving idea. Candidate headline (test it): "cheap generator + good sampler beats big model + random pairing" — inverts the scaling instinct.

PUBLICATION STRATEGY (2026-07-25) — where this goes so others find it

Goal is DISCOVERABILITY + PRIORITY, not prestige. Sequencing: get a result → post the preprint → ship the benchmark → chase venues at leisure. The preprint does the discovery work; the venues do the ratification work; only the benchmark does the compounding work.

1. arXiv FIRST, the moment a result exists — this is the actual discovery layer. Evidence: in all five of our own prior-art sweeps, essentially everything we found (HindSight, CKM, Jansen/AI2, Alien Space, FARS, the whole 2026 frontier) was found ON ARXIV, most never having reached a venue. The next person doing this work will look there. It also plants the priority flag — Google Research is already using "Dreaming" for a related LLM mechanism (2606.03979). Categories: cs.CL + cs.AI + cs.HC (cross-list; cs.HC is how the PKM/personal-informatics audience finds it). This is also the BRAND.md keystone: arXiv author page + Google Scholar profile + a notability reference, seeded in one move.

2. The BENCHMARK is the citation engine, not the paper. Papers get read once; a benchmark gets USED, and every use is a citation. Ship the eval harness (permutation null + held-out-slice test + scoring rubric) as a NAMED, runnable artifact with a leaderboard. Then anyone claiming their idea-generator produces novel ideas has to run against it or explain why not. This is the highest-citation act available to us.

3. TWO papers, TWO venues — the contributions are different and belong in different rooms. - Paper A — method + benchmark (ML venue). Targets: NeurIPS Datasets & Benchmarks track (exists precisely for "we built the thing everyone should measure against"; PlanBench went there) · ACL/EMNLP via ACL Rolling Review — big advantage: ROLLING submission, no annual cliff · ICLR (deadline typically ~September for the following spring). - Paper B — the human study (HCI venue). Targets: CHI (deadline typically ~September) or IMWUT/UbiComp (QUARTERLY deadlines, very flexible). Why this matters more than it sounds: the flat-null study we must answer (Strömel et al., N=273, insight p=.64) was published AT CHI. Publishing our result in the same venue directly answers a specific published finding in the room where it landed — a far stronger citation hook than burying it in an ML conference whose readers never saw the null. - ⚠️ Verify all deadlines before planning around them (my dates are approximate). Best guess as of Jul 2026: NeurIPS 2026 has passed; CHI and ICLR ~September are plausibly catchable this cycle; ACL Rolling Review and IMWUT have no cliff at all.

4. Discovery mechanics beyond venues: - A short, distinctive, greppable NAME. People find work by searching the term; a generic name is invisible. (Also why the name decision is load-bearing, not cosmetic.) - Hugging Face Papers (hf.co/papers/submit) — where practitioners actually browse. - GitHub repo linked from the paper — puts us in code-search paths as well as literature search. - Get picked up by the SURVEYS. This field is producing taxonomies every few months and they scrape arXiv. Being on arXiv early with a clear name means the next survey files us as a CATEGORY — that is how you become the reference point instead of a footnote.

Registry + daemon wait on no venue. Ship them regardless.

POSITIONING — post-prior-art-sweep (2026-07-24). READ BEFORE WRITING ANY CLAIM.

The claim, in one sentence: a persistent daemon accumulating state over a PRIVATE corpus, generating connections via a generator→critic cascade, with novelty checked against PUBLIC literature. Lead with the compound (private corpus × persistence); it is a different KIND of system, not a variant of the published pipelines. Every serious system (SciMON, ResearchAgent, Nova, CoI, MOOSE-Chem, Deep Ideation, BioDisco, Co-Scientist, FARS, Alien Space, SciMuse) runs on PUBLISHED literature. SciMuse personalizes — but from the researcher's published record.

Differentiator audit: (a) temporal holdout = CLAIMED (MOOSE-Chem, BioDisco, HindSight) · (b) old-cutoff models = CLAIMED (same) · (c) base-rate correction = PARTIAL, keep the narrow decoy-pair version · (d) private corpus = UNCLAIMED · (e) persistent background operation = UNCLAIMED for ideation (FARS's 417-hour campaign is a long BATCH run; per the Always-On Agents survey 2606.30306 the defining property is durable state across invocations, not literal continuity — claim membership, not invention).

The reviewer's first objection, have the answer ready: a private corpus destroys the field's only credible eval (you can't do future-paper matching on ideas nobody else can see). Answer = the hybrid: private corpus for generation, public literature for novelty checking and temporal validation. That combination is unclaimed and defensible.

THREE NUMBERS THAT JUSTIFY THE PROJECT (put them in the paper's intro)

  1. HindSight: ρ = −0.29 (p<0.01) — LLM-judged novelty is NEGATIVELY correlated with anticipating real future research; optimizing against LLM-judge scores yields impressive-sounding vacuity.
  2. Si et al. (ICLR 2025): 4,000 generated ideas → ~200 unique (5% survival) — brute-force generation does NOT expand the idea space. Direct empirical support for the sampler thesis: the PAIRING SUBSTRATE is the only remaining lever. Alien Space supplies the mechanism (LLMs "recombine high-density regions of the literature when prompted for novel ideas").
  3. FARS: automated review 5.00 vs 88 human reviewers 3.23 — LLM-judge inflation measured at deployment scale. Plus the Ideation–Execution Gap (2506.20803): AI ideas rated novel pre-execution scored BELOW human ideas post-execution.

FORCED DESIGN CHANGE — the judge cannot be an opinion

The field has discredited exactly what we first specified (an LLM scoring novelty on a rubric). Therefore: novel-vs-corpus and novel-vs-world become RETRIEVAL/ENTAILMENT checks against actual documents, not model opinions; the LLM's role is verifying grounding, not voting on novelty. And Julian's morning verdicts become the PRIMARY ground truth, not a nice-to-have — which conveniently makes the product the eval instrument. Never claim "our ideas score higher on novelty"; that framing is dead on arrival post-Ideation–Execution-Gap.

STEAL LIST (from the sweep)

  • Co-Scientist Meta-review agent — recurring review patterns fed forward into the NEXT round's prompts. Cheapest published self-improvement mechanism; maps exactly onto our LEARNINGS loop.
  • Co-Scientist Proximity agent — similarity graph for dedup/diversity; directly counters Si's 5% collapse.
  • SciMON novelty loop as a HARD GATE — regenerate until dissimilar-enough from retrieved neighbours (not a score, a gate).
  • idea-mill's two-stage blind connection — generate the cross-domain observation WITHOUT seeing the target problem, then apply it. Prevents pattern-matching straight to the answer.
  • ResearchAgent's entity-centric co-occurrence store — closest published pairing substrate.
  • Alien Space's "idea atoms" — the unit of connection (validates concept cards).
  • FARS's auditable-corpus discipline — publish intermediate artifacts, not curated successes.
  • Vault Daydream's recency weighting + pair-history dedup.
  • Graphiti-style structural heuristic — propose bridges between graph communities with no existing path.

COMPETITIVE FACTS

  • Vault Daydream (glebis/claude-skills, 329★): vault pairs → Sonnet generators + Haiku critics → ≥7.0 threshold → writes insight notes back. Run-on-demand, no distance forcing, threshold-only QC, no evaluation. → our claim is "first EVALUATED, distance-controlled" system.
  • zby/DayDreamingDayDreaming (Jul–Oct 2025): piloted pre-cutoff-model temporal novelty testing; stopped with two stated walls — "I didn't implement the search algorithm — I manually selected combinations" and "a domain-agnostic novelty verifier is the fundamental research bottleneck." Those two sentences ARE this project. Cite him.
  • Consensus of both serious builders: generation is easy, VERIFICATION is the wall. Goedecke: you can't scale past what the operator can personally judge → personal corpora are secretly ideal, because the user IS the world expert on their own vault.
  • Anthropic shipped a Dreams API (May 2026) and OpenClaw/Letta/MiMo all ship consolidation → consolidation is COMMODITIZED; build none of it, differentiate 100% on divergent generation. Copy their contract shape (async job, immutable input, new output store, human-review gate).
  • Gwern has been silent — essay unmodified since 2025-07-14, has linked NO implementation. Do not build distribution plans on him.
  • Cautionary tale: Napkin (most connection-forward PKM product) SHUT DOWN. Similarity-based connection surfacing wasn't enough. The dreams must be genuinely good → the judge is everything.
  • Demand receipts: 553★ OpenClaw auto-dream plugin; top Smart Connections review asks for exactly our feature ("I wish it would identify gaps and prompt me").

THE REFRAME (2026-07-24, post-creativity-science sweep) — THIS IS THE PROJECT

Not a better generator. A SELECTION ENGINE for machine-generated ideas, deployed over a private corpus where the owner is the ground-truth expert. Anchor quote — Anthropic's own physicist, doing real QCD research with Claude (Schwartz, "Vibe physics," anthropic.com/research/vibe-physics, Mar 2026): "LLMs are profoundly creative. They simply lack a sense of which paths might be fruitful before walking them." DeepMind concurs structurally ("AI agents are conjecture machines… the binding constraint is the validation bottleneck," Jul 2026). Both serious DDL builders died at the verifier wall. HindSight proved LLM judges anti-correlate with real value. → Generation is commoditized; SELECTION is the frontier, and a private corpus is the one setting where a competent human ground truth is available for free.

CREATIVITY-SCIENCE CORRECTIONS — read before making any claim

  • Sampling > capability is CONTESTED, do not claim it flatly. FunSearch (Nature 2024) is the strong support: deliberately used a FAST WEAK model, ~10⁶ samples, "results not too sensitive to the exact choice of LLM." But AlphaEvolve (2506.13131) reverses its own lab's finding: "performs increasingly better as the underlying LLM improves"; Table 1 goes from "small LLMs, no benefit from larger + 10⁶ samples" to "benefits from SOTA LLMs + 10³ samples" — a ~1000× sample reduction bought by capability. Verbalized Sampling: "more capable models benefit MORE." Chen & Ding 2023: sampling helped everything EXCEPT GPT-4 (opposite interaction sign — unresolved). Defensible restatement: a search/evaluation loop is necessary for verified novelty at ANY capability level; better models convert into drastically cheaper search, not single-shot discovery. Never conflate "sampling strategy" with "temperature" (Peeperkorn 2024 killed temperature).
  • ⚠️ MODEL COLLAPSE is a structural threat to the writeback loop. Shumailov et al. (Nature 2024): recursive self-training destroys distributional TAILS FIRST — and remote associations ARE the tails. Our compounding loop is that loop. Non-negotiable mitigations: (a) segregate generated material, never silently fold into ground truth; (b) anchor every cycle in fresh real inputs; (c) human verification gate. DreamCoder is the existence proof of doing it right — trains on fantasies AND replays of real solved tasks, every fantasy grounded by an actual solver run (3.7%→79.6% on text editing; learned 93% of 60 physics laws from 100–200 real tasks).
  • ⚠️ The sleep biology is weak — stop leaning on it. Wagner 2004 "Sleep inspires insight" LARGELY FAILED TO REPLICATE (Brodt 2018 literally titled "Incubation, not sleep, aids problem-solving"; Schönauer 2018 null across all conditions). → NEW MANDATORY BASELINE: a cron job that simply re-asks the question tomorrow. If we can't beat incubation, the daemon is over-engineered. What survives: Hoel's ARCHITECTURAL argument — a system doing dropout-style noise injection during live operation would fail at its job, therefore "a dedicated offline period is needed." Also solid: sleep abstracts gist while preserving category structure (Liu et al. 2025, Comms Biology); REM lowers the threshold for WEAK/remote associates (Stickgold 1999).
  • ✅ VERIFIED WHITE SPACE: nobody has connected Hoel's overfitted-brain hypothesis to LLMs. Zero mentions in Hoel's paper (predates ChatGPT); exactly two arXiv full-text hits for abs:"overfitted brain", neither about LLMs; Hoel's own AI writing runs the opposite direction. This is our theoretical frame, claimable honestly.
  • Best quantitative support for recombination itself: PAD (Deperrois et al., eLife 2022) — ablating the MIXING of multiple memories in the REM phase costs 4.4 pts (CIFAR-10) and 18 pts (SVHN). The recombination, not merely the offline replay, does measurable work.
  • ⚠️ Most under-appreciated risk (Chen/Zhao/Cohan 2607.01233): LLMs already over-produce "bridge-like opportunities and synthesis methods" — a combination-daemon may amplify the exact mode LLMs are already over-indexed on, while real human diversity lives in other framings. Mitigation: measure output-shape diversity, not just novelty. Also: models hallucinate ALIKE (shared imagination, 2407.16604) → multi-model ensembling is NOT a diversity fix.
  • ⚠️ PRIORITY RISK: Google Research published a "Sleep"/"Dreaming" stage for LLMs (2606.03979, Jun 2026 — RL-generated synthetic curriculum); 2601.06115 proposes an offline high-temperature "dream layer" for agents. The word is being claimed. Move.
  • The bar to beat is higher than it looks: Si et al. showed plain over-generate-and-rank already beat 100+ expert researchers on novelty with no dreaming at all. Our likely win is diversity/anti-convergence and SELECTION quality, not raw novelty.

SAMPLER DESIGN CHANGE — Uzzi's shape, not a scalar distance band

Uzzi (Science 2013, 17.9M papers) is two-dimensional, and it is the better-evidenced policy: high-impact work = conventional CORE + an injection of ATYPICAL (papers high on both are ~2× more likely to land in the top 5% of citations). So the sampler should draw an anchor from dense familiar territory + ONE remote element, rather than pairing two mid-distance concepts. This also sidesteps the falsified scalar-band claim entirely, and it is cheap to A/B against banded pairing (add as arm B6). Also note Kauffman's adjacent possible is a NEAR-distance prior (one step from what exists) — it cuts against far-transfer, so cite it as framing, never as evidence for reaching far.

CLAIM DISCIPLINE — adopt Tao's Erdős taxonomy for every novelty claim we make

Tao et al.'s community ledger (github.com/teorth/erdosproblems wiki) classifies AI contributions: 1(a) AI-independent, no comparable literature · 1(b) AI solution + literature found afterwards · 1(c) AI building on known literature · 1(d) AI+human · 2(a) literature search · 2(b) formalization · 2(c) rewriting · 2(d) computation. Classify every "dream hit" against this before calling it novel. The cautionary case: Oct 2025, GPT-5 was publicized as solving ~10 open Erdős problems; it had located existing literature solutions (Hassabis: "embarrassing"). That is the exact failure mode our retrieval-grounded judge must catch. Tao's own read is the honest frame: AI is "becoming capable enough to pick off the lowest hanging fruit… precisely the category most likely to have been solved in the literature already" — but the capability trend "bodes well for scanning the long tail of underexamined problems." A personal corpus IS a long tail of underexamined material.

C2 RESTATED (use this wording, don't overclaim)

"A generate→evaluate→select loop is necessary for verified novelty at any capability level; better models make search cheaper, not unnecessary." Concede the AlphaEvolve reversal explicitly — a critic will find it otherwise. Cleanest supporting datapoint: Shen 2026, strategy alone moved novel-solution rate 1.6% → >50% at fixed model and temperature.

UPDATED EXPERIMENT MATRIX (supersedes the four-arm version below)

B0 raw "be creative" · B1 random pairing (sampler ablation) · B2 reflection-style summarization (Generative Agents) · B3 banded distance-forcing · B4 incubation control (cron re-ask, no recombination) · B5 over-generate-and-rank (the Si et al. bar) · B6 Uzzi anchor+remote (conventional core + one atypical injection). Primary curve: usable-idea yield (novelty × feasibility) vs distance band, including the >7/10 tail. Secondary: lift over matched decoy null · false-positive rate (nobody reports one) · CROSS-RUN and CROSS-MODEL diversity — mandatory, not optional (Argument Collapse: humans 65.3% unique arguments vs LLMs 3.4%; Shared Imagination: models hallucinate alike, so multi-model ensembling is NOT a diversity fix; Tang & Yang 2026: AI research agents measurably NARROW exploration) · cost per surviving idea. Selection rule: coarse BINARY keep/discard, never 1–10 Likert (expert agreement collapses at fine granularity; LLM judges score 43–53% consistency vs 50% random). Never use an internal novelty score as a fitness function for iteration — only as triage.

HIGHEST-PRIORITY READS BEFORE WRITING THE PAPER

BioDisco 2508.01285 (full text — closest to our eval design) · Alien Space 2603.01092 (determines how much of (c) survives) · HindSight 2603.15164 · Ideation–Execution Gap 2506.20803 · ResearchStudio-Idea 2607.04439 (newest, Jul 2026) · Krenn's pre-LLM semantic-network forecasting.

Order-of-magnitude levers (condensed)

Own the verifiable eval (benchmark) + ship the daemon (dependency: PKM crowd + Letta-class embedding) + one documented real hit with timestamped logs (story) + months-long compounding curve (moat) × distribution already built (Julian's channels). (Corrected: an earlier draft said "Gwern links serious implementations of his proposals." He has linked NONE — essay untouched since 2025-07-24. Do not plan distribution around him.)

Build plan

  • Weekend (go/no-go gate) — TESTS THE SELECTION THESIS, NOT THE GENERATION ONE. Generation is demonstrably not the bottleneck and every corpse in this field died on verification, so the MVP must be instrumented to answer "can we tell good from bad?" — not "can we produce output?". Build: core + Claude-memory adapter over Julian's own memory → first morning.md. THE GATE IS A BLIND TEST: night one emits real dreams mixed with permuted/scrambled controls (same concepts, wrong pairings). Julian scores them without knowing which is which.
  • Can't tell them apart → we learned it in a weekend for a few dollars, and the honest negative result is itself publishable.
  • Real ones consistently stand out → we have signal nobody else has, and the way we proved it is the thing that makes it publishable. This design exists because of the CHI '24 null: engagement moved, insight didn't, and self-report could not distinguish them. Never gate on "did you like it?"
  • Week 2: fertile-zone experiment on a public corpus → Figure 1 exists. Obsidian adapter + uvx packaging.
  • Weeks 3–4: retro-eval (old-cutoff generator, arXiv slice, matched null).
  • Week 4+: registry launch + build-in-public (X, Show HN, r/ObsidianMD, Obsidian forum, hf.co/papers, Gwern served on a plate). Codebase + Zotero adapters follow.
  • Then: writeup → arXiv (+ workshop if catchable).
  • Model split: research/design on Fable 5 xhigh; build on Opus 5 (settings default).

Risks (honest)

  1. Hit-rate quality is THE gamble — whether survivors are actually good is the experiment. For: idea-mill got real novel ideas from a worse setup; we fix exactly what its author said was broken (fact extraction → concept cards; filtering → cascade + rubric). Downside case: benchmark + daemon + honest negative-result paper — smaller but real win.
  2. Critic quality ceiling — judge gets the expensive model + strictest prompt; empty > boring.
  3. Confabulation — NONE-permission + falsifiable-implication requirement + case-control null.
  4. Hallucination compounding in writeback — provenance lineage + unverified-until-endorsed.
  5. Privacy optics — local-first, BYO key, local embeddings, no telemetry. Non-negotiable.
  6. Fable-specific: never prompt to extract model reasoning (reasoning_extraction refusal category — dreams are OUTPUT text, always); handle stop_reason: "refusal"; Fable needs 30-day retention (no ZDR) → judge must be swappable.

Success criteria

  • Weekend: ≥1 "huh, I hadn't connected those" dream.
  • Launch month: strangers running it on their own vaults; first unsolicited screenshot.
  • 6 months: eval published; registry checkable; a framework or plugin ecosystem embeds it.

Open decisions (Julian's)

  • Final name (working: daydreamd).
  • Retro-eval corpus domain (rec: arXiv cs slice). Everything else above is decided-by-default; future sessions should not re-litigate it.