Skip to content

Adversarial sanity review of the daydreamd approach

Produced 2026-09-13 by a hostile-but-fair reviewer agent (Claude Fable 5.1, fork of the founding session) against BRIEF v3 (2026-07-25). Reproduced verbatim. Line references point at research/BRIEF-v3-2026-07-25.md.


Audit of the daydreamd approach (brief v3, 2026-07-25) before build/publish.

1. The ten hardest objections

O1. "Your ground truth is one person's opinion." (Track C, N=1) The primary ground truth is "Julian's morning verdicts" (brief L313-315). A reviewer will call this an anecdote with a daemon attached. Brief's answer: "the user IS the world expert on their own vault" (L343). That justifies why the judge is competent, not why the result generalizes. Neutralizer: the blind permutation test converts N=1 opinion into a within-subject discrimination experiment with a known chance level. Report d' / AUC, not "liked it". Then recruit 5-10 opt-in vault owners for the same protocol before any generalization claim. Until then every claim is "in a single-owner case study".

O2. "Permuted controls are trivially distinguishable for the wrong reason." Permuted pairs (concept A from note X + concept B from note Y, re-shuffled) may yield dreams that are incoherent, so the human spots the control by fluency, not by insight. That's the CHI '24 lesson (engagement ≠ insight, L80-84) turned against you. Brief does not address this. Neutralizer: controls must be generated by the same pipeline from random (non-distance-forced) pairs, so both arms are fluent LLM prose; the only difference is the sampling policy. That is B1 (random pairing, L224), not a label permutation. Keep Wainrib-style label permutation (L124-129) as the statistical null, but the stimulus control must be B1. The brief conflates the two (see §2).

O3. "Novelty-vs-world is unmeasurable over a private corpus." Brief pre-empts this (L293-296: hybrid, private generation / public checking). Fine in principle, but "recall probe" (L170) is exactly the LLM-opinion judge the brief itself declares dead (L310-316). Neutralizer: novelty-vs-world must be retrieval over a real index (Semantic Scholar / arXiv / web search API), reported as nearest-neighbour similarity with a threshold, plus Tao's 1(a)-2(d) classification (L420-427) per hit. Publish the retrieval traces.

O4. "Scaffold engineering is worth 1.5% of variance; you are shipping a scaffold." (Ríos-García, L99-105) Brief accepts this and pivots to "value from CORPUS and EVALUATION" (L104-105). But the experiment matrix B0-B6 is all scaffold ablations. Neutralizer: add a base-model axis (Haiku / Sonnet / Opus / a 70B open model) crossed with 2 samplers, and report variance decomposition. If the sampler effect is <5% of model effect, say so; that is still a publishable finding ("sampler doesn't matter, cost does").

O5. "You have no false-positive rate and neither does anyone else, so you will look the same." Brief flags this (L130, L446). Neutralizer: pre-register the FP definition: a "hit" the retrieval check later shows already exists in the corpus or in the literature. Report precision at the judge's threshold with a Wilson interval.

O6. "Concept extraction is the failure point (Goedecke hand-wrote YAML)." (L164-166) "Cached + incremental" is not evidence extraction works. Neutralizer: sample 50 concept cards, have a human (Julian) mark each as faithful / distorted / hallucinated, report the rate. If distortion >10%, the whole downstream is contaminated. This is a 20-minute job and must be in the first paper.

O7. "Incubation beats you." (B4, L387-390) Correct baseline, but the brief specifies it as "cron re-ask the question tomorrow", which presumes a question. Dreams have no question. Neutralizer: B4 = "given note X alone, propose one implication you have not written down" (single-note reflection, Generative-Agents style, = B2). If B2/B4 match B3 on usable-yield, the recombination claim dies and you report that.

O8. "Model collapse via writeback." (L377-383) Brief has mitigations. A reviewer will ask for a measurement: tail-diversity of dreams across nights with and without writeback. Neutralizer: track type-token / embedding-dispersion of dreams per night over ≥14 nights, writeback on vs off. Cheap, and no one has that curve.

O9. "Contamination: the generator has seen the corpus owner's public writing." Julian's notes overlap his X/LinkedIn posts, which may be in training data. Neutralizer: for the retro track, old-cutoff generator + 6-month margin (L119-121). For the private track, state the limitation plainly and run the Co-Scientist-style "unwritten answer" test (L93-97) with 3-5 questions Julian knows the answer to but never wrote.

O10. "Why a daemon? A cron job with a script is the same thing." The "always-on" membership claim (L288-291) is weak; the brief admits "claim membership, not invention". Neutralizer: do not lead with the daemon. Lead with the evaluation protocol (blind permutation + held-out slice + retrieval-grounded novelty). The daemon is packaging.

2. Internal contradictions and over-claims

  • L124-129 vs L467-478. Null design is defined as label permutation (permute dream→note assignments, Wainrib style), but the weekend gate says night one emits "permuted/scrambled controls (same concepts, wrong pairings)", which is a stimulus control. These are different experiments with different questions. Fix: stimulus control = B1 random pairs, statistical null = label permutation. State both.
  • L136-149 vs L156-161 ("Architecture (locked)"). Decision 3 corrects the fertile-zone claim to novelty × feasibility, yet the "locked" pipeline still says "banded distance sampler… default to the measured peak band" and L411-418 replaces the scalar band with Uzzi anchor+remote. Three sampler designs coexist. Pick B6 as default, B3 as ablation.
  • L200-204 cost claim. "200-dream night ≈ 1M tokens ≈ $2" and "$1.40" are both quoted as headline. 200 dreams × (card pair ~600 tok + output ~300) ≈ 180k tokens for generation, not 1M; the 1M presumably includes critic + judge passes but that's not shown. Do not publish a cost number until it is measured from logs.
  • L169-173 judge rubric vs L310-316. Rubric axes include "novel vs THE WORLD (recall probe)", an LLM-opinion check, one section after declaring LLM novelty opinions dead. Replace with retrieval.
  • L449-452. "Binary keep/discard, never Likert" is right, but the primary curve is "usable-idea yield (novelty × feasibility)" which requires two graded scores. Define: yield = P(keep) per band, where keep requires both "not in corpus" (retrieval) and "owner would act on it" (binary human). Drop the product-of-scores language.
  • L221-222 vs L500. Success criterion "≥1 'huh' dream" is a self-report, exactly what L84 forbids. Replace with the discrimination statistic in §3.
  • L44 "HindSight ρ=−0.29 is the evidence the join matters." It is evidence LLM-judged novelty is misleading; it says nothing about the triad or private corpora. Cite it for what it is.
  • L107-111 lead-time distribution applies only to Track A; the micro-experiment cannot produce it. Don't mention lead time in v0.1.

3. Minimum viable experiment (one laptop, one evening)

Corpus. 40-60 markdown notes from ~/.claude/projects/*/memory (frozen copy, hashed, never modified). Strip credentials files. Record SHA of the snapshot.

Concept cards. One Haiku-class pass per note, schema:

{id, source_note, claim (≤30 words, atomic), entities[], why_it_matters (≤20 words), confidence: high|med|low}
Cap 3-6 cards per note → ~200 cards. Human check: 50 random cards marked faithful/distorted/hallucinated. Report the rate (O6).

Embeddings. Local all-MiniLM-L6-v2 (or bge-small) on claim. Cosine distance matrix over cards. Exclude pairs from the same note.

Pair sampling. Distance bands by quantile of the pair-distance distribution: Q1 (near), Q2-3 (mid), Q4 (far), top-5% (tail). Arms: - B3 banded: 25 pairs per band × 4 bands = 100 - B6 anchor+remote: 25 pairs where anchor = card from the densest cluster, remote = from Q4 - B1 random: 100 pairs uniform - B4 single-note reflection: 50 single cards, prompt "state one implication not written in this note" Total ≈ 275 generation calls.

Generator prompt (identical across arms), Sonnet or Haiku, temperature default:

You will see two claims from one person's notes. Most pairs are unrelated. If no genuine, non-obvious connection exists, output exactly NONE. Otherwise output JSON: {connection: ≤40 words, mechanism: why it holds, testable_implication: one concrete thing the note owner could check or do this week, needs: which claim supplies what}. Do not restate either claim. Do not invent facts absent from the claims.

Log NONE rate per arm and band (free honesty metric, L153-154).

Critic (cheap, Haiku): binary. Kill if: restates a claim; implication not checkable; connection is generic ("both involve users"). Report kill rate per arm.

Corpus-novelty gate (retrieval, not opinion): embed each survivor's connection; if cosine to any existing card or note chunk > 0.85, mark "already in corpus" and kill. Report rate.

Blind human scoring. Pool all survivors, shuffle, strip arm labels, present in batches of 20 as markdown checkboxes. Julian marks each: (a) KEEP: I would act on this or write it down (binary), (b) KNOWN: I already had this thought (binary). Scoring done before unblinding; the key file is sealed by SHA before scoring begins (commit the hash, reveal the file after).

Statistics. - Primary: keep-rate B3/B6 vs B1 (Fisher exact, one-sided). With 100 vs 100 survivors-before-critic, you can detect 30% vs 12% at α=.05 with ~80% power. If critic kills 85%, you have ~15 per arm and the test is underpowered; so either score pre-critic output (recommended for v0.1: score everything) or raise pairs to 200 per arm. - Label-permutation null: shuffle arm labels 10,000×, recompute the keep-rate gap; report the permutation p. - Secondary: keep-rate by distance band (Cochran-Armitage trend), NONE-rate by band, "KNOWN" rate by arm, B4 keep-rate vs B3. - "Signal" = permutation p < .05 on B3-or-B6 vs B1 AND B3-or-B6 > B4. Anything else is reported as null, in full.

Tables in paper.md. - T1: corpus stats (notes, cards, card fidelity rate). - T2: per arm: pairs, NONE%, critic-kill%, corpus-duplicate%, survivors, human KEEP%, KNOWN%. - T3: KEEP% by distance band (B3), with counts. - T4: permutation null result + Fisher p. - T5: cost: measured tokens and $ per arm and per KEEP.

Cost. 275 generations × ~800 tok in / 250 out ≈ 0.3M tokens; critic 275 × 500 ≈ 0.14M; cards 60 × 1.5k ≈ 0.1M. ≈0.55M tokens, under $2 on Sonnet, ~$0.30 on Haiku; $0 marginal on claude -p via Max. Human scoring: ~275 items × 20s ≈ 90 min. Do the scoring in two sittings to test intra-rater consistency on 30 repeated items (report κ).

4. Honest abstract sentences after the micro-experiment only

Allowed: - "In a single-owner case study over N private notes, distance-forced pairing produced a higher blind keep-rate than random pairing (x% vs y%, permutation p = z), and [did / did not] exceed single-note reflection." - "Owner-blind keep decisions, not LLM novelty scores, are the ground truth; LLM roles are limited to generation, a binary coherence critic, and retrieval-based duplicate detection." - "We release the protocol (permutation null, sealed key, retrieval-grounded corpus-novelty check) so the result can be replicated on any private corpus." - "The generator declined to connect x% of pairs; the rate rose with embedding distance."

Not allowed until Tracks A/B/multi-user exist: - "novel ideas", "discovers", "anticipates research", any lead-time number, any cost-per-idea headline not read from logs, "first", "benchmark", any generalization beyond one corpus, anything about sleep biology.

5. Naming

daydreamd collides semantically with every neighbour (Gwern's day-dreaming loop, Vault Daydream, zby/DayDreamingDayDreaming, Google "Dreaming", Anthropic Dreams API). A survey author will file you under "daydreaming implementations", not as a category. Search hits: 23 GitHub repos already contain "daydream" in the name. Recommend renaming before the first commit.

Checked (PyPI/npm/GitHub name search): - oneird — free on PyPI, npm; 1 GitHub repo. Oneiric + daemon, greppable, one word. Top pick. - hypnagog — free everywhere, 41 loosely matching repos (mostly "hypnagogia" art projects). Evocative, slightly hard to spell. - nocturnd — free everywhere. Neutral, no "dream" collision. - somnid — free everywhere. Short, blank slate. - dreamgap — free everywhere, but keeps the "dream" root; use only if you want the collision as SEO.

Avoid remd, dreamt, remcycle (npm taken).


Disposition (2026-09-13, Julian): name kept as daydreamd (free on PyPI, npm, crates.io; only an idle 2013 GitHub username). All other recommendations adopted in PREREGISTRATION.md and paper/paper.md.