ORBITALNET
research/LOG.md

Lab log

What was built in each phase, which tests pass, what surprised us, and what we distrust about every result.

research/ LOG

Running log of what was built, what passed, what surprised us, and decisions made on ambiguous points. Append-only, one section per phase.

Phase 3 — real-data counterfactual replay: infrastructure (2026-10-11)

Status: pipeline built and tested; NO results yet. Raw data download is incomplete (see "Open items"), so nothing has been run at scale.

What it is

Real ZTF Bright Transient Survey events (spectroscopic class = truth, real r-band light curve, real RA/Dec, real date) replayed against the Phase 2 Starlink fleet propagated from Space-Track TLEs archived at the time. Only the planners' measurements are simulated (GP mean through the real points + limiting-magnitude instrument noise). Same 5 strategies as Phase 2, unchanged.

Decisions

  • Hypothesis set: modeled = SN Ia / SN II / SN Ibc; withheld (real h0) = TDE, SLSN-I/II, SN Ibn. User chose "both studies": this SN study is the main one; a small case study on the manuscript classes is still to do.
  • Class models are empirical templates (median peak-normalized curve) built ONLY from events first detected before 2020-01-01, i.e. before the fleet existed. Default chosen by me when the user said "start the next phase" without picking between this and sncosmo; flagged in chat.
  • Split: template = pre-2020; post-2020 events halved by date into validation (older) and test (newer). Test is locked in code until PREREGISTRATION.md exists.
  • Selection bug fixed: first selection took the EARLIEST 300 events per class, so almost all supernovae predated the fleet. Now per modeled class: up to 300 pre-2020 + 300 post-2020, evenly spaced in time; all withheld.
  • No truth extrapolation: candidate times after an event's last real detection are infeasible for every strategy.
  • No stale orbits: a cell is infeasible if the nearest archived TLE containing that satellite is > 7 days from the observation time.
  • h0 GP: signal amplitude marginalized over the same amplitude grid as the modeled classes (real fluxes span ~3 dex; a fixed var_f would make h0 evidence track brightness instead of shape). Lengthscale 10 d, fixed, tunable on validation only.
  • Candidate times 0.5-60 d (16 log-spaced), budget 5, 8 satellites, limiting mags 19/20/21: same structure as Phase 2, timescale widened for SNe.
  • Sun exclusion still NOT modeled (now possible since dates are real; future work).

Validation tuning protocol (written BEFORE any tuning run was seen)

First validation run (seed 1, no stretch, GP lengthscale 10 d; committed 180660f): P's AUROC(P(h0): withheld vs modeled) = 0.582, and P decided h0 on 33% of modeled SNe. Diagnosis: a single fixed median template per class is too rigid for real SN diversity, so the flexible h0 GP wins by default.

Change: optional per-event template stretch s (f = A * template(t/s)), log-uniform over 7 values in [0.5, 2.0]; marginalized like amplitude. B2/P decay terms now use posterior-mean flux (identical to before for amplitude-only models; verified max diff 3.6e-15 on Phase 2 beliefs).

Grid (validation split, seed 1 only): stretch in {off, [0.5, 2.0]} x GP lengthscale in {5, 10, 20} days. Selection rule, fixed now: pick the config with the highest P AUROC; ties (within 0.01) broken by lower P confident-wrong rate on MODELED events. The chosen config is then frozen in PREREGISTRATION.md before the test split is run once. All configs' results are kept, not just the winner.

Losses (reported by the loader, never silent)

  • ~35% of downloaded objects have < 8 r-band detections and are dropped.
  • 30 TLE weeks Feb-Aug 2021 came back EMPTY although the fleet was in orbit (likely Space-Track throttling). Events there currently get no contact.

Open items

  1. Re-run fetch_bts (~1,100 new light curves for the post-2020 pool).
  2. Re-run fetch_tles for new event weeks, after deleting the empty 2021 files.
  3. Validation run, then PREREGISTRATION.md, then test run.
  4. Viewer replay tab; manuscript-class case study.

Phase 2 — single-planner experiment with real orbits (2026-10-01)

Two passes: own design, then realigned to the manuscript

Phase 2 was built twice. The first pass (documented in the earlier parts of this entry's git history) used toy exponential/plateau/oscillating classes, a plain-EIG utility with no decay/adequacy weighting, and 5 scenarios varying fleet size — all explicitly labeled "this project's own design" because no manuscript existed yet. Mid-session the user supplied Hypothesis-Discriminating Observation Campaigns for Distributed Space Observatories.pdf. Its Part IV defines H1-H4, the utility function (eq. 3-4), the exact class list, and the scenario list precisely — all different from the first pass. The first pass's validation results (null control passed on all 5 of the old scenarios) were discarded rather than used for a preregistration, since preregistering against a superseded design would have been pointless. Phase 2 was then rebuilt to match the manuscript. This section documents the SECOND (current) build.

Built

  • transients/models.py: replaced the toy P2 classes with the manuscript's 4 modeled (kilonova, shock-cooling/Type IIb early phase, GRB afterglow broken power-law, M-dwarf flare) + 3 withheld (FBOT/AT2018cow-like, TDE, synthetic "nonsense") classes, as physically-motivated analytic approximations — NOT the published templates (sncosmo etc.) the manuscript asks for, since this branch has no network access beyond requirements-research.txt. Documented inline and here; replacing these with real forward models is the single highest-value follow-up. Also added modeled_flux_fn_p2_perturbed(cls, factor) (time-dilates a class's shape by factor) for the model-error scenario.
  • inference/mutual_information.py: added expected_information_gain_batch (Phase 1 already had the scalar version); this was already done before the manuscript arrived but is load-bearing for Phase 2's utility.
  • planners/phase2.py: rewrote the baselines to match the manuscript's definitions — B0 is now a deterministic fixed-cadence dispatcher (was random), B2 is now a "classifier-entropy" disagreement heuristic over point-estimate amplitudes with no Bayesian marginal-likelihood machinery (was best-SNR greedy). B1 (CNP reimplementation) and B3 (closed-world EIG, renamed internally to share a _class_only_eig_batch helper with P) needed no change. P now computes U = I_class + beta*I_Z + gamma*D:
    • I_class: exact, via _class_only_eig_batch (closed-world only, same computation B3 uses).
    • I_Z (adequacy term): exact, by reusing the already-built open-world cell/weight machinery with a 2-group split (all modeled cells vs. the single h0 cell) instead of the old per-class+h0 grouping — this is actually simpler than what was there before.
    • D (decay term): approximated, not exact. The manuscript's I_k(Y_a) is a per-class mutual-information quantity that would need one more EIG quadrature per class per candidate per step — at ~1000 events x 3 scenarios this was not affordable on top of I_class and I_Z. Used a Fisher-information-style SNR² proxy instead: D ≈ Σ_k b_t(h_k) · [SNR_k(t)² − SNR_k(t')²]₊ where t' is the next time the same satellite is visible (or SNR treated as 0 if there is no later opportunity, i.e. maximal urgency). SNR² is the right order for Fisher information of an amplitude-like parameter under Gaussian noise, so this is a defensible proxy, not an arbitrary one, but it is not what eq. 4 literally specifies. Tested in tests/test_phase2_utility.py (zero when nothing is lost by waiting, positive when the class is fading).
    • cost(a): fixed at 1 (uniform). The manuscript's exposure/slew cost model is not implemented.
    • beta=1.0, gamma=0.01 are fixed defaults chosen so I_class/I_Z (nats, O(0-1.4) for K=4) and D (SNR² units, empirically larger) land on a comparable scale — picked for scale-matching before seeing results, not tuned against outcomes. The manuscript's own statistics section says to tune beta/gamma on a validation split; that calibration step is not yet automated here (see "What we distrust" below).
    • Mode switch (manuscript 5.5): implemented as a fallback at the experiment-script level (phase2_single_planner.run_one_event), not inside the planner — once P(h0) exceeds MODE_SWITCH_TAU=0.9 for two consecutive steps, remaining steps for that event use fixed_cadence_planner instead of P's own EIG utility, approximating "maximize coverage" with "take whatever's next." A time_coverage_fraction metric (fraction of the full candidate-time range spanned by chosen measurements) stands in for the manuscript's wavelength-coverage preservation metric, since no wavelength bands are modeled.
  • experiments/scenarios.py: replaced the fleet-size-based S1-S5 with Nominal / Crowding / Model-error (3 scenarios); Communication sweep and Node loss are reserved for Phase 4 (see PLAN.md section 6 for why). Crowding randomly masks a fraction of (time, satellite) cells once per event, shared across strategies. Model error swaps the BELIEF's (not the true event's) flux resolver for the time-dilated perturbed one.
  • experiments/phase2_single_planner.py: widened CANDIDATE_TIMES_DAYS from [0.2, 10] days (12 points) to [0.01, 20] days (16 points) — the manuscript's classes span minutes (flare) to weeks (TDE), which the old grid could not represent. CONFIDENT_THRESHOLD raised to 0.95 (manuscript H1). Added the mode-switch/preservation logic described above.

Tests passed

pytest research/ — 33/33 passed, including:

  • test_phase2_planners.py: updated for the renamed B0/B2 (fixed-cadence picks the deterministic earliest-time/lowest-satellite cell; classifier-entropy matches its own documented disagreement formula independently recomputed in the test).
  • test_phase2_utility.py (new): the perturbed-model resolver equals the true model evaluated at t/factor; _next_visible_index finds the correct next-visible slot per satellite; the decay-term proxy is exactly zero when waiting loses nothing and strictly positive when the class is fading.

What we distrust about this result

  • The decay term is a proxy, not eq. 4's literal quantity (see above) — H2 (irreversibility) should be read as testing "does SOME decay-aware term help fast classes more," not as a direct test of the manuscript's exact formulation.
  • beta/gamma are fixed, not calibrated on a validation split as the manuscript's statistics section asks — a real calibration pass (sweep a small grid on validation seeds, pick by some criterion, freeze before test seeds) should happen before any of this is reported as more than a qualitative existence check.
  • The forward models are this project's own analytic approximations of the named classes' qualitative shapes (timescale family, single/double-peaked, power-law vs. exponential), not the published templates the manuscript calls for. Good enough to test the PLANNER logic; not good enough to claim astrophysical realism.
  • cost(a)=1 for every action means the utility's division by cost is a no-op right now; the manuscript's resource-cost model (exposure, slew) is unimplemented.
  • Crowding and model-error are this project's own operationalizations of the manuscript's scenario names (cell-masking; time-dilated belief), not validated against any other source — documented as such in scenarios.py.

Phase 1 — Existence proof (2026-10-01)

Built

  • transients/models.py: two modeled classes ("fast" tau=1.0d, "slow" tau=4.0d, pure exponential decay) + one withheld class ("bump": Gaussian rise + exp decay), amplitude drawn log-uniform per event.
  • transients/noise.py: fixed-sigma (SIGMA0=1.0) additive Gaussian noise, pre-drawn per candidate time so every strategy sees identical noise at a given time for a given event (paired comparison).
  • inference/bayes.py: exact discrete joint posterior over (class, amplitude grid of 25 log-spaced points), closed-form Bayesian update, class-marginal posterior, per-class evidence (marginalized over amplitude).
  • inference/mutual_information.py: expected information gain via exact 1D trapezoidal quadrature over the observation y (no Monte Carlo).
  • inference/openworld.py: fixed-hyperparameter RBF-kernel GP (lengthscale=1.2d, var_f=100) as the h0 model; top-level Bayes-factor posterior over {fast, slow, h0}.
  • planners/: random (B0-style), B3_closed_world_eig (ignores h0), P_open_world_eig (includes h0 and the GP predictive in its EIG search).
  • experiments/phase1_existence_proof.py: runs all three strategies on IDENTICAL events/noise per seed, writes config/summary/worked_example JSON (committed) and events_raw.jsonl (gitignored), updates progress.json live.
  • app/viewer.py + app/data.py + app/plotting.py: read-only Streamlit viewer, 6 tabs (one per phase), provenance line (seed/git hash/config hash) on every result page, DEMO warning on tag="demo" runs.
  • launch.py + launch_research.bat: dependency check, background experiment subprocess, viewer on port 8502 (never 8501), browser open, clean shutdown on Ctrl+C.

Tests passed

pytest research/ — 9/9 passed:

  • Bayesian update vs. independently-computed (scipy.stats.norm) closed-form posterior, single update and sequential updates.
  • Posterior concentration sanity check under near-zero noise.
  • EIG == 0 for indistinguishable predictives (known answer).
  • EIG -> ln(2) in the noiseless, well-separated, uniform-binary-prior limit (known answer).
  • EIG strictly between 0 and ln(2) for partial separation (sanity bound).
  • GP log marginal likelihood vs. independently-computed (scipy.stats.norm) closed form for a single observation (known answer).
  • GP behaves as its prior with zero observations.

Quick demo result (seed=0, n_events=100, tag=demo)

strategy confident-wrong (modeled truth) confident-wrong (withheld truth) accuracy (withheld, h0 credited)
random 0.000 0.879 0.000
B3_closed_world_eig 0.000 1.000 0.000
P_open_world_eig 0.000 0.121 0.788

B3 is confidently wrong on every single withheld-class event (it has no way to say "neither"); P drops that to 12% and correctly attributes 79% of withheld events to h0. Neither strategy is ever confidently wrong when the truth is actually a modeled class. This is the qualitative result the research goal predicts, from real code on a real (if small) sample — not asserted a priori.

Surprises

  • Performance bug found during the quick-demo run: the initial expected_information_gain implementation looped in Python over all 2001 quadrature points, calling a Python-level group_entropy (itself another Python loop) at each one. 5 events took 3m42s. Root cause: no vectorization across the quadrature axis. Fixed by vectorizing both the quadrature loop and the per-group logsumexp (loop only over the tiny number of groups, 2-3, never over the 2001 y-points or the event/strategy loops). After the fix: 5 events in 5.2s, 100 events in 159.5s (quick demo), pytest suite 8.2s -> 0.48s. This would have made Phase 2+ (1000+ events, more candidates, orbits) completely infeasible if left unfixed — worth flagging because it was a silent correctness-preserving-but-unusably-slow bug, exactly the kind of thing that would otherwise surface much later as "Phase 2 is taking days."

Decisions made on ambiguous points

  • Streamlit import in research/app/: the branch rule in CLAUDE.md says new work in research/ "must NOT import redis, streamlit, websockets, or any live-stack module." The user's own follow-up instruction explicitly asked for research/app/ to be "a Streamlit app... still headless at the core," and only restated the redis/websockets ban for it. Read this as: the core research packages (transients/, inference/, planners/, experiments/) stay streamlit-free and fully unit-testable; research/app/ is a narrow, explicitly-requested exception that only reads result files and never computes. Flagged to the user rather than silently assumed.
  • launch_research.bat lives at the repo root, outside research/, because a "one-click launcher" only works there, and the user asked for it by that exact name. Also flagged as a narrow exception to "work only inside research/".
  • Phase 1 noise model: fixed-sigma additive Gaussian (not yet the limiting-magnitude model — that's explicitly a Phase 2 item per the original brief). Chosen so the Bayesian update and MI estimator have exact Gaussian likelihoods, which is what makes the known-answer tests possible.
  • GP hyperparameters are fixed, not optimized (lengthscale=1.2d, var_f=100): a notebook-sized existence proof doesn't need marginal- likelihood hyperparameter optimization, and fixing them keeps Phase 1 fully deterministic given a seed. Revisit in Phase 2+ if GP evidence turns out to be sensitive to this choice.
  • Confident-wrong threshold = 0.9, n_obs_budget = 5, amplitude grid = 25 log-spaced points, 20 log-spaced candidate times from 0.2 to 10 days: reasonable round-number choices for a notebook-sized demo, not tuned against the outcome (chosen before the first run, unchanged after seeing results).

What we distrust about this result

  • n=100 is a demo sample; the confident-wrong-rate numbers above have real sampling noise (e.g. 0.879 vs 0.788 are not precise to three decimals at this n). Phase 2's acceptance criterion (1000 events) is what the paired comparison should actually be judged on.
  • The GP's fixed hyperparameters were chosen by eye, not fit or cross-validated. If Phase 2's instrument noise model makes light curves noisier or shorter in duration, these may need revisiting — and if they do, that revisit must happen on validation seeds only, per the branch preregistration rule, not on whatever seed produces a nicer plot.
  • "Confident-wrong" is a threshold-based summary (0.9) of what is really a continuous miscalibration story; Phase 2's AUROC-of-P(h0) metric is a more complete picture and should be treated as the primary evidence over this single-threshold number.
research/PLAN.md

Research plan

How the live OrbitalNet auction maps to the baselines, what is reused, and the design of each phase.

research/ — Open-World Sequential Measurement Planning: PLAN

Status: draft, awaiting approval. No code has been written yet.

0. Research goal (restated)

Choose, at each step, which spacecraft makes which measurement next, to discriminate among K modeled physical explanations of a fading transient plus an explicit "unmodeled" hypothesis h0, where measurement value decays at a hypothesis-dependent rate, across a constellation with intermittent inter-satellite links.

1. How the existing CNP auction maps to baseline B1

The live-stack auction lives entirely in consensus_engine.py::elect_plane_leaders(), re-run every 5s tick over everything in MISSIONS_LEDGER. Its logic, stripped of Redis bookkeeping, is:

  1. P0 gatekeeper (hard filter): a satellite is eligible iff is_task_locked == 0 AND payload_type == required_sensor AND haversine(sat, target) <= zone_radius + 1000.
  2. Score: auction_score = w_proximity * normalize(distance) + w_battery * soc — a weighted sum of exactly two normalized terms, no sequential/information-theoretic reasoning anywhere.
  3. Dispatch: sort eligible bidders by score, take the top required_nodes as a single-shot greedy team (no iteration — one auction tick picks the whole team).
  4. Rolling enclave: each later tick re-checks the same gatekeeper (now "is this member still in range") and drops/reopens on failure.

This is a one-shot, feasibility-gated, highest-static-score dispatcher — it never asks "which measurement teaches me the most." That is exactly the shape of B1: a reimplementation of CNP nearest-capable dispatch, specialized to our setting as:

  • Gatekeeper → visibility/feasibility test (can this spacecraft observe the transient's sky position right now, given its orbit and a minimum elevation/occlusion constraint) replacing the hardware+haversine gatekeeper.
  • Score → a resource/proximity score (e.g. favor the spacecraft with the best geometry/SNR and freest schedule) replacing proximity+battery, computed the same way (weighted sum, no information term).
  • Dispatch → at each decision epoch, argmax over currently-feasible spacecraft; no belief, no expected information gain, no hypothesis tracking. This preserves the "dispatch on static score, ignore what you'd learn" character of the original CNP logic, which is the point of having it as a baseline.

The rolling-enclave handoff (drop out-of-range, reopen) is not reused verbatim for B1 (B1 picks one spacecraft per measurement, not a persistent team), but the same mechanic reappears legitimately in Phase 4 as the model for intermittent inter-satellite contact windows.

2. Reuse vs. headless rewrite

Can reuse (structure/approach, not the running code as-is)

  • SGP4 propagation: physics_engine.py::load_satellites() (TLE parsing via Satrec.twoline2rv) and the sat.sgp4(jd, fr) call pattern. We reuse the approach and a trimmed subset of satellites.txt, but call it synchronously inside a seeded simulation function — not the while True: time.sleep(1) loop that writes to Redis.
  • eci_to_latlon: useful as-is for ground-track geometry if any ground-station visibility is needed, but it is not a visibility/line-of-sight test — see below.
  • Gatekeeper → weighted-score shape from scoring_engine.py (evaluate_gatekeepers → calculate_base_capability → apply_risk_decay): this three-stage pattern (hard filter, weighted linear score, decay term) is a good scaffold for the baselines' resource/adequacy terms and for the hypothesis-dependent value-decay term in the planner. The actual numbers (thresholds like reaction_wheel_rpm >= 0.95, conjunction_prob) are demo flavor and are not reused — they have no bearing on photometric measurement quality.
  • Auction forensics logging shape (AUCTION_LOGS entries in consensus_engine.py): the idea of logging every bidder, not just the winner, per decision — reused as the schema for research/results/*.json so experiments are auditable, but written to files per the branch rules instead of a Redis hash.

Must be rewritten headless (no equivalent exists, or existing code is non-physical)

  • Visibility / contact graph: the only "visibility" check in the live stack is ground-target surface distance via haversine with a flat +1000 km slew margin — it ignores orbital altitude, slant range, elevation angle, and Earth occlusion entirely. For both ground-to-transient visibility (Phase 2) and inter-satellite link contact windows (Phase 4) we need a real line-of-sight/elevation/Earth-occlusion test. This does not exist in the repo and must be written new in orbitsim/geometry.py.
  • Sensor/payload model: classify_satellite() assigns sensor type by MD5 hash of the satellite's name purely to make the Starlink-only demo dataset look varied. It carries no physical information and must not leak (even by copy-paste) into the noise model. Phase 2's instrument noise model is derived from limiting magnitudes, built fresh.
  • State/control flow: every live-stack engine is an infinite loop with Redis as the only interface (side effects, not return values) — this is untestable and non-reproducible by construction. All research code is plain functions over explicit arguments/returns so it can be unit tested and driven by --seed.
  • Bayesian hypothesis update, mutual-information/EIG estimator, GP auxiliary model for h0, decentralized belief merge: no existing equivalent anywhere in the repo — entirely new (inference/, decentral/).

3. Module layout

research/
  PLAN.md                       (this file)
  PREREGISTRATION.md            (added later, gates test-seed access — Phase 2+)
  requirements-research.txt     (new deps go here only; requirements.txt untouched)
  orbitsim/
    tle.py                      # headless TLE load + SGP4 propagate (reuses physics_engine.py's approach)
    geometry.py                 # ECI/ECEF transforms, elevation, Earth-occlusion line-of-sight
    contacts.py                 # builds per-seed visibility/contact schedules (ground+ISL)
  transients/
    models.py                   # K parametric light-curve classes + withheld-class generator
    noise.py                    # instrument noise from limiting magnitude
  inference/
    bayes.py                    # closed-world posterior update over K hypotheses
    mutual_information.py       # EIG / MI estimator
    openworld.py                # GP auxiliary model + h0 posterior / adequacy term
  planners/
    baselines.py                # B0 (random/round-robin), B1 (CNP reimpl.), B2, B3
    eig_planner.py               # P: closed-world expected-information-gain planner
    openworld_planner.py          # P + h0 adequacy/decay term
  decentral/                    # Phase 4 only
    belief_merge.py
    marginal_auction.py
  experiments/
    phase1_existence_proof.py
    phase2_single_planner.py
    phase3_ablations.py
    phase4_decentralized.py
    phase5_real_data_check.py
  results/                      # JSON/CSV outputs only, nothing hand-edited
  figures/                      # regenerated only from results/*.json
  tests/
    test_bayes_update.py
    test_mutual_information.py
    test_belief_merge_equivalence.py   # Phase 4 exactness proof
    ...

All packages under research/ are pure numpy/scipy — no redis, streamlit, websockets, or other live-stack import, per the branch rules in CLAUDE.md. Each experiment script takes --seed and --n-events and writes only to research/results/.

4. Risks

  • Superficial B1 mapping: the live auction's "visibility" is a flat-Earth haversine hack. If orbitsim/geometry.py's real occlusion/elevation test and the reused haversine-style gatekeeper disagree in spirit, B1 risks being a strawman rather than a faithful reimplementation of "what the CNP logic actually does." Mitigate by reviewing B1 against the mapping in §1 before Phase 2 is scored.
  • Accidental live-stack coupling: pulling logic from physics_engine.py / scoring_engine.py by copy-paste risks dragging in a redis/hal_simulator import transitively. Enforced by a lint/test step (e.g. grep or an import-time check in tests/) that fails if anything under research/ imports a banned module.
  • MI/EIG estimator correctness: expected-information-gain and GP marginal-likelihood estimators are easy to get subtly wrong (Monte Carlo variance, numerical stability). Phase 1's acceptance criterion (a known-answer test case) is the main defense; the toy case must be chosen so the true MI is analytically computable, not just plausible.
  • Belief-merge exactness (Phase 4): "merging via measurement IDs reproduces the centralized posterior exactly" only holds if measurements are not double-counted and log-odds combine additively under conditional independence. Any planner path that lets a spacecraft see a duplicate or correlated measurement ID breaks this silently; the Phase 4 test must include a duplicate-ID case, not just the happy path.
  • Scope/sequencing: six phases is a lot of surface area. Phase 6 (firmware/testbed) is explicitly design-only until Phases 1-4 pass — this plan does not create any firmware code or hardware-facing scaffolding now.
  • Seed discipline: validation vs. test seeds must stay separated before PREREGISTRATION.md is committed. Risk is an experiment script using an un-flagged seed range during Phase 2/3 tuning; mitigate by making the seed-role split an explicit, checked argument rather than a convention to remember.

5. Open questions before Phase 1 starts

  • Confirm the exact set of strategies expected in Phase 1 output (closed-world EIG planner + GP/h0 planner only, per the prompt) vs. Phase 2's fuller strategy list (B0, B1, B2, B3, P) — Phase 1 does not need B1 since there are no orbits yet.
  • Confirm results/ and figures/ are committed (so reviewers can see without rerunning) or gitignored (regenerate-only) — affects repo hygiene but not the code.

Nothing above requires a decision to start Phase 1; flagging for awareness.

6. Phase 2 realignment to the manuscript (2026-10-01)

Phase 2 was originally built (classes, utility, scenarios, baselines) as this project's own design, documented above as such, because no manuscript existed in the repo yet. The user then supplied Hypothesis-Discriminating Observation Campaigns for Distributed Space Observatories.pdf ("the manuscript"), which specifies these choices precisely. Phase 2 was rebuilt to match it. See research/LOG.md's Phase 2 entry for the full list of what changed and the simplifications made where exact fidelity was not tractable at this compute budget (decay term, exposure/slew cost, wavelength bands).

Scenario split across phases: the manuscript lists five scenarios (Nominal, Communication sweep, Node loss, Crowding, Model error) without assigning them to a phase. Communication sweep and Node loss are about multi-spacecraft belief staleness, which does not exist until Phase 4 (decentralization). Phase 2 (a single centralized planner) runs Nominal, Crowding, and Model error; Phase 4 runs Communication sweep and Node loss alongside its own contact-fraction sweep, which subsumes them.

Rendered from research/LOG.md (updated 2026-10-11) · research/PLAN.md (updated 2026-10-11) at build time. Edit the markdown, not this page.