Draft manuscript · Space Autonomy & Time-Domain Astrophysics
Planning follow-up campaigns for fading transients under intermittent connectivity
Short form: Differential Diagnosis for the Sky
Can a constellation work out why something is fading — before the evidence fades along with it?
The problem
02:14 UTCA wide-field survey flags a new transient. At first light it could be a kilonova, a shock-cooling supernova, a relativistic afterglow, or a stellar flare — in the first hour, the photometry of all four looks almost the same.
The measurement that would actually tell them apart exists for a few hours, maybe a day, and then it's gone for good. There is no retake.
Six to eight satellites can see the sky region between them. None has the full picture. None can wait for a ground operator to decide. They have to work out, on their own and from whatever they currently believe, which of them should point which instrument where — next, right now — and keep doing that as the picture changes and the window closes.
That decision — not the detection, not the classification afterward, but the choice of next measurement under a closing deadline and a hypothesis set that might be incomplete — is what this project plans for. OrbitalNet is the distributed substrate that makes the campaign executable; it is not itself the contribution. The planning layer on top of it is.
Fast optical transients such as kilonovae, shock-cooling supernovae, relativistic afterglows and stellar flares are often indistinguishable in their first hours, yet the measurements that separate them lose value within hours to days as the sources fade. Current follow-up is organised by recipes chosen per assumed class, by classifier-uncertainty ranking of objects, or by distributed allocation toward fixed objectives. None of these plans a sequence of heterogeneous measurements whose purpose is to discriminate competing physical explanations of one event before the discriminating signal disappears, and none accounts for the possibility that the true explanation lies outside the hypothesis set. We formulate follow-up as open-world sequential model discrimination with time-decaying measurement value, executed by spacecraft that hold unequal, intermittently synchronised beliefs. The planner maintains a posterior over K physical models plus an explicit unmodeled alternative represented by a flexible auxiliary model, selects measurements by a utility combining expected discrimination, expected evidence of model inadequacy and opportunity decay, and switches from discrimination to evidence preservation when inadequacy becomes probable. Allocation across spacecraft uses a marginal-gain auction that tolerates stale beliefs. We evaluate on a simulated heterogeneous low-Earth-orbit constellation with physically grounded visibility, slew, sensitivity and contact constraints, using injected transients from published light-curve models and withheld classes as ground-truth anomalies, and replicate the communication layer on an ESP32 hardware-in-the-loop testbed. We test whether the planner reaches correct identification earlier, recovers more early-time information, and flags withheld classes more reliably than reactive, uncertainty-sampling and closed-world information-gain baselines under identical budgets.
We show that a distributed space observatory that keeps an explicit “none of the above” explanation and accounts for fading signals recovers more of the physically unrecoverable information about fading transients than closed-world or reactive planners under the same constraints.
When an unexplained transient appears, a constellation of heterogeneous instruments must decide, repeatedly and without waiting for the ground, which spacecraft should take which measurement next. The right choice depends on which explanations remain plausible, how fast each explanation predicts the discriminating signal will vanish, what other spacecraft are already doing, and whether any existing explanation fits at all.
The measurement that separates two hypotheses is often the one available only briefly — an early ultraviolet colour, the first spectrum before a photosphere cools, a near-infrared epoch before a kilonova fades. A wrong or late choice is not a delay; it is a permanent loss, because wide-field surveys now produce far more candidates than follow-up capacity can serve, which makes the choice of measurement, not just of target, the binding constraint. This is not hypothetical: automated pipelines have mis-flagged a Type IIb supernova as a kilonova candidate for days before spectroscopic confirmation ruled it out.
Fig. 1 · A measurement's power to tell hypotheses apart shrinks with the source, then hits zero.
Schematic curve for intuition only — not a measured or simulated light curve.
Hypothesis-discriminating observation exists for Earth science and rover geology, but with a closed hypothesis set and a single planner. Distributed spacecraft allocation exists for fixed coverage objectives, with no hypothesis state at all. Misspecification-robust experimental design exists for single experimenters without deadlines. Astronomical follow-up optimisation targets training-set construction or one-shot spectroscopic selection. No located work combines open-world hypothesis discrimination, a fading-value objective, multi-instrument sequencing and decentralised execution in one campaign planner — and none measures what the combination recovers.
Can a distributed, resource-constrained space observatory plan sequential multi-instrument measurements that discriminate competing physical explanations of a fading transient, and recognise when none of them applies, better than reactive and closed-world planners under identical physical and communication constraints?
The table below names, for each closest prior system, what it does, what it assumes, and the specific capability it stops short of — the delta this project occupies. All entries are draft claims pending a full read of the cited sources (§References).
| System | What it does | What it does not do |
|---|---|---|
| JPL POISE | Chooses observations that best discriminate competing hypotheses and re-tasks Earth-observing assets under cost and feasibility constraints, given domain models. | No explicit unmodeled hypothesis; no decentralized planning under intermittent links. |
| SHM‑POMDP | Plans belief-space observations over science hypotheses for a single rover in a static, closed hypothesis structure. | Single agent; no time-decaying targets; no model-inadequacy signal. |
| NASA DSA / Starling | Four spacecraft share perspectives and allocate ionospheric measurements without ground direction, surviving member loss, toward a fixed shared objective. | No hypothesis-discrimination objective; proves distributed allocation flies, not epistemic planning. |
| Astudillo et al. 2020 | Information-theoretic choice of which objects receive a spectroscopic follow-up slot. | One follow-up type; population-level selection; no timing or sequencing. |
| RESSPECT / Fink active learning | Chooses spectroscopic targets to improve a population classifier's training set. | Goal is a better classifier, not per-event discrimination; does not model misspecification. |
| NMMA‑Astro‑COLIBRI | Fits early photometry to competing kilonova/supernova/afterglow models and reports Bayes factors as data arrive. | Does not choose the next measurement; comparison is after the fact. |
| Robust BED (EGIG / GBOED) | Designs experiments that remain informative, or detect misspecification, when the model may be wrong. | Single experimenter, no hard deadlines, no distributed assets — the mechanism this project borrows, applied to a new operational setting. |
| GRANDMA / TOM Toolkit / AEON | Coordinates global human follow-up networks for transient alerts. | Humans decide the campaign logic; no autonomous hypothesis-driven planner. |
The pattern across the broader landscape (twelve candidate questions were traced to a prior-art bottleneck before this gap was selected) matters more than any single row: every lower-stack capability — detect, rank, allocate, retarget — has prior art. What recurs unresolved is the layer where a distributed system must reason about which explanation is true, under a closing window, while allowing that its explanations may all be wrong.
Four falsifiable hypotheses follow from the research question. Each names the result that would refute it; a refuted hypothesis — H4 in particular — is still a publishable finding and will be reported as a refutation, not tuned away.
The proposed planner (P) reaches 95% posterior on the true modeled class in less median time than every baseline, using no more measurements.
Refuted if a baseline matches or beats P's median time-to-95% at equal budget.
Removing the opportunity-decay term reduces recovered early-time information more for fast-evolving classes (kilonova, shock-cooling) than for slow ones.
Refuted if the decay-ablation's information loss is the same (or worse) for slow classes.
On withheld classes, the planner's model-inadequacy probability P(h₀) separates unmodeled from modeled events better than classifier-entropy thresholding, at a matched false-alarm rate.
Refuted if AUROC(P(h₀)) ≤ AUROC(entropy threshold) on held-out events.
Performance degrades gracefully as inter-satellite contact fraction falls; the stale-belief auction retains a measurable fraction of centralized-oracle performance where independent greedy planners duplicate measurements.
Refuted if the auction collapses to independent-greedy performance at all tested contact fractions.
OrbitalNet, the distributed substrate, supplies exactly three things to the planner: who can see the target when, who can talk to whom when, and a message layer that survives link gaps. The contribution sits entirely in the belief and utility layers built on top.
Fig. 2 · Satellites bid their marginal utility over whatever stale beliefs they currently hold; only one has an open ground contact.
Animated schematic of the mechanism, not a recording of simulated motion.
loop: step 7 ⟳ step 2
If P(h₀) > τ for two consecutive updates ⟶ switch to preservation mode: discrimination terms drop out, the planner maximizes time/wavelength coverage of the auxiliary model, and onboard retention priority for this event's raw data is raised.
The event's true class h is drawn from K physical models plus one unmodeled alternative h₀. Each modeled class k carries parameters θk (explosion time, distance, ejecta mass, temperature, and similarly class-specific quantities) and a forward model predicting the flux an instrument would measure:
The unmodeled alternative is represented deliberately by a flexible auxiliary model — a Gaussian process over time and wavelength with broad priors — so that it never wins by fitting a known class well. It wins only when every known class fits badly:
Fig. 3 · A posterior that hasn't converged — kilonova and shock-cooling are still close, h₀ stays live.
Illustrative toy bars for intuition — synthetic, not output from the planner or any simulation.
An action a = (spacecraft i, instrument m, band/mode λ, start time t, exposure Δ) is feasible only if the orbital layer reports the target visible and outside Sun, Earth-limb and South Atlantic Anomaly exclusion, the slew fits the schedule, and power and storage budgets hold. Measurement noise σa is read from the instrument's limiting-magnitude curve, so a faded source automatically yields an uninformative measurement — the physical mechanism behind the decay term below.
A candidate measurement's score combines three terms, each a mutual-information or decision-theoretic quantity, divided by its resource cost c(a):
The first term is the mutual information between the measurement and the class label: how much it separates the remaining explanations. Z is a binary indicator of whether the truth is h₀; the second term is how much the measurement would reveal about model adequacy, following the misspecification-indicator idea from robust Bayesian experimental design (§2). The third term is opportunity decay: the expected information lost if this measurement is postponed to the next feasible slot, averaged over the current posterior — classes that fade fast make early measurements urgent only while they remain plausible:
For a fixed class variable and measurements conditionally independent given the class, mutual information is monotone submodular, so greedy batch selection is within a factor (1 − 1/e) of optimal — but only for a centralized planner with perfectly shared beliefs. The proposed auction instead lets each spacecraft bid its marginal gain given whatever commitments it has heard about. Stale beliefs break the optimality guarantee; measuring by how much is H4, not a result to hide.
If the posterior on h₀ exceeds a threshold τ for two consecutive updates, the utility changes: discrimination terms drop out, the planner maximizes coverage in time and wavelength of the auxiliary model, and onboard retention priority for all raw data on the event is raised. τ is fixed on a validation split to hit a chosen false-alarm rate on modeled classes, never on the test split.
Fig. 4 · The one-way switch that trades discrimination for evidence preservation once the open-world term dominates.
The moving dot paces the transition condition; it is not a timer or a measured delay.
Each spacecraft keeps a grid or particle posterior per class. Messages carry per-class log-evidence and a compact list of committed future measurements, never raw data. Merging two beliefs that share earlier measurements must avoid double counting; tagging every measurement with a unique ID and summing only unseen evidence handles this exactly for discrete classes — the property exercised directly by the decentralization experiments (§6).
Every strategy sees the same injected events, the same orbits, the same instrument budgets and the same contact schedule; only the decision rule changes. That pairing is what turns the simulator into evidence.
| Role | Classes | Why |
|---|---|---|
| Modeled (K=4) | Kilonova; shock-cooling / Type IIb early phase; GRB afterglow (broken power-law); M-dwarf flare. | All plausible early confusions; published analytic or template models exist for each. |
| Withheld (truth = h₀) | Fast blue optical transient (AT2018cow-like); tidal disruption event; one synthetic "nonsense" class with physically odd colour evolution. | Ground truth for the open-world test; the planner never sees these models. |
Forward models are taken from published light-curve codes or templates (e.g. sncosmo templates, kilonova grids, flare templates), each separately cited, so reviewers can check the physics first. Distance and host extinction are sampled from realistic ranges and classes weighted by rough relative rates, to avoid a toy balanced problem.
Six to eight low-Earth-orbit spacecraft from the OrbitalNet propagator, mixed payloads: two ultraviolet imagers, two optical multi-band imagers, one near-infrared imager, one low-resolution spectrograph, optionally one X-ray monitor. Each instrument carries a limiting-magnitude-versus-exposure curve, a slew rate and duty-cycle limits. Visibility uses SGP4 with Earth occultation, Sun avoidance and South Atlantic Anomaly exclusion; inter-satellite links follow line-of-sight and range, sampled into a contact schedule.
| Label | Strategy | Role |
|---|---|---|
| B0 | Fixed recipe: two optical bands every N hours, a spectrum once brighter than a threshold. | Human-style practice. |
| B1 | Reactive nearest-capable dispatcher (the existing project baseline). | Shows what the pre-existing system does. |
| B2 | Classifier-entropy sampling: measure whatever most reduces predicted-class entropy; no physics likelihoods. | Active-learning-style baseline. |
| B3 | Closed-world expected information gain, centralized, no decay term, no h₀. | POIE-style baseline — the key comparison. |
| P | Full method: h₀ + decay + stale-belief auction. | Proposed. |
| O | Centralized, perfect communication, full method. | Upper bound. |
Each ablation removes exactly one term so a reader can see which part earns its keep: drop h₀ entirely; drop the decay term (γ = 0); drop the inadequacy term while keeping h₀ in the belief (β = 0); replace the auction with independent greedy planners; vary K.
At least 1,000 injected events per scenario, fixed seeds shared across strategies. Strategies are compared with paired tests (Wilcoxon signed-rank on per-event metrics), medians reported with bootstrap 95% intervals, and multiple comparisons corrected (Holm). Effect sizes are reported alongside p-values, not in place of them. The mode-switch threshold τ and the utility weights β, γ are tuned on a validation split of events, never on the test split.
Archival data cannot answer what a different measurement choice would have shown, since only the measurements actually taken exist. Real public light curves (e.g. from ZTF via a broker such as Fink or ALeRCE) are used for two narrower purposes only: checking that the forward models and noise model match reality, and a retrospective replay on densely observed events, where the planner may choose only among epochs that genuinely exist. This limitation is stated plainly rather than implied away.
The decision layer is replicated on physical hardware to show it survives real asynchronous radios and microcontroller-class compute, with results close to simulation — not to reproduce flight-qualified radiation tolerance, real RF link budgets, attitude control or flight software, which the testbed explicitly does not claim.
Fig. 5 · Each node runs the planner's belief and bidding logic on-device; links blink independently to stand in for intermittent, lossy ESP-NOW contact.
Topology sketch of the planned testbed — link blinking is illustrative, not a captured radio trace.
| Element | Hardware | Constraint reproduced |
|---|---|---|
| Spacecraft nodes (6–8) | ESP32 boards running belief update, utility and bidding on-device. | Limited compute and memory; tests whether the planner fits a microcontroller-class budget. |
| Crosslinks | ESP-NOW radio, gated on and off by the OrbitalNet contact schedule replayed in real time (or accelerated). | Intermittent links, message loss, asynchrony. |
| Clocks | Independent node clocks, synchronized only via messages. | Clock drift and ordering errors. |
| Sky oracle & ground | Raspberry Pi serving measurement outcomes from the simulated universe; acts as ground station. | Separates ground truth from the nodes' own beliefs. |
| Faults | Scheduled power-cycling of nodes. | Node loss and rejoin. |
Nothing in this section is demonstrated; these are the predictions the experiments above are built to confirm or refute.
| Type | Threat | Mitigation |
|---|---|---|
| Internal | Tuning weights on the test events. | Separate validation set; fixed seeds; metrics pre-registered before final runs. |
| Internal | Baselines implemented weakly. | B3 receives the same likelihoods and budget as P; baselines tuned as carefully as P. |
| Construct | Posterior contraction may not equal "scientific knowledge." | Also report accuracy of recovered physical parameters against injected truth. |
| External | Simulator realism. | Validate noise and light curves against real broker data; state which physics is omitted. |
| External | Hardware testbed does not generalise to flight. | Claim limited to decision-layer behaviour under asynchrony and compute limits. |
| Conclusion | Many pairwise comparisons. | Holm correction; effect sizes; confidence intervals. |
The release will include the simulator, class forward models (or scripts that fetch them), orbital elements, contact schedules, every strategy's code, the analysis notebooks producing each figure, ESP32 firmware and the wiring diagram, under an open licence with a tagged, DOI-archived version.
Planned extensions: scaling to larger hypothesis libraries; shadow-mode integration with a real alert stream alongside a partner follow-up network (planner advises, humans act); learned, amortized design policies to reduce onboard compute; inadequacy-driven onboard retention as a second paper; and, for an eventual flight path, radiation-tolerant compute, real link budgets and a hosted-payload or rideshare demonstration.
These entries were located by web search on 24 September 2026; only abstracts or summaries were read while drafting, not full texts. They are leads, not confirmed citations — each must be read in full and its author list verified before use in any submitted version.