← OrbitalNet research

Draft manuscript · Space Autonomy & Time-Domain Astrophysics

Open-World, Deadline-Aware Hypothesis Discrimination for Distributed Space Observatories

Planning follow-up campaigns for fading transients under intermittent connectivity

Short form: Differential Diagnosis for the Sky

Can a constellation work out why something is fading — before the evidence fades along with it?

Draft · 24 September 2026 · Target venues & submission roadmap · Status: pre-registration, unrun

The problem

02:14 UTCA wide-field survey flags a new transient. At first light it could be a kilonova, a shock-cooling supernova, a relativistic afterglow, or a stellar flare — in the first hour, the photometry of all four looks almost the same.

The measurement that would actually tell them apart exists for a few hours, maybe a day, and then it's gone for good. There is no retake.

Six to eight satellites can see the sky region between them. None has the full picture. None can wait for a ground operator to decide. They have to work out, on their own and from whatever they currently believe, which of them should point which instrument where — next, right now — and keep doing that as the picture changes and the window closes.

That decision — not the detection, not the classification afterward, but the choice of next measurement under a closing deadline and a hypothesis set that might be incomplete — is what this project plans for. OrbitalNet is the distributed substrate that makes the campaign executable; it is not itself the contribution. The planning layer on top of it is.

Abstract

Fast optical transients such as kilonovae, shock-cooling supernovae, relativistic afterglows and stellar flares are often indistinguishable in their first hours, yet the measurements that separate them lose value within hours to days as the sources fade. Current follow-up is organised by recipes chosen per assumed class, by classifier-uncertainty ranking of objects, or by distributed allocation toward fixed objectives. None of these plans a sequence of heterogeneous measurements whose purpose is to discriminate competing physical explanations of one event before the discriminating signal disappears, and none accounts for the possibility that the true explanation lies outside the hypothesis set. We formulate follow-up as open-world sequential model discrimination with time-decaying measurement value, executed by spacecraft that hold unequal, intermittently synchronised beliefs. The planner maintains a posterior over K physical models plus an explicit unmodeled alternative represented by a flexible auxiliary model, selects measurements by a utility combining expected discrimination, expected evidence of model inadequacy and opportunity decay, and switches from discrimination to evidence preservation when inadequacy becomes probable. Allocation across spacecraft uses a marginal-gain auction that tolerates stale beliefs. We evaluate on a simulated heterogeneous low-Earth-orbit constellation with physically grounded visibility, slew, sensitivity and contact constraints, using injected transients from published light-curve models and withheld classes as ground-truth anomalies, and replicate the communication layer on an ESP32 hardware-in-the-loop testbed. We test whether the planner reaches correct identification earlier, recovers more early-time information, and flags withheld classes more reliably than reactive, uncertainty-sampling and closed-world information-gain baselines under identical budgets.

Experiment status
Falsifiable hypotheses4
Modeled transient classes4
Withheld (open-world) classes3
Strategies compared6
Simulated events run0
Real-data replay events run0
ResultsPENDING
This counts what's designed, not what's demonstrated. It updates when the experiment is actually run — nothing above that line is a result.

We show that a distributed space observatory that keeps an explicit “none of the above” explanation and accounts for fading signals recovers more of the physically unrecoverable information about fading transients than closed-world or reactive planners under the same constraints.

1Introduction

When an unexplained transient appears, a constellation of heterogeneous instruments must decide, repeatedly and without waiting for the ground, which spacecraft should take which measurement next. The right choice depends on which explanations remain plausible, how fast each explanation predicts the discriminating signal will vanish, what other spacecraft are already doing, and whether any existing explanation fits at all.

The measurement that separates two hypotheses is often the one available only briefly — an early ultraviolet colour, the first spectrum before a photosphere cools, a near-infrared epoch before a kilonova fades. A wrong or late choice is not a delay; it is a permanent loss, because wide-field surveys now produce far more candidates than follow-up capacity can serve, which makes the choice of measurement, not just of target, the binding constraint. This is not hypothetical: automated pipelines have mis-flagged a Type IIb supernova as a kilonova candidate for days before spectroscopic confirmation ruled it out.

A closing discrimination window A schematic fading light curve with a shaded window of high discriminating value that shrinks and then disappears as the source fades. time since trigger → brightness discriminating window signal below instrument noise — no measurement left to take

Fig. 1 · A measurement's power to tell hypotheses apart shrinks with the source, then hits zero.

Schematic curve for intuition only — not a measured or simulated light curve.

Research gap

Hypothesis-discriminating observation exists for Earth science and rover geology, but with a closed hypothesis set and a single planner. Distributed spacecraft allocation exists for fixed coverage objectives, with no hypothesis state at all. Misspecification-robust experimental design exists for single experimenters without deadlines. Astronomical follow-up optimisation targets training-set construction or one-shot spectroscopic selection. No located work combines open-world hypothesis discrimination, a fading-value objective, multi-instrument sequencing and decentralised execution in one campaign planner — and none measures what the combination recovers.

Research question

Can a distributed, resource-constrained space observatory plan sequential multi-instrument measurements that discriminate competing physical explanations of a fading transient, and recognise when none of them applies, better than reactive and closed-world planners under identical physical and communication constraints?

2Related work

The table below names, for each closest prior system, what it does, what it assumes, and the specific capability it stops short of — the delta this project occupies. All entries are draft claims pending a full read of the cited sources (§References).

Table 1 · Closest prior work
SystemWhat it doesWhat it does not do
JPL POISEChooses observations that best discriminate competing hypotheses and re-tasks Earth-observing assets under cost and feasibility constraints, given domain models.No explicit unmodeled hypothesis; no decentralized planning under intermittent links.
SHM‑POMDPPlans belief-space observations over science hypotheses for a single rover in a static, closed hypothesis structure.Single agent; no time-decaying targets; no model-inadequacy signal.
NASA DSA / StarlingFour spacecraft share perspectives and allocate ionospheric measurements without ground direction, surviving member loss, toward a fixed shared objective.No hypothesis-discrimination objective; proves distributed allocation flies, not epistemic planning.
Astudillo et al. 2020Information-theoretic choice of which objects receive a spectroscopic follow-up slot.One follow-up type; population-level selection; no timing or sequencing.
RESSPECT / Fink active learningChooses spectroscopic targets to improve a population classifier's training set.Goal is a better classifier, not per-event discrimination; does not model misspecification.
NMMA‑Astro‑COLIBRIFits early photometry to competing kilonova/supernova/afterglow models and reports Bayes factors as data arrive.Does not choose the next measurement; comparison is after the fact.
Robust BED (EGIG / GBOED)Designs experiments that remain informative, or detect misspecification, when the model may be wrong.Single experimenter, no hard deadlines, no distributed assets — the mechanism this project borrows, applied to a new operational setting.
GRANDMA / TOM Toolkit / AEONCoordinates global human follow-up networks for transient alerts.Humans decide the campaign logic; no autonomous hypothesis-driven planner.

The pattern across the broader landscape (twelve candidate questions were traced to a prior-art bottleneck before this gap was selected) matters more than any single row: every lower-stack capability — detect, rank, allocate, retarget — has prior art. What recurs unresolved is the layer where a distributed system must reason about which explanation is true, under a closing window, while allowing that its explanations may all be wrong.

3Hypotheses

Four falsifiable hypotheses follow from the research question. Each names the result that would refute it; a refuted hypothesis — H4 in particular — is still a publishable finding and will be reported as a refutation, not tuned away.

  1. H1 · Discrimination

    The proposed planner (P) reaches 95% posterior on the true modeled class in less median time than every baseline, using no more measurements.

    Refuted if a baseline matches or beats P's median time-to-95% at equal budget.

  2. H2 · Irreversibility

    Removing the opportunity-decay term reduces recovered early-time information more for fast-evolving classes (kilonova, shock-cooling) than for slow ones.

    Refuted if the decay-ablation's information loss is the same (or worse) for slow classes.

  3. H3 · Open world

    On withheld classes, the planner's model-inadequacy probability P(h₀) separates unmodeled from modeled events better than classifier-entropy thresholding, at a matched false-alarm rate.

    Refuted if AUROC(P(h₀)) ≤ AUROC(entropy threshold) on held-out events.

  4. H4 · Distribution

    Performance degrades gracefully as inter-satellite contact fraction falls; the stale-belief auction retains a measurable fraction of centralized-oracle performance where independent greedy planners duplicate measurements.

    Refuted if the auction collapses to independent-greedy performance at all tested contact fractions.

4Contributions

  1. A formulation of transient follow-up as open-world sequential model discrimination with hypothesis-dependent measurement decay.
  2. A decentralized allocation rule for epistemic (discrimination) objectives under intermittent inter-satellite links.
  3. A discrimination-to-preservation mode switch driven by model-inadequacy evidence.
  4. An open simulator and benchmark with withheld-class tests, plus a hardware-in-the-loop communication testbed.

5System formulation

OrbitalNet, the distributed substrate, supplies exactly three things to the planner: who can see the target when, who can talk to whom when, and a message layer that survives link gaps. The contribution sits entirely in the belief and utility layers built on top.

Constellation and ground link schematic Three satellites in orbit with intermittent crosslinks bid in a marginal-gain auction; one satellite has a ground contact window. GROUND HORIZON bid → ← bid SAT A SAT B SAT C GROUND contact MARGINAL-GAIN AUCTION

Fig. 2 · Satellites bid their marginal utility over whatever stale beliefs they currently hold; only one has an open ground contact.

Animated schematic of the mechanism, not a recording of simulated motion.

1Orbital layerSGP4 visibility, slew, eclipse
⟶
2Feasible actionsper spacecraft, per epoch
⟶
3Utility enginediscrimination + inadequacy + decay
⟶
4Marginal-gain auctionstale-belief tolerant
⟶
5Execute measurement—
⟶
6Local Bayesian updateper-class posterior
⟶
7Belief exchangewhenever contact allows

loop: step 7 ⟳ step 2

If P(h₀) > τ for two consecutive updates  ⟶  switch to preservation mode: discrimination terms drop out, the planner maximizes time/wavelength coverage of the auxiliary model, and onboard retention priority for this event's raw data is raised.

5.1 Hypothesis space

The event's true class h is drawn from K physical models plus one unmodeled alternative h₀. Each modeled class k carries parameters θk (explosion time, distance, ejecta mass, temperature, and similarly class-specific quantities) and a forward model predicting the flux an instrument would measure:

ℋ = {h1, …, hK} ∪ {h0},    y | hk, θk, a  ∼  𝒩 ( fk(θk, λa, ta),  σa2(fk) )
eq. 1

The unmodeled alternative is represented deliberately by a flexible auxiliary model — a Gaussian process over time and wavelength with broad priors — so that it never wins by fitting a known class well. It wins only when every known class fits badly:

y | h0, a  ∼  𝒢𝒫 ( m(t,λ),  kbroad ((t,λ),(t′,λ′)) )
eq. 2
Toy belief distribution over hypotheses Illustrative bars for four modeled classes and the open-world alternative, gently varying to suggest a posterior that is still unsettled. h₁ kilonova h₂ SC-SN h₃ afterglow h₄ flare h₀ “none”

Fig. 3 · A posterior that hasn't converged — kilonova and shock-cooling are still close, h₀ stays live.

Illustrative toy bars for intuition — synthetic, not output from the planner or any simulation.

5.2 Actions and feasibility

An action a = (spacecraft i, instrument m, band/mode λ, start time t, exposure Δ) is feasible only if the orbital layer reports the target visible and outside Sun, Earth-limb and South Atlantic Anomaly exclusion, the slew fits the schedule, and power and storage budgets hold. Measurement noise σa is read from the instrument's limiting-magnitude curve, so a faded source automatically yields an uninformative measurement — the physical mechanism behind the decay term below.

5.3 Utility

A candidate measurement's score combines three terms, each a mutual-information or decision-theoretic quantity, divided by its resource cost c(a):

U(a | bt)  =  I(H1:K; Ya | bt) + β I(Z; Ya | bt) + γ D(a, bt) c(a)
eq. 3

The first term is the mutual information between the measurement and the class label: how much it separates the remaining explanations. Z is a binary indicator of whether the truth is h₀; the second term is how much the measurement would reveal about model adequacy, following the misspecification-indicator idea from robust Bayesian experimental design (§2). The third term is opportunity decay: the expected information lost if this measurement is postponed to the next feasible slot, averaged over the current posterior — classes that fade fast make early measurements urgent only while they remain plausible:

D(a, bt)  =  ∑k bt(hk)  [  Ik(Ya) − Ik(Ya′) ]+ ,   a′ = same measurement at the next feasible time
eq. 4

5.4 Multi-spacecraft allocation

For a fixed class variable and measurements conditionally independent given the class, mutual information is monotone submodular, so greedy batch selection is within a factor (1 − 1/e) of optimal — but only for a centralized planner with perfectly shared beliefs. The proposed auction instead lets each spacecraft bid its marginal gain given whatever commitments it has heard about. Stale beliefs break the optimality guarantee; measuring by how much is H4, not a result to hide.

5.5 Mode switch

If the posterior on h₀ exceeds a threshold τ for two consecutive updates, the utility changes: discrimination terms drop out, the planner maximizes coverage in time and wavelength of the auxiliary model, and onboard retention priority for all raw data on the event is raised. τ is fixed on a validation split to hit a chosen false-alarm rate on modeled classes, never on the test split.

Mode-switch state diagram Planner starts in discrimination mode and switches to preservation mode once model-inadequacy probability exceeds the threshold twice in a row. DISCRIMINATION maximize I(H;Y) + β·I(Z;Y) P(h₀) > τ, 2 updates running PRESERVATION maximize time/λ coverage,

Fig. 4 · The one-way switch that trades discrimination for evidence preservation once the open-world term dominates.

The moving dot paces the transition condition; it is not a timer or a measured delay.

5.6 Belief update and communication

Each spacecraft keeps a grid or particle posterior per class. Messages carry per-class log-evidence and a compact list of committed future measurements, never raw data. Merging two beliefs that share earlier measurements must avoid double counting; tagging every measurement with a unique ID and summing only unseen evidence handles this exactly for discrete classes — the property exercised directly by the decentralization experiments (§6).

6Experimental design

Every strategy sees the same injected events, the same orbits, the same instrument budgets and the same contact schedule; only the decision rule changes. That pairing is what turns the simulator into evidence.

6.1 Simulated universe

Table 2 · Injected event classes
RoleClassesWhy
Modeled (K=4)Kilonova; shock-cooling / Type IIb early phase; GRB afterglow (broken power-law); M-dwarf flare.All plausible early confusions; published analytic or template models exist for each.
Withheld (truth = h₀)Fast blue optical transient (AT2018cow-like); tidal disruption event; one synthetic "nonsense" class with physically odd colour evolution.Ground truth for the open-world test; the planner never sees these models.

Forward models are taken from published light-curve codes or templates (e.g. sncosmo templates, kilonova grids, flare templates), each separately cited, so reviewers can check the physics first. Distance and host extinction are sampled from realistic ranges and classes weighted by rough relative rates, to avoid a toy balanced problem.

6.2 Simulated constellation

Six to eight low-Earth-orbit spacecraft from the OrbitalNet propagator, mixed payloads: two ultraviolet imagers, two optical multi-band imagers, one near-infrared imager, one low-resolution spectrograph, optionally one X-ray monitor. Each instrument carries a limiting-magnitude-versus-exposure curve, a slew rate and duty-cycle limits. Visibility uses SGP4 with Earth occultation, Sun avoidance and South Atlantic Anomaly exclusion; inter-satellite links follow line-of-sight and range, sampled into a contact schedule.

6.3 Strategies compared

Table 3 · Strategies under identical budgets
LabelStrategyRole
B0Fixed recipe: two optical bands every N hours, a spectrum once brighter than a threshold.Human-style practice.
B1Reactive nearest-capable dispatcher (the existing project baseline).Shows what the pre-existing system does.
B2Classifier-entropy sampling: measure whatever most reduces predicted-class entropy; no physics likelihoods.Active-learning-style baseline.
B3Closed-world expected information gain, centralized, no decay term, no h₀.POIE-style baseline — the key comparison.
PFull method: h₀ + decay + stale-belief auction.Proposed.
OCentralized, perfect communication, full method.Upper bound.

6.4 Metrics

  • Time to identification — hours until the posterior on the true class exceeds 0.95, censored at the event's detectability limit.
  • Accuracy at deadline — correct top class and log score at a fixed horizon (e.g. 72 h).
  • Early-time information recovered — posterior contraction of each class's physical parameters (e.g. ejecta mass, explosion time) relative to prior; the knowledge metric.
  • Open-world detection — area under the ROC curve of P(h₀) separating withheld from modeled events; false-alarm rate on modeled events at the chosen τ.
  • Preservation — for withheld events, fraction of the event's detectable lifetime covered in at least three wavelength regions.
  • Cost — measurements used, slew time, messages sent, bytes exchanged.

6.5 Scenarios

  1. Nominal — full constellation, contact fraction as propagated.
  2. Communication sweep — contacts artificially thinned from 100% down to 10% of the propagated schedule.
  3. Node loss — one or two spacecraft removed mid-campaign.
  4. Crowding — several events competing at once for the same instruments.
  5. Model error — forward models perturbed from the injected truth by realistic systematics, testing h₀ false alarms.

6.6 Ablations

Each ablation removes exactly one term so a reader can see which part earns its keep: drop h₀ entirely; drop the decay term (γ = 0); drop the inadequacy term while keeping h₀ in the belief (β = 0); replace the auction with independent greedy planners; vary K.

6.7 Statistics

At least 1,000 injected events per scenario, fixed seeds shared across strategies. Strategies are compared with paired tests (Wilcoxon signed-rank on per-event metrics), medians reported with bootstrap 95% intervals, and multiple comparisons corrected (Holm). Effect sizes are reported alongside p-values, not in place of them. The mode-switch threshold τ and the utility weights β, γ are tuned on a validation split of events, never on the test split.

7Real-data validation

Archival data cannot answer what a different measurement choice would have shown, since only the measurements actually taken exist. Real public light curves (e.g. from ZTF via a broker such as Fink or ALeRCE) are used for two narrower purposes only: checking that the forward models and noise model match reality, and a retrospective replay on densely observed events, where the planner may choose only among epochs that genuinely exist. This limitation is stated plainly rather than implied away.

8Hardware-in-the-loop testbed

The decision layer is replicated on physical hardware to show it survives real asynchronous radios and microcontroller-class compute, with results close to simulation — not to reproduce flight-qualified radiation tolerance, real RF link budgets, attitude control or flight software, which the testbed explicitly does not claim.

Hardware-in-the-loop testbed topology Several ESP32 nodes linked by ESP-NOW radio, with a Raspberry Pi acting as sky oracle and ground station. ESP32 node A ESP32 node B ESP32 node C Raspberry Pi sky oracle / ground belief + bid logic belief + bid logic belief + bid logic

Fig. 5 · Each node runs the planner's belief and bidding logic on-device; links blink independently to stand in for intermittent, lossy ESP-NOW contact.

Topology sketch of the planned testbed — link blinking is illustrative, not a captured radio trace.

Table 4 · Testbed elements
ElementHardwareConstraint reproduced
Spacecraft nodes (6–8)ESP32 boards running belief update, utility and bidding on-device.Limited compute and memory; tests whether the planner fits a microcontroller-class budget.
CrosslinksESP-NOW radio, gated on and off by the OrbitalNet contact schedule replayed in real time (or accelerated).Intermittent links, message loss, asynchrony.
ClocksIndependent node clocks, synchronized only via messages.Clock drift and ordering errors.
Sky oracle & groundRaspberry Pi serving measurement outcomes from the simulated universe; acts as ground station.Separates ground truth from the nodes' own beliefs.
FaultsScheduled power-cycling of nodes.Node loss and rejoin.

9Expected results, limitations & validity

Nothing in this section is demonstrated; these are the predictions the experiments above are built to confirm or refute.

9.1 Expected results

  • P is expected to beat B0–B2 clearly on time to identification, since none of them use physics likelihoods to choose discriminating bands.
  • P versus B3 is the real test: expected similar performance on modeled classes, better early-time parameter recovery for fast classes (the decay term), and much better behaviour on withheld classes, where B3 is expected to confidently misclassify them.
  • A plausible failure mode: h₀ raises false alarms when forward models are imperfect (Scenario 5). If so, the false-alarm/detection trade-off curve is itself a useful, reportable result.
  • The auction is expected to approach the centralized oracle (O) at high contact fractions and degrade toward independent-greedy at low ones; where the crossover sits is unknown and worth reporting either way.

9.2 Limitations

  • Hypothesis sets are small and hand-chosen; real follow-up faces dozens of classes and subclasses.
  • Forward models are simplified; systematics between templates and reality can masquerade as "unmodeled."
  • No real spacecraft constellation today carries this payload mix; the constellation is a design study.
  • The "none of the above" hypothesis detects inadequacy; it does not invent a new explanation. The system participates in discovery by preserving evidence and flagging it, not by theorising.
  • Computing mutual information in real time on microcontrollers may require coarse approximations that change rankings.
Table 5 · Threats to validity
TypeThreatMitigation
InternalTuning weights on the test events.Separate validation set; fixed seeds; metrics pre-registered before final runs.
InternalBaselines implemented weakly.B3 receives the same likelihoods and budget as P; baselines tuned as carefully as P.
ConstructPosterior contraction may not equal "scientific knowledge."Also report accuracy of recovered physical parameters against injected truth.
ExternalSimulator realism.Validate noise and light curves against real broker data; state which physics is omitted.
ExternalHardware testbed does not generalise to flight.Claim limited to decision-layer behaviour under asynchrony and compute limits.
ConclusionMany pairwise comparisons.Holm correction; effect sizes; confidence intervals.

10Reproducibility & future work

The release will include the simulator, class forward models (or scripts that fetch them), orbital elements, contact schedules, every strategy's code, the analysis notebooks producing each figure, ESP32 firmware and the wiring diagram, under an open licence with a tagged, DOI-archived version.

Planned extensions: scaling to larger hypothesis libraries; shadow-mode integration with a real alert stream alongside a partner follow-up network (planner advises, humans act); learned, amortized design policies to reduce onboard compute; inadequacy-driven onboard retention as a second paper; and, for an eventual flight path, radiation-tolerant compute, real link budgets and a hosted-payload or rideshare demonstration.

RReferences

These entries were located by web search on 24 September 2026; only abstracts or summaries were read while drafting, not full texts. They are leads, not confirmed citations — each must be read in full and its author list verified before use in any submitted version.

  1. JPL POISE project page and the POISE i-SAIRAS 2020 paper.
  2. Guirguis et al., POMDPs for Autonomous Science Exploration (SHM-POMDP), arXiv:2608.03155; cites Candela et al., IROS 2017, on Science Hypothesis Maps.
  3. NASA: What Is Distributed Spacecraft Autonomy?, and the DSA Starling 1.0 results (SmallSat).
  4. An information-theoretic approach to deciding spectroscopic follow-ups (Astudillo et al., AJ 2020).
  5. Kennamer et al., active learning with RESSPECT.
  6. Real-time active learning with the Fink broker, PASA 2025.
  7. NMMA-Astro-COLIBRI, arXiv:2608.17568.
  8. Improving robustness to model misspecification in Bayesian experimental design (OpenReview).
  9. Metrics for Bayesian optimal experiment design under model misspecification, arXiv:2304.07949.
  10. Robust experimental design via generalised Bayesian inference (GBOED), arXiv:2511.07671.
  11. GRANDMA observations of ZTF/Fink transients; TOM Toolkit and AEON.
  12. On-orbit space AI survey, arXiv:2604.16518.
  13. Real-time detection of anomalies in large-scale transient surveys, MNRAS 2022.
  14. Krause & Guestrin on submodularity of information gain (still to locate and cite directly).
  15. The consensus-based bundle algorithm (still to locate and cite directly).
  16. ASE/EO-1 autonomous science papers, Chien et al. (still to locate and cite directly).
  17. The specific light-curve model paper for each injected class (still to select and cite per class).