Skip to content

JUNO Failover

Current role

JUNO is the CMCC-hosted backup path for MED-WAV. It has its own model output and a dedicated upload script, mds-medmfc-upload-v9-juno.sh.

The intended logic is:

flowchart TD
    J([JUNO cycle]) --> P[Read copernicus01 progress]
    P --> S{copernicus01 state?}
    S -->|reSuccess = 1| STOP([Stop: copernicus01 succeeded])
    S -->|nepSleep = 0| WAIT[Sleep 5 minutes]
    WAIT --> P
    S -->|nepSleep = 1| TAKE[JUNO takes over]
    S -->|Progress log unavailable| TAKE
    TAKE --> UP[Upload to MDS]

Current decision rules

The current JUNO logic interprets copernicus01 state approximately as follows:

copernicus01 state JUNO action
reSuccess=1 Stop; primary delivery succeeded
nepSleep=0 copernicus01 still running; wait 5 minutes and retry
nepSleep=1 without success Take over
copernicus01 progress cannot be retrieved Take over

Operational windows

Cycle 00 Cycle 12
copernicus01 idle/fallback window 18:00–24:00 UTC 02:00–10:00 UTC
JUNO wait window 20:00–02:00 next day 04:00–11:00 UTC
TDT 20:00 UTC 06:00 UTC

These windows are part of the legacy failover mechanism and should not be treated as a substitute for explicit cycle ownership.

Critical coordination limitation

The JUNO script currently has the lines that would publish its own progress log back to copernicus01 disabled/commented.

Therefore:

  • JUNO can observe copernicus01;
  • copernicus01 cannot reliably observe a successful JUNO takeover;
  • the intended bidirectional coordination is incomplete.

If JUNO takes over and copernicus01 later finds its files, Neptune may continue and upload the same cycle again. This can overwrite the same filenames and produce overlapping DNT activity.

This is the principal duplicate-delivery / race-condition risk in the current architecture.

Operational implication

Until explicit ownership is implemented, successful JUNO delivery should be confirmed through:

  1. JUNO operational evidence;
  2. DNT response;
  3. independent MDS/S3 validation.

A passive monitor should raise FALLBACK_ACTIVE and, if HCMR activity continues for the same cycle, DUPLICATE_RISK.

Security note

Legacy operational scripts contain embedded credentials. Secrets must not be copied into this repository. The refactor should move them to protected node-specific configuration or a secret-management mechanism.

Planned replacement

The v2 architecture replaces the progress-log-based failover with a coordinator-visible state model and a dedicated should_i_upload() decision point, while preserving autonomous operation if the coordinator is unavailable.