Skip to content

Planned Changes

Goal

The refactor must improve observability and coordination without silently changing the production semantics that already work.

The design is divided into three phases.

Phase A — common upload implementation

Phase A refactors the two NRT upload paths into:

  • shared medwav-common.sh logic;
  • node-specific protected configuration;
  • thin copernicus01 and JUNO orchestrators;
  • a dry-run/test harness;
  • a normalized structured status event function.

The functional baseline should remain equivalent to production: DNT content, exit behavior and operational notifications must be reproducible before coordination logic is changed.

Structured events

A proposed common event interface is:

emit_status(product, node, cycle, phase, result, detail, action_needed)

Events are written as JSONL and can feed three consumers:

  1. operational notification/email;
  2. web monitoring/dashboard;
  3. future coordinator heartbeat/state.

This avoids maintaining three unrelated interpretations of the same production state.

Decision hook

A single should_i_upload() decision point is introduced in Phase A but remains behaviorally equivalent to the current rules. Phase B can then replace the decision implementation without rewriting the entire transfer process.

Phase B — coordinator and monitoring

Phase B introduces:

  • a common HTTPS coordinator endpoint;
  • local agents/reporters;
  • normalized cycle ownership/state;
  • monitoring and event history;
  • fallback arbitration;
  • the fix for the copernicus01/JUNO race condition.

Authentication is planned around a protected pre-shared/HMAC-style mechanism rather than exposing shell access between the production nodes.

The coordinator remains advisory/fault-tolerant: loss of the coordinator must not automatically prevent production delivery.

flowchart TD
    CO[Coordinator<br/>Cycle State & Ownership]
    H[copernicus01]
    J[JUNO]
    M[(Event / State Store)]
    UI[Dashboard]
    N[Notifications]
    C[Copernicus Marine MDS]

    H -->|HTTPS status| CO
    J -->|HTTPS status| CO
    CO -->|should_i_upload decision| H
    CO -->|should_i_upload decision| J

    H --> C
    J --> C

    CO --> M
    M --> UI
    CO --> N

    H -.->|Coordinator unavailable:<br/>defined autonomous mode| C
    J -.->|Coordinator unavailable:<br/>defined autonomous mode| C

Phase C — upstream production integration

A later phase may extend observability into the production chain and wind/forcing dependencies. It is intentionally separated from the upload refactor.

Harness and regression strategy

The planned harness executes the real orchestrators while mocking transfer/mail/verification commands or targeting a non-production test endpoint.

Key goals:

  • compare generated DNT against a known-good golden reference;
  • test both nodes safely in parallel with production;
  • test retries, missing files, stale mounts and response behavior;
  • validate double-dissemination behavior before live deployment.

Double-dissemination readiness

The refactor treats datasets=(...) as a list rather than a single value. This is required for transition periods in which both the current and new dataset TAG must be delivered in the same production cycle.

The safer execution model is upload + DNT + response per dataset/TAG sequentially, isolating failure of one TAG from the other.

Change sequence

flowchart LR
    A[Lock current baseline] --> B[Shared library + config]
    B --> C[Dry-run harness]
    C --> D[Structured events]
    D --> E[Passive monitoring]
    E --> F[Alerts and evidence]
    F --> G[Coordinator / ownership]
    G --> H[Upstream integration]

No phase should make monitoring a hidden prerequisite for production.