Skip to content

Incident Response

Objective

The incident process separates three questions:

  1. Is production data available?
  2. Did HCMR/JUNO successfully deliver it?
  3. Did MDS ingest and publish it?

This prevents a transfer problem from being misdiagnosed as a model-production problem, and vice versa.

Decision flow

flowchart TD
    A[Delivery problem detected] --> B{21 expected files<br/>available locally?}
    B -->|No| U[Upstream / production issue]
    B -->|Yes| C{Uploader still running?}
    C -->|Yes| P[Inspect current phase / progress]
    C -->|No| M{MDS contains expected data?}

    P --> M
    M -->|Correct| DONE[No recovery required]
    M -->|Missing / partial| R[Recovery / CLEANER]
    M -->|Extra obsolete data| D[DNT Delete / cleanup]

    R --> X{Can delivery be restored<br/>before TDT?}
    D --> X
    X -->|Yes| V[Verify DNT response + S3/CMT]
    X -->|No| INC[Formal incident reporting]

Common failure classes

Upstream/model output unavailable

Symptoms include zero files or incomplete production directories. The upload layer cannot repair missing scientific output; escalate toward the production/upstream chain.

Storage/NFS problem

A stale or missing NetApp mount can make the uploader report that the base directory does not exist even though the model output exists elsewhere.

FTP/network transfer failure

Connection timeouts or repeated failed transfers require retry within the operational policy. If TDT becomes threatened, treat the problem as an operational incident.

Integrity failure

An MD5/checksum mismatch means the object must be re-transferred. A corrupted delivery must not be accepted as successful.

DNT response missing or negative

If the files were transferred but the response is missing, delayed or reports Ingested=False, inspect MDS evidence and the response error before deciding whether to retry the whole delivery or only a subset.

Neptune/JUNO overlap

If JUNO takes over and Neptune continues, verify the final MDS state and stop unnecessary duplicate activity where operationally appropriate.

Recovery mechanisms

The existing operational toolbox includes:

  • automated retry in the primary scripts;
  • JUNO fallback;
  • CLEANER-style recovery for partial or inconsistent MDS states;
  • manual analysis upload for historical gaps after multi-day incidents;
  • explicit DNT Delete for obsolete forecast files left behind after failures.

Recovery should always be followed by DNT-response validation and remote publication verification.

Reporting threshold

The NRT Target Delivery Times are:

  • cycle 00: 20:00 UTC;
  • cycle 12: 06:00 UTC.

A problem that cannot be resolved automatically/local-operationally before the relevant delivery commitment should enter the formal incident-reporting process.