Incident Response
Objective
The incident process separates three questions:
- Is production data available?
- Did HCMR/JUNO successfully deliver it?
- Did MDS ingest and publish it?
This prevents a transfer problem from being misdiagnosed as a model-production problem, and vice versa.
Decision flow
flowchart TD
A[Delivery problem detected] --> B{21 expected files<br/>available locally?}
B -->|No| U[Upstream / production issue]
B -->|Yes| C{Uploader still running?}
C -->|Yes| P[Inspect current phase / progress]
C -->|No| M{MDS contains expected data?}
P --> M
M -->|Correct| DONE[No recovery required]
M -->|Missing / partial| R[Recovery / CLEANER]
M -->|Extra obsolete data| D[DNT Delete / cleanup]
R --> X{Can delivery be restored<br/>before TDT?}
D --> X
X -->|Yes| V[Verify DNT response + S3/CMT]
X -->|No| INC[Formal incident reporting]
Common failure classes
Upstream/model output unavailable
Symptoms include zero files or incomplete production directories. The upload layer cannot repair missing scientific output; escalate toward the production/upstream chain.
Storage/NFS problem
A stale or missing NetApp mount can make the uploader report that the base directory does not exist even though the model output exists elsewhere.
FTP/network transfer failure
Connection timeouts or repeated failed transfers require retry within the operational policy. If TDT becomes threatened, treat the problem as an operational incident.
Integrity failure
An MD5/checksum mismatch means the object must be re-transferred. A corrupted delivery must not be accepted as successful.
DNT response missing or negative
If the files were transferred but the response is missing, delayed or reports Ingested=False, inspect MDS evidence and the response error before deciding whether to retry the whole delivery or only a subset.
Neptune/JUNO overlap
If JUNO takes over and Neptune continues, verify the final MDS state and stop unnecessary duplicate activity where operationally appropriate.
Recovery mechanisms
The existing operational toolbox includes:
- automated retry in the primary scripts;
- JUNO fallback;
- CLEANER-style recovery for partial or inconsistent MDS states;
- manual analysis upload for historical gaps after multi-day incidents;
- explicit DNT Delete for obsolete forecast files left behind after failures.
Recovery should always be followed by DNT-response validation and remote publication verification.
Reporting threshold
The NRT Target Delivery Times are:
- cycle 00: 20:00 UTC;
- cycle 12: 06:00 UTC.
A problem that cannot be resolved automatically/local-operationally before the relevant delivery commitment should enter the formal incident-reporting process.