Operations and Troubleshooting
Operational success criteria
A cycle is not fully successful because the upload process exited or because the 21 files left the local server.
Verify in order:
- the correct cycle started;
- all 21 expected files became available;
- all intended transfers completed;
- the DNT/XML was created and uploaded;
- a DNT response was received;
- validation/ingestion succeeded;
- independent publication verification succeeds when available.
Quick diagnostic matrix
| Symptom | Likely area | Next action |
|---|---|---|
| 0 files / Files not found | upstream production | check production directory and model chain |
| Base directory not found / stale NFS | HCMR storage | restore mount; reassess cycle |
| Connection timed out | network or MDS FTP | retry within policy; escalate if TDT at risk |
| Checksum mismatch | transfer integrity | resend affected file(s) |
| Files not delivered | FTP transfer | recovery upload / CLEANER path |
| No DNT response | MDS processing | inspect response area and S3/CMT evidence |
| Ingested=False | MDS/application error | inspect error and recovery path |
| JUNO succeeds while copernicus01 keeps running | coordination race | stop duplicate path if appropriate; verify published data |
MDS-state decision
| Observed MDS state | Typical action |
|---|---|
| Correct 21 files for the intended bulletin | no recovery required |
| Fewer than expected / old bulletin | recovery/CLEANER workflow |
| Extra or obsolete forecast files | CLEANER or explicit DNT delete |
| Missing historical analysis files after multi-day incident | manual analysis recovery |
New run appears immediately completed
Likely cause: stale success state from a previous progress snapshot.
Monitoring must associate progress state with the current process lifetime and ignore state older than the current run.
Process stops without ingestion success
A stopped process is not proof of successful delivery. Inspect the DNT response and, if necessary, public MDS evidence.
Files do not reach 21/21
Keep the underlying state as WAITING_FILES and show the current count. If the operational threshold is exceeded, add a DELAYED condition without hiding the underlying state.
JUNO becomes active
Expose FALLBACK_ACTIVE and continue observing HCMR. If both nodes remain eligible to deliver the same cycle, raise DUPLICATE_RISK.
Monitoring is unavailable
Production continues. Monitoring failure is handled independently and must not alter the production success path.
TDT and incident escalation
The key NRT deadlines are:
- cycle 00: 20:00 UTC
- cycle 12: 06:00 UTC
If delivery cannot be restored within the operational window and TDT is at risk or missed, the event should move from local troubleshooting to the formal incident-reporting process.
Evidence priority
When dashboard information disagrees with production evidence, investigate in this order:
- DNT response / ingestion result;
- current-run production transfer logs and generated DNT;
- independent CMT/S3 publication evidence;
- normalized monitoring state;
- dashboard presentation.
The dashboard is a view of evidence, not the source of truth.