Skip to content

Operations and Troubleshooting

Operational success criteria

A cycle is not fully successful because the upload process exited or because the 21 files left the local server.

Verify in order:

  1. the correct cycle started;
  2. all 21 expected files became available;
  3. all intended transfers completed;
  4. the DNT/XML was created and uploaded;
  5. a DNT response was received;
  6. validation/ingestion succeeded;
  7. independent publication verification succeeds when available.

Quick diagnostic matrix

Symptom Likely area Next action
0 files / Files not found upstream production check production directory and model chain
Base directory not found / stale NFS HCMR storage restore mount; reassess cycle
Connection timed out network or MDS FTP retry within policy; escalate if TDT at risk
Checksum mismatch transfer integrity resend affected file(s)
Files not delivered FTP transfer recovery upload / CLEANER path
No DNT response MDS processing inspect response area and S3/CMT evidence
Ingested=False MDS/application error inspect error and recovery path
JUNO succeeds while copernicus01 keeps running coordination race stop duplicate path if appropriate; verify published data

MDS-state decision

Observed MDS state Typical action
Correct 21 files for the intended bulletin no recovery required
Fewer than expected / old bulletin recovery/CLEANER workflow
Extra or obsolete forecast files CLEANER or explicit DNT delete
Missing historical analysis files after multi-day incident manual analysis recovery

New run appears immediately completed

Likely cause: stale success state from a previous progress snapshot.

Monitoring must associate progress state with the current process lifetime and ignore state older than the current run.

Process stops without ingestion success

A stopped process is not proof of successful delivery. Inspect the DNT response and, if necessary, public MDS evidence.

Files do not reach 21/21

Keep the underlying state as WAITING_FILES and show the current count. If the operational threshold is exceeded, add a DELAYED condition without hiding the underlying state.

JUNO becomes active

Expose FALLBACK_ACTIVE and continue observing HCMR. If both nodes remain eligible to deliver the same cycle, raise DUPLICATE_RISK.

Monitoring is unavailable

Production continues. Monitoring failure is handled independently and must not alter the production success path.

TDT and incident escalation

The key NRT deadlines are:

  • cycle 00: 20:00 UTC
  • cycle 12: 06:00 UTC

If delivery cannot be restored within the operational window and TDT is at risk or missed, the event should move from local troubleshooting to the formal incident-reporting process.

Evidence priority

When dashboard information disagrees with production evidence, investigate in this order:

  1. DNT response / ingestion result;
  2. current-run production transfer logs and generated DNT;
  3. independent CMT/S3 publication evidence;
  4. normalized monitoring state;
  5. dashboard presentation.

The dashboard is a view of evidence, not the source of truth.