Monitoring and Status Model
Monitoring objectives
For every cycle the system should answer:
- Did the cycle start?
- Are all required production files available?
- Were the files transferred?
- Was the DNT/XML submitted and was a response received?
- Was the delivery successfully ingested?
- Is the expected product independently visible remotely?
Existing monitoring direction
The initial monitor observes the existing production process and records lifecycle events such as job_started, job_stopped and monitor_error. It is independent of the upload script.
Polling every minute can miss short processes. A shorter systemd interval or explicit start/end event emission from the production script/wrapper is more reliable.
Normalized states
| State | Meaning |
|---|---|
| SCHEDULED | Expected cycle has not started |
| STARTED | Workflow started |
| WAITING_FILES | Waiting for the complete input set |
| FILES_READY | All 21 files are present |
| UPLOADING | Product transfer is active |
| FILES_UPLOADED | File upload phase completed |
| DNT_CREATED | Delivery notification generated |
| DNT_UPLOADED | Delivery notification submitted |
| WAITING_RESPONSE | Waiting for ingestion response |
| INGESTED | Response confirms Ingested=True |
| FAILED | Terminal error detected |
| DELAYED | Timing threshold exceeded |
| FALLBACK_ACTIVE | JUNO has taken over |
| DUPLICATE_RISK | More than one path may act on the cycle |
Event model
Every meaningful transition and important check should become an append-only event containing at least timestamp, cycle, source, event type, severity, message and optional details.
Sources can include HCMR, JUNO, coordinator and verifier.
Monitoring-side communication is best-effort. Failure to POST an event must never abort the production workflow.
Separate health dimensions
The dashboard should retain separate states for production-file readiness, transfer, DNT/notification, ingestion, public verification and HCMR/JUNO coordination instead of collapsing everything into one green/red status.