Skip to content

Monitoring and Status Model

Monitoring objectives

For every cycle the system should answer:

  1. Did the cycle start?
  2. Are all required production files available?
  3. Were the files transferred?
  4. Was the DNT/XML submitted and was a response received?
  5. Was the delivery successfully ingested?
  6. Is the expected product independently visible remotely?

Existing monitoring direction

The initial monitor observes the existing production process and records lifecycle events such as job_started, job_stopped and monitor_error. It is independent of the upload script.

Polling every minute can miss short processes. A shorter systemd interval or explicit start/end event emission from the production script/wrapper is more reliable.

Normalized states

State Meaning
SCHEDULED Expected cycle has not started
STARTED Workflow started
WAITING_FILES Waiting for the complete input set
FILES_READY All 21 files are present
UPLOADING Product transfer is active
FILES_UPLOADED File upload phase completed
DNT_CREATED Delivery notification generated
DNT_UPLOADED Delivery notification submitted
WAITING_RESPONSE Waiting for ingestion response
INGESTED Response confirms Ingested=True
FAILED Terminal error detected
DELAYED Timing threshold exceeded
FALLBACK_ACTIVE JUNO has taken over
DUPLICATE_RISK More than one path may act on the cycle

Event model

Every meaningful transition and important check should become an append-only event containing at least timestamp, cycle, source, event type, severity, message and optional details.

Sources can include HCMR, JUNO, coordinator and verifier.

Monitoring-side communication is best-effort. Failure to POST an event must never abort the production workflow.

Separate health dimensions

The dashboard should retain separate states for production-file readiness, transfer, DNT/notification, ingestion, public verification and HCMR/JUNO coordination instead of collapsing everything into one green/red status.