Some checks failed
pre-commit / pre-commit (pull_request) Has been cancelled
The previous design required two passes of the 10-minute sweep before a maintenance request was created, so an outage took at least ~10 extra minutes (up to 20) to be reported. Split the monitoring into two crons with distinct responsibilities: - the 10-minute sweep only discovers outages, flags the KO streak via ``http_first_ko_at`` and auto-resolves recovered services; - a new 2-minute confirmation cron re-checks only the flagged services and creates a request once the service has been continuously KO for ``HTTP_KO_CONFIRMATION_DELAY`` (lowered from 5 to 2 minutes). This keeps the transient-outage filter while removing the extra sweep latency. Services already under an open request are skipped by the confirmation cron and remain handled by the sweep.