0e32152278032cfbbdf72c2138375fb37f20dd48
Some checks failed
pre-commit / pre-commit (pull_request) Has been cancelled
The previous design required two passes of the 10-minute sweep before a maintenance request was created, so an outage took at least ~10 extra minutes (up to 20) to be reported. Split the monitoring into two crons with distinct responsibilities: - the 10-minute sweep only discovers outages, flags the KO streak via ``http_first_ko_at`` and auto-resolves recovered services; - a new 2-minute confirmation cron re-checks only the flagged services and creates a request once the service has been continuously KO for ``HTTP_KO_CONFIRMATION_DELAY`` (lowered from 5 to 2 minutes). This keeps the transient-outage filter while removing the extra sweep latency. Services already under an open request are skipped by the confirmation cron and remain handled by the sweep.
Description
Languages
Python
85.9%
JavaScript
11.8%
Shell
2.3%