[IMP] maintenance_service_http_monitoring : speed up KO confirmation #10

Merged
stephansainleger merged 1 commits from 18.0-improve-http-tests into 18.0 2026-09-18 21:17:05 +00:00

The previous design required two passes of the 10-minute sweep before a
maintenance request was created, so an outage took at least ~10 extra
minutes (up to 20) to be reported.

Split the monitoring into two crons with distinct responsibilities:

  • the 10-minute sweep only discovers outages, flags the KO streak via
    http_first_ko_at and auto-resolves recovered services;
  • a new 2-minute confirmation cron re-checks only the flagged services
    and creates a request once the service has been continuously KO for
    HTTP_KO_CONFIRMATION_DELAY (lowered from 5 to 2 minutes).

This keeps the transient-outage filter while removing the extra sweep
latency. Services already under an open request are skipped by the
confirmation cron and remain handled by the sweep.

The previous design required two passes of the 10-minute sweep before a maintenance request was created, so an outage took at least ~10 extra minutes (up to 20) to be reported. Split the monitoring into two crons with distinct responsibilities: - the 10-minute sweep only discovers outages, flags the KO streak via ``http_first_ko_at`` and auto-resolves recovered services; - a new 2-minute confirmation cron re-checks only the flagged services and creates a request once the service has been continuously KO for ``HTTP_KO_CONFIRMATION_DELAY`` (lowered from 5 to 2 minutes). This keeps the transient-outage filter while removing the extra sweep latency. Services already under an open request are skipped by the confirmation cron and remain handled by the sweep.
stephansainleger added 1 commit 2026-09-18 21:14:19 +00:00
[IMP] maintenance_service_http_monitoring : speed up KO confirmation
Some checks failed
pre-commit / pre-commit (pull_request) Has been cancelled
0e32152278
The previous design required two passes of the 10-minute sweep before a
maintenance request was created, so an outage took at least ~10 extra
minutes (up to 20) to be reported.

Split the monitoring into two crons with distinct responsibilities:

- the 10-minute sweep only discovers outages, flags the KO streak via
  ``http_first_ko_at`` and auto-resolves recovered services;
- a new 2-minute confirmation cron re-checks only the flagged services
  and creates a request once the service has been continuously KO for
  ``HTTP_KO_CONFIRMATION_DELAY`` (lowered from 5 to 2 minutes).

This keeps the transient-outage filter while removing the extra sweep
latency. Services already under an open request are skipped by the
confirmation cron and remain handled by the sweep.
stephansainleger merged commit 0e32152278 into 18.0 2026-09-18 21:17:05 +00:00
stephansainleger deleted branch 18.0-improve-http-tests 2026-09-18 21:17:05 +00:00
Sign in to join this conversation.
No description provided.