[IMP] maintenance_service_http_monitoring : speed up KO confirmation
Some checks failed
pre-commit / pre-commit (pull_request) Has been cancelled

The previous design required two passes of the 10-minute sweep before a
maintenance request was created, so an outage took at least ~10 extra
minutes (up to 20) to be reported.

Split the monitoring into two crons with distinct responsibilities:

- the 10-minute sweep only discovers outages, flags the KO streak via
  ``http_first_ko_at`` and auto-resolves recovered services;
- a new 2-minute confirmation cron re-checks only the flagged services
  and creates a request once the service has been continuously KO for
  ``HTTP_KO_CONFIRMATION_DELAY`` (lowered from 5 to 2 minutes).

This keeps the transient-outage filter while removing the extra sweep
latency. Services already under an open request are skipped by the
confirmation cron and remain handled by the sweep.
This commit is contained in:
Stéphan Sainléger
2026-09-18 23:13:26 +02:00
parent e48a3443f1
commit 0e32152278
4 changed files with 130 additions and 29 deletions

View File

@@ -11,7 +11,7 @@ except ImportError:
_logger = logging.getLogger(__name__)
HTTP_CHECK_TIMEOUT = 20 # seconds
HTTP_KO_CONFIRMATION_DELAY = timedelta(minutes=5)
HTTP_KO_CONFIRMATION_DELAY = timedelta(minutes=2)
class ServiceInstance(models.Model):
@@ -46,7 +46,9 @@ class ServiceInstance(models.Model):
Perform HTTP check for each record and return the KO recordset.
Writes last_http_status_code, last_http_check_date and http_status_ok on every
checked record. Does NOT create maintenance.request — the cron only opens one
checked record, and maintains http_first_ko_at (the start of the current KO
streak, reset as soon as the service is OK again). Does NOT create
maintenance.request — that decision belongs to cron_confirm_http_ko_services,
once the service has been continuously KO for at least
HTTP_KO_CONFIRMATION_DELAY, to avoid flagging transient outages (e.g. a short
server overload) as real incidents.
@@ -115,12 +117,12 @@ class ServiceInstance(models.Model):
@api.model
def cron_check_http_services(self):
"""
Check all active services with a URL.
Discovery sweep: check all active services with a URL.
A service must be continuously KO for at least HTTP_KO_CONFIRMATION_DELAY
before a maintenance.request is created — this tolerates transient outages
(e.g. a temporary server overload) regardless of how often this cron runs.
Services that had an open request and are now OK are auto-resolved.
This cron only detects outages — it flags the start of a KO streak through
http_first_ko_at (via check_http_status) and auto-resolves services that
recovered. Maintenance requests are created by the separate, faster
cron_confirm_http_ko_services once the outage is confirmed.
"""
domain = [
("active", "=", True),
@@ -137,13 +139,39 @@ class ServiceInstance(models.Model):
and not s.http_maintenance_request.stage_id.done
)
ko_services = services.check_http_status()
services.check_http_status()
# Auto-resolve services that recovered
recovered = services_with_open_request.filtered(lambda s: s.http_status_ok)
if recovered:
recovered._close_http_maintenance_request()
@api.model
def cron_confirm_http_ko_services(self):
"""
Confirmation pass: re-check services currently flagged KO and create their
maintenance.request once the outage is confirmed.
Only services without an open request are considered (re-checking a service
already under an open request is pointless). A service must be continuously
KO for at least HTTP_KO_CONFIRMATION_DELAY: a service that recovered in the
meantime resets http_first_ko_at via check_http_status and no request is
created. Services already under an open request are handled by the slower
discovery cron, which also auto-resolves them when they recover.
"""
domain = [
("active", "=", True),
("service_url", "!=", False),
("equipment_id", "!=", False),
("http_first_ko_at", "!=", False),
("http_maintenance_request", "=", False),
]
services = self.search(domain).filtered(
lambda s: not s.equipment_id.maintenance_mode
)
ko_services = services.check_http_status()
confirmed_ko = ko_services.filtered(
lambda s: s.last_http_check_date - s.http_first_ko_at
>= HTTP_KO_CONFIRMATION_DELAY