[IMP] maintenance_service_http_monitoring : speed up KO confirmation
Some checks failed
pre-commit / pre-commit (pull_request) Has been cancelled

The previous design required two passes of the 10-minute sweep before a
maintenance request was created, so an outage took at least ~10 extra
minutes (up to 20) to be reported.

Split the monitoring into two crons with distinct responsibilities:

- the 10-minute sweep only discovers outages, flags the KO streak via
  ``http_first_ko_at`` and auto-resolves recovered services;
- a new 2-minute confirmation cron re-checks only the flagged services
  and creates a request once the service has been continuously KO for
  ``HTTP_KO_CONFIRMATION_DELAY`` (lowered from 5 to 2 minutes).

This keeps the transient-outage filter while removing the extra sweep
latency. Services already under an open request are skipped by the
confirmation cron and remain handled by the sweep.
This commit is contained in:
Stéphan Sainléger
2026-09-18 23:13:26 +02:00
parent e48a3443f1
commit 0e32152278
4 changed files with 130 additions and 29 deletions

View File

@@ -39,13 +39,19 @@ By default, maintenance mode lasts 4 hours. To change this:
## Cron Jobs ## Cron Jobs
Two scheduled actions are installed: Three scheduled actions are installed:
1. **HTTP Service Monitoring: check all services** 1. **HTTP Service Monitoring: check all services**
- Runs every 15 minutes - Runs every 10 minutes
- Checks HTTP status of all active service instances with URLs - Discovery sweep: checks HTTP status of all active service instances with URLs
and flags the start of a KO streak
2. **HTTP Service Monitoring: deactivate expired maintenance mode** 2. **HTTP Service Monitoring: confirm KO services**
- Runs every 2 minutes
- Re-checks only the services currently flagged KO (without an open request) and
creates a maintenance request once the outage is confirmed
3. **HTTP Service Monitoring: deactivate expired maintenance mode**
- Runs every 15 minutes - Runs every 15 minutes
- Automatically disables maintenance mode when the end time is reached - Automatically disables maintenance mode when the end time is reached
@@ -98,14 +104,18 @@ On service instances, you can see:
## Automatic Maintenance Requests ## Automatic Maintenance Requests
When a service fails HTTP checks: When a service fails HTTP checks:
- The 10-minute discovery sweep flags the outage (``http_first_ko_at``) but does
**not** create a request yet
- The 2-minute confirmation cron re-checks the flagged services only. A request is
created once the service has been continuously KO for at least 2 minutes
(``HTTP_KO_CONFIRMATION_DELAY``), so a short transient outage is not flagged
- A corrective maintenance request is created per failing service, named - A corrective maintenance request is created per failing service, named
``[HTTP KO] {service_url}`` ``[HTTP KO] {service_url}``
- The request description includes the error detail: the HTTP status code, - The request description includes the error detail: the HTTP status code,
or a network error label (timeout / DNS / SSL) when no HTTP response was received or a network error label (timeout / DNS / SSL) when no HTTP response was received
- No duplicate is created as long as an open request already exists for that service - No duplicate is created as long as an open request already exists for that service
- A **double-check** is performed before creating the request: the service is retested - The confirmation cron skips services that already have an open request; those are
after 2 seconds. A maintenance request is only created if the service fails **both** handled by the discovery sweep
checks, reducing noise from transient HTTP errors
When a service recovers (returns HTTP 200 after having an open request): When a service recovers (returns HTTP 200 after having an open request):
- The open maintenance request is automatically moved to the first **done** stage - The open maintenance request is automatically moved to the first **done** stage

View File

@@ -8,6 +8,14 @@
<field name="interval_number">10</field> <field name="interval_number">10</field>
<field name="interval_type">minutes</field> <field name="interval_type">minutes</field>
</record> </record>
<record id="ir_cron_http_service_confirmation" model="ir.cron">
<field name="name">HTTP Service Monitoring : confirm KO services</field>
<field name="model_id" ref="maintenance_server_data.model_service_instance" />
<field name="state">code</field>
<field name="code">model.cron_confirm_http_ko_services()</field>
<field name="interval_number">2</field>
<field name="interval_type">minutes</field>
</record>
<record id="ir_cron_maintenance_mode_expiry" model="ir.cron"> <record id="ir_cron_maintenance_mode_expiry" model="ir.cron">
<field <field
name="name" name="name"

View File

@@ -11,7 +11,7 @@ except ImportError:
_logger = logging.getLogger(__name__) _logger = logging.getLogger(__name__)
HTTP_CHECK_TIMEOUT = 20 # seconds HTTP_CHECK_TIMEOUT = 20 # seconds
HTTP_KO_CONFIRMATION_DELAY = timedelta(minutes=5) HTTP_KO_CONFIRMATION_DELAY = timedelta(minutes=2)
class ServiceInstance(models.Model): class ServiceInstance(models.Model):
@@ -46,7 +46,9 @@ class ServiceInstance(models.Model):
Perform HTTP check for each record and return the KO recordset. Perform HTTP check for each record and return the KO recordset.
Writes last_http_status_code, last_http_check_date and http_status_ok on every Writes last_http_status_code, last_http_check_date and http_status_ok on every
checked record. Does NOT create maintenance.request — the cron only opens one checked record, and maintains http_first_ko_at (the start of the current KO
streak, reset as soon as the service is OK again). Does NOT create
maintenance.request — that decision belongs to cron_confirm_http_ko_services,
once the service has been continuously KO for at least once the service has been continuously KO for at least
HTTP_KO_CONFIRMATION_DELAY, to avoid flagging transient outages (e.g. a short HTTP_KO_CONFIRMATION_DELAY, to avoid flagging transient outages (e.g. a short
server overload) as real incidents. server overload) as real incidents.
@@ -115,12 +117,12 @@ class ServiceInstance(models.Model):
@api.model @api.model
def cron_check_http_services(self): def cron_check_http_services(self):
""" """
Check all active services with a URL. Discovery sweep: check all active services with a URL.
A service must be continuously KO for at least HTTP_KO_CONFIRMATION_DELAY This cron only detects outages — it flags the start of a KO streak through
before a maintenance.request is created — this tolerates transient outages http_first_ko_at (via check_http_status) and auto-resolves services that
(e.g. a temporary server overload) regardless of how often this cron runs. recovered. Maintenance requests are created by the separate, faster
Services that had an open request and are now OK are auto-resolved. cron_confirm_http_ko_services once the outage is confirmed.
""" """
domain = [ domain = [
("active", "=", True), ("active", "=", True),
@@ -137,13 +139,39 @@ class ServiceInstance(models.Model):
and not s.http_maintenance_request.stage_id.done and not s.http_maintenance_request.stage_id.done
) )
ko_services = services.check_http_status() services.check_http_status()
# Auto-resolve services that recovered # Auto-resolve services that recovered
recovered = services_with_open_request.filtered(lambda s: s.http_status_ok) recovered = services_with_open_request.filtered(lambda s: s.http_status_ok)
if recovered: if recovered:
recovered._close_http_maintenance_request() recovered._close_http_maintenance_request()
@api.model
def cron_confirm_http_ko_services(self):
"""
Confirmation pass: re-check services currently flagged KO and create their
maintenance.request once the outage is confirmed.
Only services without an open request are considered (re-checking a service
already under an open request is pointless). A service must be continuously
KO for at least HTTP_KO_CONFIRMATION_DELAY: a service that recovered in the
meantime resets http_first_ko_at via check_http_status and no request is
created. Services already under an open request are handled by the slower
discovery cron, which also auto-resolves them when they recover.
"""
domain = [
("active", "=", True),
("service_url", "!=", False),
("equipment_id", "!=", False),
("http_first_ko_at", "!=", False),
("http_maintenance_request", "=", False),
]
services = self.search(domain).filtered(
lambda s: not s.equipment_id.maintenance_mode
)
ko_services = services.check_http_status()
confirmed_ko = ko_services.filtered( confirmed_ko = ko_services.filtered(
lambda s: s.last_http_check_date - s.http_first_ko_at lambda s: s.last_http_check_date - s.http_first_ko_at
>= HTTP_KO_CONFIRMATION_DELAY >= HTTP_KO_CONFIRMATION_DELAY

View File

@@ -12,8 +12,8 @@ EQUIPMENT_HTTP_REQUESTS = (
".models.maintenance_equipment.http_requests" ".models.maintenance_equipment.http_requests"
) )
# Comfortably past HTTP_KO_CONFIRMATION_DELAY (currently 10 minutes). # Comfortably past HTTP_KO_CONFIRMATION_DELAY (currently 2 minutes).
PAST_CONFIRMATION_DELAY = timedelta(minutes=20) PAST_CONFIRMATION_DELAY = timedelta(minutes=10)
def _mock_response(status_code): def _mock_response(status_code):
@@ -42,7 +42,7 @@ class TestHttpMonitoring(TransactionCase):
) )
def _backdate_first_ko(self, *service_instances): def _backdate_first_ko(self, *service_instances):
"""Simulate that the current KO streak started long enough ago to be confirmed.""" """Simulate a KO streak old enough ago to be confirmed."""
for service_instance in service_instances: for service_instance in service_instances:
service_instance.write( service_instance.write(
{"http_first_ko_at": fields.Datetime.now() - PAST_CONFIRMATION_DELAY} {"http_first_ko_at": fields.Datetime.now() - PAST_CONFIRMATION_DELAY}
@@ -81,7 +81,7 @@ class TestHttpMonitoring(TransactionCase):
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests: with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(500) mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services() self.env["service.instance"].cron_confirm_http_ko_services()
request = self.service_instance.http_maintenance_request request = self.service_instance.http_maintenance_request
self.assertTrue(request) self.assertTrue(request)
@@ -118,16 +118,19 @@ class TestHttpMonitoring(TransactionCase):
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests: with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(500) mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services() # confirmed, request created # confirmed -> request created
self.env["service.instance"].cron_confirm_http_ko_services()
request_1 = self.service_instance.http_maintenance_request request_1 = self.service_instance.http_maintenance_request
self.assertTrue(request_1) self.assertTrue(request_1)
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests: with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services() # still KO # already handled -> skipped
self.env["service.instance"].cron_confirm_http_ko_services()
# Service already has an open request -> the confirmation cron skips it
mock_requests.get.assert_not_called()
self.assertEqual(self.service_instance.http_maintenance_request, request_1) self.assertEqual(self.service_instance.http_maintenance_request, request_1)
self.assertEqual( self.assertEqual(
self.env["maintenance.request"].search_count( self.env["maintenance.request"].search_count(
@@ -222,7 +225,7 @@ class TestHttpMonitoring(TransactionCase):
): ):
mock_requests.get.return_value = _mock_response(500) mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services() self.env["service.instance"].cron_confirm_http_ko_services()
mock_http.post.assert_called_once() mock_http.post.assert_called_once()
call_kwargs = mock_http.post.call_args call_kwargs = mock_http.post.call_args
@@ -249,7 +252,7 @@ class TestHttpMonitoring(TransactionCase):
): ):
mock_requests.get.return_value = _mock_response(500) mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services() self.env["service.instance"].cron_confirm_http_ko_services()
self.assertTrue(self.service_instance.http_maintenance_request) self.assertTrue(self.service_instance.http_maintenance_request)
mock_http.post.assert_not_called() mock_http.post.assert_not_called()
@@ -283,7 +286,7 @@ class TestHttpMonitoring(TransactionCase):
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests: with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(200) mock_requests.get.return_value = _mock_response(200)
mock_requests.exceptions.RequestException = Exception mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services() self.env["service.instance"].cron_confirm_http_ko_services()
self.assertTrue(self.service_instance.http_status_ok) self.assertTrue(self.service_instance.http_status_ok)
self.assertEqual(self.service_instance.last_http_status_code, 200) self.assertEqual(self.service_instance.last_http_status_code, 200)
@@ -314,7 +317,7 @@ class TestHttpMonitoring(TransactionCase):
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests: with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(503) mock_requests.get.return_value = _mock_response(503)
mock_requests.exceptions.RequestException = Exception mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services() self.env["service.instance"].cron_confirm_http_ko_services()
self.assertEqual(mock_requests.get.call_count, 1) self.assertEqual(mock_requests.get.call_count, 1)
self.assertFalse(self.service_instance.http_status_ok) self.assertFalse(self.service_instance.http_status_ok)
@@ -345,7 +348,7 @@ class TestHttpMonitoring(TransactionCase):
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests: with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(500) mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services() self.env["service.instance"].cron_confirm_http_ko_services()
req1 = self.service_instance.http_maintenance_request req1 = self.service_instance.http_maintenance_request
req2 = service_instance2.http_maintenance_request req2 = service_instance2.http_maintenance_request
@@ -374,11 +377,11 @@ class TestHttpMonitoring(TransactionCase):
self._backdate_first_ko(self.service_instance) self._backdate_first_ko(self.service_instance)
# Second cron run: streak confirmed -> request created # Confirmation cron run: streak confirmed -> request created
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests: with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(500) mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services() self.env["service.instance"].cron_confirm_http_ko_services()
request = self.service_instance.http_maintenance_request request = self.service_instance.http_maintenance_request
self.assertTrue(request) self.assertTrue(request)
@@ -411,3 +414,55 @@ class TestHttpMonitoring(TransactionCase):
# Calling close directly must not raise # Calling close directly must not raise
self.service_instance._close_http_maintenance_request() self.service_instance._close_http_maintenance_request()
self.assertFalse(self.service_instance.http_maintenance_request) self.assertFalse(self.service_instance.http_maintenance_request)
# ------------------------------------------------------------------
# Test 17 -- Confirmation cron ignores services without a KO streak
# ------------------------------------------------------------------
def test_confirmation_cron_ignores_unflagged_services(self):
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(200)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_confirm_http_ko_services()
mock_requests.get.assert_not_called()
self.assertFalse(self.service_instance.http_maintenance_request)
# ------------------------------------------------------------------
# Test 18 -- Confirmation cron does not create a request before
# HTTP_KO_CONFIRMATION_DELAY has elapsed
# ------------------------------------------------------------------
def test_confirmation_cron_waits_for_confirmation_delay(self):
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services()
self.assertTrue(self.service_instance.http_first_ko_at)
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_confirm_http_ko_services()
self.assertFalse(self.service_instance.http_maintenance_request)
# ------------------------------------------------------------------
# Test 19 -- Confirmation cron ignores services whose equipment is in
# maintenance mode
# ------------------------------------------------------------------
def test_confirmation_cron_skips_maintenance_mode(self):
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services()
self.assertTrue(self.service_instance.http_first_ko_at)
self.equipment.write({"maintenance_mode": True})
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_confirm_http_ko_services()
mock_requests.get.assert_not_called()
self.assertFalse(self.service_instance.http_maintenance_request)