[IMP] maintenance_service_http_monitoring : speed up KO confirmation #10

Merged
stephansainleger merged 1 commits from 18.0-improve-http-tests into 18.0 2026-09-18 21:17:05 +00:00
4 changed files with 130 additions and 29 deletions

View File

@@ -39,13 +39,19 @@ By default, maintenance mode lasts 4 hours. To change this:
## Cron Jobs
Two scheduled actions are installed:
Three scheduled actions are installed:
1. **HTTP Service Monitoring: check all services**
- Runs every 15 minutes
- Checks HTTP status of all active service instances with URLs
- Runs every 10 minutes
- Discovery sweep: checks HTTP status of all active service instances with URLs
and flags the start of a KO streak
2. **HTTP Service Monitoring: deactivate expired maintenance mode**
2. **HTTP Service Monitoring: confirm KO services**
- Runs every 2 minutes
- Re-checks only the services currently flagged KO (without an open request) and
creates a maintenance request once the outage is confirmed
3. **HTTP Service Monitoring: deactivate expired maintenance mode**
- Runs every 15 minutes
- Automatically disables maintenance mode when the end time is reached
@@ -98,14 +104,18 @@ On service instances, you can see:
## Automatic Maintenance Requests
When a service fails HTTP checks:
- The 10-minute discovery sweep flags the outage (``http_first_ko_at``) but does
**not** create a request yet
- The 2-minute confirmation cron re-checks the flagged services only. A request is
created once the service has been continuously KO for at least 2 minutes
(``HTTP_KO_CONFIRMATION_DELAY``), so a short transient outage is not flagged
- A corrective maintenance request is created per failing service, named
``[HTTP KO] {service_url}``
- The request description includes the error detail: the HTTP status code,
or a network error label (timeout / DNS / SSL) when no HTTP response was received
- No duplicate is created as long as an open request already exists for that service
- A **double-check** is performed before creating the request: the service is retested
after 2 seconds. A maintenance request is only created if the service fails **both**
checks, reducing noise from transient HTTP errors
- The confirmation cron skips services that already have an open request; those are
handled by the discovery sweep
When a service recovers (returns HTTP 200 after having an open request):
- The open maintenance request is automatically moved to the first **done** stage

View File

@@ -8,6 +8,14 @@
<field name="interval_number">10</field>
<field name="interval_type">minutes</field>
</record>
<record id="ir_cron_http_service_confirmation" model="ir.cron">
<field name="name">HTTP Service Monitoring : confirm KO services</field>
<field name="model_id" ref="maintenance_server_data.model_service_instance" />
<field name="state">code</field>
<field name="code">model.cron_confirm_http_ko_services()</field>
<field name="interval_number">2</field>
<field name="interval_type">minutes</field>
</record>
<record id="ir_cron_maintenance_mode_expiry" model="ir.cron">
<field
name="name"

View File

@@ -11,7 +11,7 @@ except ImportError:
_logger = logging.getLogger(__name__)
HTTP_CHECK_TIMEOUT = 20 # seconds
HTTP_KO_CONFIRMATION_DELAY = timedelta(minutes=5)
HTTP_KO_CONFIRMATION_DELAY = timedelta(minutes=2)
class ServiceInstance(models.Model):
@@ -46,7 +46,9 @@ class ServiceInstance(models.Model):
Perform HTTP check for each record and return the KO recordset.
Writes last_http_status_code, last_http_check_date and http_status_ok on every
checked record. Does NOT create maintenance.request — the cron only opens one
checked record, and maintains http_first_ko_at (the start of the current KO
streak, reset as soon as the service is OK again). Does NOT create
maintenance.request — that decision belongs to cron_confirm_http_ko_services,
once the service has been continuously KO for at least
HTTP_KO_CONFIRMATION_DELAY, to avoid flagging transient outages (e.g. a short
server overload) as real incidents.
@@ -115,12 +117,12 @@ class ServiceInstance(models.Model):
@api.model
def cron_check_http_services(self):
"""
Check all active services with a URL.
Discovery sweep: check all active services with a URL.
A service must be continuously KO for at least HTTP_KO_CONFIRMATION_DELAY
before a maintenance.request is created — this tolerates transient outages
(e.g. a temporary server overload) regardless of how often this cron runs.
Services that had an open request and are now OK are auto-resolved.
This cron only detects outages — it flags the start of a KO streak through
http_first_ko_at (via check_http_status) and auto-resolves services that
recovered. Maintenance requests are created by the separate, faster
cron_confirm_http_ko_services once the outage is confirmed.
"""
domain = [
("active", "=", True),
@@ -137,13 +139,39 @@ class ServiceInstance(models.Model):
and not s.http_maintenance_request.stage_id.done
)
ko_services = services.check_http_status()
services.check_http_status()
# Auto-resolve services that recovered
recovered = services_with_open_request.filtered(lambda s: s.http_status_ok)
if recovered:
recovered._close_http_maintenance_request()
@api.model
def cron_confirm_http_ko_services(self):
"""
Confirmation pass: re-check services currently flagged KO and create their
maintenance.request once the outage is confirmed.
Only services without an open request are considered (re-checking a service
already under an open request is pointless). A service must be continuously
KO for at least HTTP_KO_CONFIRMATION_DELAY: a service that recovered in the
meantime resets http_first_ko_at via check_http_status and no request is
created. Services already under an open request are handled by the slower
discovery cron, which also auto-resolves them when they recover.
"""
domain = [
("active", "=", True),
("service_url", "!=", False),
("equipment_id", "!=", False),
("http_first_ko_at", "!=", False),
("http_maintenance_request", "=", False),
]
services = self.search(domain).filtered(
lambda s: not s.equipment_id.maintenance_mode
)
ko_services = services.check_http_status()
confirmed_ko = ko_services.filtered(
lambda s: s.last_http_check_date - s.http_first_ko_at
>= HTTP_KO_CONFIRMATION_DELAY

View File

@@ -12,8 +12,8 @@ EQUIPMENT_HTTP_REQUESTS = (
".models.maintenance_equipment.http_requests"
)
# Comfortably past HTTP_KO_CONFIRMATION_DELAY (currently 10 minutes).
PAST_CONFIRMATION_DELAY = timedelta(minutes=20)
# Comfortably past HTTP_KO_CONFIRMATION_DELAY (currently 2 minutes).
PAST_CONFIRMATION_DELAY = timedelta(minutes=10)
def _mock_response(status_code):
@@ -42,7 +42,7 @@ class TestHttpMonitoring(TransactionCase):
)
def _backdate_first_ko(self, *service_instances):
"""Simulate that the current KO streak started long enough ago to be confirmed."""
"""Simulate a KO streak old enough ago to be confirmed."""
for service_instance in service_instances:
service_instance.write(
{"http_first_ko_at": fields.Datetime.now() - PAST_CONFIRMATION_DELAY}
@@ -81,7 +81,7 @@ class TestHttpMonitoring(TransactionCase):
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services()
self.env["service.instance"].cron_confirm_http_ko_services()
request = self.service_instance.http_maintenance_request
self.assertTrue(request)
@@ -118,16 +118,19 @@ class TestHttpMonitoring(TransactionCase):
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services() # confirmed, request created
# confirmed -> request created
self.env["service.instance"].cron_confirm_http_ko_services()
request_1 = self.service_instance.http_maintenance_request
self.assertTrue(request_1)
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services() # still KO
# already handled -> skipped
self.env["service.instance"].cron_confirm_http_ko_services()
# Service already has an open request -> the confirmation cron skips it
mock_requests.get.assert_not_called()
self.assertEqual(self.service_instance.http_maintenance_request, request_1)
self.assertEqual(
self.env["maintenance.request"].search_count(
@@ -222,7 +225,7 @@ class TestHttpMonitoring(TransactionCase):
):
mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services()
self.env["service.instance"].cron_confirm_http_ko_services()
mock_http.post.assert_called_once()
call_kwargs = mock_http.post.call_args
@@ -249,7 +252,7 @@ class TestHttpMonitoring(TransactionCase):
):
mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services()
self.env["service.instance"].cron_confirm_http_ko_services()
self.assertTrue(self.service_instance.http_maintenance_request)
mock_http.post.assert_not_called()
@@ -283,7 +286,7 @@ class TestHttpMonitoring(TransactionCase):
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(200)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services()
self.env["service.instance"].cron_confirm_http_ko_services()
self.assertTrue(self.service_instance.http_status_ok)
self.assertEqual(self.service_instance.last_http_status_code, 200)
@@ -314,7 +317,7 @@ class TestHttpMonitoring(TransactionCase):
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(503)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services()
self.env["service.instance"].cron_confirm_http_ko_services()
self.assertEqual(mock_requests.get.call_count, 1)
self.assertFalse(self.service_instance.http_status_ok)
@@ -345,7 +348,7 @@ class TestHttpMonitoring(TransactionCase):
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services()
self.env["service.instance"].cron_confirm_http_ko_services()
req1 = self.service_instance.http_maintenance_request
req2 = service_instance2.http_maintenance_request
@@ -374,11 +377,11 @@ class TestHttpMonitoring(TransactionCase):
self._backdate_first_ko(self.service_instance)
# Second cron run: streak confirmed -> request created
# Confirmation cron run: streak confirmed -> request created
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services()
self.env["service.instance"].cron_confirm_http_ko_services()
request = self.service_instance.http_maintenance_request
self.assertTrue(request)
@@ -411,3 +414,55 @@ class TestHttpMonitoring(TransactionCase):
# Calling close directly must not raise
self.service_instance._close_http_maintenance_request()
self.assertFalse(self.service_instance.http_maintenance_request)
# ------------------------------------------------------------------
# Test 17 -- Confirmation cron ignores services without a KO streak
# ------------------------------------------------------------------
def test_confirmation_cron_ignores_unflagged_services(self):
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(200)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_confirm_http_ko_services()
mock_requests.get.assert_not_called()
self.assertFalse(self.service_instance.http_maintenance_request)
# ------------------------------------------------------------------
# Test 18 -- Confirmation cron does not create a request before
# HTTP_KO_CONFIRMATION_DELAY has elapsed
# ------------------------------------------------------------------
def test_confirmation_cron_waits_for_confirmation_delay(self):
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services()
self.assertTrue(self.service_instance.http_first_ko_at)
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_confirm_http_ko_services()
self.assertFalse(self.service_instance.http_maintenance_request)
# ------------------------------------------------------------------
# Test 19 -- Confirmation cron ignores services whose equipment is in
# maintenance mode
# ------------------------------------------------------------------
def test_confirmation_cron_skips_maintenance_mode(self):
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_check_http_services()
self.assertTrue(self.service_instance.http_first_ko_at)
self.equipment.write({"maintenance_mode": True})
with patch(SERVICE_INSTANCE_REQUESTS) as mock_requests:
mock_requests.get.return_value = _mock_response(500)
mock_requests.exceptions.RequestException = Exception
self.env["service.instance"].cron_confirm_http_ko_services()
mock_requests.get.assert_not_called()
self.assertFalse(self.service_instance.http_maintenance_request)