[IMP] maintenance_service_http_monitoring : speed up KO confirmation
Some checks failed
pre-commit / pre-commit (pull_request) Has been cancelled
Some checks failed
pre-commit / pre-commit (pull_request) Has been cancelled
The previous design required two passes of the 10-minute sweep before a maintenance request was created, so an outage took at least ~10 extra minutes (up to 20) to be reported. Split the monitoring into two crons with distinct responsibilities: - the 10-minute sweep only discovers outages, flags the KO streak via ``http_first_ko_at`` and auto-resolves recovered services; - a new 2-minute confirmation cron re-checks only the flagged services and creates a request once the service has been continuously KO for ``HTTP_KO_CONFIRMATION_DELAY`` (lowered from 5 to 2 minutes). This keeps the transient-outage filter while removing the extra sweep latency. Services already under an open request are skipped by the confirmation cron and remain handled by the sweep.
This commit is contained in:
@@ -39,13 +39,19 @@ By default, maintenance mode lasts 4 hours. To change this:
|
||||
|
||||
## Cron Jobs
|
||||
|
||||
Two scheduled actions are installed:
|
||||
Three scheduled actions are installed:
|
||||
|
||||
1. **HTTP Service Monitoring: check all services**
|
||||
- Runs every 15 minutes
|
||||
- Checks HTTP status of all active service instances with URLs
|
||||
- Runs every 10 minutes
|
||||
- Discovery sweep: checks HTTP status of all active service instances with URLs
|
||||
and flags the start of a KO streak
|
||||
|
||||
2. **HTTP Service Monitoring: deactivate expired maintenance mode**
|
||||
2. **HTTP Service Monitoring: confirm KO services**
|
||||
- Runs every 2 minutes
|
||||
- Re-checks only the services currently flagged KO (without an open request) and
|
||||
creates a maintenance request once the outage is confirmed
|
||||
|
||||
3. **HTTP Service Monitoring: deactivate expired maintenance mode**
|
||||
- Runs every 15 minutes
|
||||
- Automatically disables maintenance mode when the end time is reached
|
||||
|
||||
@@ -98,14 +104,18 @@ On service instances, you can see:
|
||||
## Automatic Maintenance Requests
|
||||
|
||||
When a service fails HTTP checks:
|
||||
- The 10-minute discovery sweep flags the outage (``http_first_ko_at``) but does
|
||||
**not** create a request yet
|
||||
- The 2-minute confirmation cron re-checks the flagged services only. A request is
|
||||
created once the service has been continuously KO for at least 2 minutes
|
||||
(``HTTP_KO_CONFIRMATION_DELAY``), so a short transient outage is not flagged
|
||||
- A corrective maintenance request is created per failing service, named
|
||||
``[HTTP KO] {service_url}``
|
||||
- The request description includes the error detail: the HTTP status code,
|
||||
or a network error label (timeout / DNS / SSL) when no HTTP response was received
|
||||
- No duplicate is created as long as an open request already exists for that service
|
||||
- A **double-check** is performed before creating the request: the service is retested
|
||||
after 2 seconds. A maintenance request is only created if the service fails **both**
|
||||
checks, reducing noise from transient HTTP errors
|
||||
- The confirmation cron skips services that already have an open request; those are
|
||||
handled by the discovery sweep
|
||||
|
||||
When a service recovers (returns HTTP 200 after having an open request):
|
||||
- The open maintenance request is automatically moved to the first **done** stage
|
||||
|
||||
Reference in New Issue
Block a user