Service outage on Tuesday, February 14 2023
Summary:
In the early morning of February 14, Chargetrip received a high volume of requests that started to back up our exchange broker queue. This resulted in our exchange broker seizing, a service slowdown, and, ultimately, a service outage. All customers were affected either with slower calculation times or a complete service outage. Chargetrip recovered all systems within three hours.
Impact:
All customers were affected by this outage by slow calculation times or denial of service. Affected services were: the routing engine, station database, vehicle database, tile service and go.chargetrip.com.
Timeline:
- Around 07:30 CET, we started receiving large volumes of messages that would ultimately seize our exchange broker.
- Around 09:00 CET, our internal warning system began sending alerts to our DevOps team.
- 09:28 We noticed our exchange broker was using an incredible amount of memory (1.4M messages in sync stations queue).
- Around 09:30 CET, our resources were increased to unblock our queue.
- 09:32 We scaled our sync daemon to 40 replicas.
- 09:33 We updated our status page.
- 09:34 We added more memory to our sync daemon.
- 09:40 We Scaled our sync daemon to 80 replicas to further distribute the load.
- 09:55 We increased the memory for our exchange broker and changed our replicas to 5 to further unblock queueing.
- At 09:55 CET, all systems were back online.
Contributing factors:
- Our Prometheus alerting system needed updated configurations and was slow to alert our engineers.
- Our monitoring system monitored calculation times and time deviations but failed to alert us of calculation errors.
Action items:
Prometheus and our monitor have been re-configured to detect memory outages and irregular volume. As a result, our recovery time for a similar incident should be below 20 minutes.