Citybreak Online

Incident report 2026-08-03

Information

INCIDENT REPORT CITYBREAK ONLINE

Systems affected
Citybreak Online

Customer/end-user impact
The incident affected all customers using the Citybreak Online platform in Production, with specific impact seen on the checkout process (basket and payment steps). During the event, the system was extremely slow, with some requests timing out or returning 503 errors. Internally, the impact included production server failures (out of capacity) and a high request queue buildup of approximately 1,000 requests. The customer-facing impact lasted for roughly 69 minutes, affecting users primarily during the booking and payment phases.

Timeline
2026-08-03, 10:15:34 Internal monitoring alert of increased response times in Citybreak Online.
2026-08-03, 10:30:28 Identification of high response times on the CBIS Export API and Server 03 failure.
2026-08-03 10:36:57 Customers are reporting the system is extremely slow and Hund is updated.
2026-08-03 10:37:57 Incident thread opened with alert that Citybreak Online is experiencing slow response times.
2026-08-03 10:37:57 Investigation is acknowledged as in progress.
2026-08-03 10:40:37 Tech reports approximately 1000 requests in queue, causing significant wait times.
2026-08-03 10:40:57 Tech informed that a bottleneck has been identified in a backend API and the team is working on a solution.
2026-08-03 10:42:00 An estimated resolution time of within the next hour is communicated.
2026-08-03 11:24:18 Tech reports servers were added and the service is back to normal.
2026-08-03 11:28:11 Tech reports more CPUs are being added to the machines to ensure stability.
2026-08-03 11:51:44 Tech reports the service has been stable for 30 minutes.

Root Cause Analysis
The root cause was an unusually high load on the CBIS Export API, driven by increased online traffic combined with a missing caching layer on the API. This high load led to server instability that cascaded across the infrastructure until additional resources were deployed.
Contributing factors included insufficient CPU capacity, a cascading failure triggered when one server failed and shifted load to others, and the lack of caching on high-traffic backend API endpoints.

Investigation/Resolution Actions
We identified high response times in Citybreak Online, driven by increased response times on the CBIS Export API and confirmed server failures (Server 03).
We attempted to disable the problematic API on HAProxy (which caused a further failure and was reverted), then reverted the load balancing strategy to leastconn.
We added extra server instances and increased CPU count on existing.
The team monitored metrics and queues, confirming 30 minutes of sustained stable operation.

Preventive Measures & Next Steps
Implement a caching layer for the CBIS Export API to better handle traffic spikes.
Audit CPU allocation for all CBIS machines and increase baseline infrastructure resources for core services.

Conclusion
On 2026-08-03, Citybreak Online in Production experienced slow response times beginning around 10:15, primarily affecting the checkout and payment experience for all platform customers. The incident was caused by unusually high load on the CBIS Export API — compounded by a missing caching layer and insufficient CPU capacity — which triggered a cascading server failure. Mitigation involved reverting load balancing to leastconn, adding server instances, and increasing CPU allocation. Service was reported stable by 11:24, with a total impact duration of approximately 1 hour 9 minutes.

The key learning is that high-traffic backend API endpoints require adequate caching and sufficient baseline infrastructure capacity to remain resilient under load spikes. Follow-up actions — including the caching layer implementation and CPU audit, will be completed and validated to prevent recurrence.

Visit Group |  Kungsgatan 34-36 |  411 19 Göteborg |  Sweden |  +46 (0)31 38 06 000
www.visitgroup.com