High Availability for a Large Online Travel Platform

Eight platform-wide sales events in over two years with zero major incidents, load tested at up to 8 times normal traffic.

Case study image
Case study image
Case study image

Challenge

A major Japanese online travel platform runs several platform-wide sales events every year. Traffic rises to 4 to 5 times the normal level during these events, and core business systems must stay stable under high concurrency. The traditional approach relied on people watching screens and investigating after the fact: operations were costly, problems were slow to pinpoint, and it was hard to find system bottlenecks before an event began.

Solution

Scope of work: Reliability engineering for core business systems, including unified monitoring, large-scale load testing, CI/CD and infrastructure automation, and automated metrics reporting, over more than two years.

  • Large-scale load testing: 3 rounds of load testing before every sales event, reaching up to 8 times normal traffic. That is well above the 4 to 5 times seen in an actual event, leaving headroom for sudden spikes.
  • Finding and fixing bottlenecks: Load testing exposed a bottleneck at the database layer. After optimization, response time for the affected endpoints fell from 1 second to under 100 milliseconds, and the problem was solved before the event began.
  • Unified monitoring platform: Monitoring data from Grafana, Prometheus, and other sources brought together to visualize system health, so key metrics can be watched in real time during events.
  • CI/CD and infrastructure automation: Built Jenkins continuous integration and automated deployment pipelines, and maintained Kubernetes clusters and infrastructure automation, improving deployment efficiency and consistency.
  • Automated metrics reporting: Developed a system that collects key performance metrics automatically and generates PDF reports, so operations data no longer has to be compiled by hand.

Results

  • 8 sales events, zero major incidents: More than 2 years of continuous reliability work kept the system stable through 8 platform-wide sales events.
  • Response time cut to under one tenth: After the database bottleneck was fixed, endpoint response time fell from 1 second to under 100 milliseconds.
  • Standardized processes: A standardized performance analysis process and an automated key-metrics reporting system are in place, and the engineering team pinpoints problems faster.

Experience building high-concurrency, high-availability systems like this is what backs our commitment to clients that systems will run reliably for the long term after launch.

More case studies

See other case studies

Video data and transaction data connected across about 500 stores, cutting the investigation of a suspicious transaction from a week to a day.

A large online produce trading platform, built from scratch and operated, ending in a sale of the business.

On-site deployment for a US robotics AI company, keeping robots running reliably on real construction sites.

Decorative globe
Contact us

The next project to get done should be yours

Write to us about the problem you want to solve. Whether it is worth doing, how to do it, and how soon it can show results: we will give you a professional and direct assessment.