Context
A live commerce platform used during live shopping events for a large Brazilian fashion brand. Events usually peaked around 6,000 concurrent users; one passed 12,000. The backend ran on Google Cloud Run after a recent migration from AWS to GCP, and the platform had never been tested at that load.
What failed
- The database saturated first, mostly on CPU and memory, and the pressure cascaded through the rest of the infrastructure.
- Capacity could not grow fast enough. Thousands of people joined at once, and between cold starts and the autoscaling configuration, new instances arrived after the demand was already there.
- The migration was recent, but the problem only showed up at this level of traffic.
Immediate response
The event was live, so stability came before diagnosis. Resources were raised as an emergency measure to keep the platform up through the event. That bought time, not an explanation.
Reproducing the failure
- After the event, we reproduced the load under controlled conditions with Taurus.
- Developers simulated thousands of concurrent users, raising the load step by step.
- Metrics and logs at each step showed which limits were reached first.
Changes
With the infrastructure team, we adjusted infrastructure and Cloud Run autoscaling, retested, and repeated. Each round looked for a balance between enough capacity to absorb thousands of users arriving at once and the cost of keeping that capacity available.
My contribution
As a Senior Developer, I worked with the infrastructure team on:
- Running the Taurus load tests and reading the results against metrics and logs.
- Analyzing where the limits were and taking part in each round of autoscaling and capacity changes.
- Retesting after every change.
Result
After the adjustments, load tests sustained about 20,000 concurrent users without exceeding the infrastructure budget defined for that scenario. This is a load-test result, not a later production event.
Engineering takeaway
A spike is a different problem from growth. Autoscaling reacts to demand that already exists, so when thousands of people arrive in the same minute, cold starts decide whether capacity shows up in time. A load test that ramps up gently will pass on a system that fails at a live event; it has to reproduce simultaneous entry. From there, performance is a trade-off between capacity, how fast it reacts, and what it costs to keep ready.