At a glance
Facebook-driven traffic on a quiz and olympiad platform reached 20,000 concurrent users against an estimated 8,000, and login and registration began timing out.
- Traced the timeouts to cross-region latency between the database and the application
- Moved the platform to an AWS ECS and EC2 backup built in advance for this situation
- Worked through an SSL verification issue during the move, starting around three in the morning
The move to the AWS backup took about 30 minutes. Roughly 2,000 registrations failed during the spike and were recovered manually afterward through the admin panel.
The situation
A quiz and olympiad platform's brand recognition on Facebook was already high, and a campaign there drove traffic the client had sized for 8,000 concurrent users. It reached 20,000. Login and registration started failing: specifically, the step after payment where a user's token comes back and they return to the platform began timing out.
How it surfaced
The failure sat in one place: the return trip after payment, not the payment itself. That is a latency signature, not a capacity one: a server that is out of compute fails broadly, a server that is waiting too long on one round trip fails at exactly the step that round trip belongs to.
What I ruled out, and why
The database was hosted in Tokyo, the default region on Supabase; Singapore was never separately considered or set up as an alternative before this happened. That was the actual gap: region was never a decision anyone made, it was just what defaulted in. The users and the application server were in Bangladesh. Under normal load the extra distance cost milliseconds nobody noticed. Under 20,000 concurrent users, every one of those round trips queued behind the last, and the token-return step (already the most latency-sensitive point in the flow) was where it broke first.
The decision and what it cost
I did not try to tune the existing setup under live load with users still arriving. An AWS ECS and EC2 backup had already been built for this kind of situation, before this traffic spike happened, and the decision was simply to use it rather than attempt a fix on infrastructure that was actively failing under a load it was never sized for.
What I did
Moving the platform to the AWS backup took about 30 minutes. An SSL verification issue slowed the cutover partway through, and the work started around three in the morning, while the campaign traffic was still live. Roughly 2,000 registrations failed to complete during the spike before the move finished.
The outcome
The move to the AWS backup took about 30 minutes from the point I started it. About 2,000 registrations failed during the spike itself; support recovered them afterward by hand through the admin panel, crediting the accounts affected. The cross-region latency problem itself was fully resolved by the move: it did not recur afterward.
What stayed changed
The backup that mattered was built before the traffic spike, not during it. An AWS ECS and EC2 environment existed and was ready to receive traffic because someone had already asked what happens if the primary setup cannot hold, before it could not. That is the same pattern as a payment gateway outage I have also documented: the response that looks fast during an incident is almost always preparation that happened earlier, not improvisation in the moment.
Related incidents
AWS ECS Deployment Failure Recovery: A Direct Fix in 35 Minutes, Not a Rollback
Migrating a client's Shopify site to a custom Next.js build on AWS, the engineer running the migration got sick mid-project and a second developer took over. A missed environment variable crashed the live site while the client's campaign was already driving traffic. An alert I have on every deployment woke me at two in the morning; I put visitors on a maintenance page first, then fixed the missing variable directly and redeployed. Service was back in 35 minutes.
- Vercel to AWS
- Env Variable Bug
- Layered Monitoring
A Launch Delayed an Hour by an Undisclosed Implementation Choice
Ahead of a SaaS launch for an international client, I asked a developer to build one API route pulling the admin dashboard's data from multiple tables at once. He built several separate calls on the same page instead, and didn't tell me. I fixed the query layer myself with indexing and a rewritten query, and told the client the launch would be about an hour late, rather than let a slow admin page surface on its own.
- SQL Indexing
- Query Rewrite
- Undisclosed Change