At a glance
A missed environment variable crashed a live site mid-migration, while the client's campaign was already driving traffic to it.
- Redirected visitors to a maintenance page before touching anything else
- Found and fixed the missing environment variable directly, without rolling back
- Met with the client afterward to explain what happened
Service was restored in 35 minutes.
The situation
A client's Shopify site was being migrated to a custom Next.js build, moving the hosting from Vercel to AWS. The engineer running the migration got sick partway through, and a second developer took over mid-project. The client had already started a marketing campaign, and visitors were arriving at the domain while the migration was still in progress.
How it surfaced
I have an email alert on every deployment, sent to myself directly rather than routed through a dashboard someone has to remember to check. Two alerts reached me independently at two in the morning: ECS's own monitoring flagged the deployment crash, and UptimeRobot, already configured to watch the site, reported it down on its own. Neither needed a client report to become real. I was working the problem before the first visitor could have complained about it.
What I ruled out, and why
The developer who took over the migration had missed an environment variable in the production configuration, which took the site down on deploy. Some of the environment values had also come from the client with an error in them (an API key among them), so the check could not stop at the deploy configuration alone. I did not treat this as a case for rolling back, because by the time I was looking at it, the missing piece was already identified rather than still unknown.
The decision and what it cost
I chose to fix the missing variable directly rather than roll back to the last working deployment. A rollback would have restored the old Shopify-era state and cost the migration progress already made; a direct fix cost only the time to apply it, because I already knew what was missing rather than needing to find out. Before either, I put visitors on a maintenance page: a site down with no explanation is worse for a campaign-driven audience than a page that says work is in progress.
What I did
Visitors went to a maintenance page first, so a live campaign was not sending traffic into a broken site while I worked. I corrected the missing environment variable and the client-supplied value that had come in wrong, redeployed the container, and brought the site back. Service was restored in 35 minutes from the alert. I called the client afterward and walked through what had happened and why, rather than let a fixed outage go unexplained.
The outcome
Service was restored in 35 minutes. The mistake belonged to the developer who took over mid-migration, but I did not make the client conversation about assigning blame: a handoff mid-project is a risk I own, not only the person executing it.
What stayed changed
Two independent alerts reaching me directly, on every deployment, are what turned a two-in-the-morning failure into a 35-minute one instead of a multi-hour one discovered by a user complaint: neither system depended on the other to catch the same thing. The maintenance page is now a standing step before any fix during business-critical traffic, not a judgment call made under pressure each time. And a developer handoff mid-migration is now a checkpoint I run myself, not an assumption that the context transferred cleanly.
Related incidents
A Production Outage During a Traffic Spike: Moved to a Pre-Built AWS Backup in 30 Minutes
A Facebook campaign on a quiz and olympiad platform drove concurrent users to 20,000 against an estimated 8,000, and login and registration started timing out: the database sat in Tokyo while users and the server were in Bangladesh, and the cross-region latency broke down under the load. I moved the platform to an AWS ECS and EC2 backup already built for exactly this situation; the move itself took about 30 minutes, at three in the morning, slowed by an SSL verification issue.
- AWS Backup
- Cross-Region Latency
- 20K Concurrent Users
Payment Gateway Failure During Launch: Recovered by a System Built Before the Outage
SSLCommerz, the payment gateway behind an education platform, went down mid-launch and stayed down for about four hours, blocking checkout for roughly 300 students. A bKash fallback was live within about an hour, but what actually kept no one lost was a manual recovery path built into the admin panel before this outage ever happened: support used it to process every stuck payment by hand.
- SSLCommerz
- bKash Fallback
- Manual Recovery