Skip to content

Cloud Infrastructure

AWS ECS Deployment Failure Recovery: A Direct Fix in 35 Minutes, Not a Rollback

Migrating a client's Shopify site to a custom Next.js build on AWS, the engineer running the migration got sick mid-project and a second developer took over. A missed environment variable crashed the live site while the client's campaign was already driving traffic. An alert I have on every deployment woke me at two in the morning; I put visitors on a maintenance page first, then fixed the missing variable directly and redeployed. Service was back in 35 minutes.

CriticalPublished

At a glance

A missed environment variable crashed a live site mid-migration, while the client's campaign was already driving traffic to it.

  • Redirected visitors to a maintenance page before touching anything else
  • Found and fixed the missing environment variable directly, without rolling back
  • Met with the client afterward to explain what happened
Result

Service was restored in 35 minutes.

The situation

A client's Shopify site was being migrated to a custom Next.js build, moving the hosting from Vercel to AWS. The engineer running the migration got sick partway through, and a second developer took over mid-project. The client had already started a marketing campaign, and visitors were arriving at the domain while the migration was still in progress.

How it surfaced

I have an email alert on every deployment, sent to myself directly rather than routed through a dashboard someone has to remember to check. Two alerts reached me independently at two in the morning: ECS's own monitoring flagged the deployment crash, and UptimeRobot, already configured to watch the site, reported it down on its own. Neither needed a client report to become real. I was working the problem before the first visitor could have complained about it.

What I ruled out, and why

The developer who took over the migration had missed an environment variable in the production configuration, which took the site down on deploy. Some of the environment values had also come from the client with an error in them (an API key among them), so the check could not stop at the deploy configuration alone. I did not treat this as a case for rolling back, because by the time I was looking at it, the missing piece was already identified rather than still unknown.

The decision and what it cost

I chose to fix the missing variable directly rather than roll back to the last working deployment. A rollback would have restored the old Shopify-era state and cost the migration progress already made; a direct fix cost only the time to apply it, because I already knew what was missing rather than needing to find out. Before either, I put visitors on a maintenance page: a site down with no explanation is worse for a campaign-driven audience than a page that says work is in progress.

What I did

Visitors went to a maintenance page first, so a live campaign was not sending traffic into a broken site while I worked. I corrected the missing environment variable and the client-supplied value that had come in wrong, redeployed the container, and brought the site back. Service was restored in 35 minutes from the alert. I called the client afterward and walked through what had happened and why, rather than let a fixed outage go unexplained.

The outcome

Service was restored in 35 minutes. The mistake belonged to the developer who took over mid-migration, but I did not make the client conversation about assigning blame: a handoff mid-project is a risk I own, not only the person executing it.

What stayed changed

Two independent alerts reaching me directly, on every deployment, are what turned a two-in-the-morning failure into a 35-minute one instead of a multi-hour one discovered by a user complaint: neither system depended on the other to catch the same thing. The maintenance page is now a standing step before any fix during business-critical traffic, not a judgment call made under pressure each time. And a developer handoff mid-migration is now a checkpoint I run myself, not an assumption that the context transferred cleanly.

Related incidents