Skip to content

Incident response · Delivery under pressure

Software Delivery Case Studies

Real leadership is easiest to judge from what happened when something broke. Each of these carries the same structure: the situation, what I ruled out and why, the call I made and what it cost, and what stayed changed afterwards.

Delivery incidents resolved

8 documented
High
Revenue Risk

Payment Gateway Failure During Launch: Recovered by a System Built Before the Outage

SSLCommerz, the payment gateway behind an education platform, went down mid-launch and stayed down for about four hours, blocking checkout for roughly 300 students. A bKash fallback was live within about an hour, but what actually kept no one lost was a manual recovery path built into the admin panel before this outage ever happened: support used it to process every stuck payment by hand.

  • SSLCommerz
  • bKash Fallback
  • Manual Recovery
View Details
High
Delivery Risk

Sprint Velocity Dropped 40% Mid-Sprint: What To Do About It

A 4-person Exprovia team's output dropped 40% mid-sprint, with 15 days left against roughly 25 days of remaining work. The cause was not effort: junior developers were taking complicated paths through problems on the codebase's critical parts, where simpler ones existed. I sat with each of them individually, worked through simpler approaches together, and spent the intervening weekend getting ahead of it. Velocity recovered in about four days.

  • 1:1 Coaching
  • Rate-Limit Bug
  • Sprint Recovery
View Details
High
Stakeholder

Managing Scope Creep with Clients: A Single-Vendor Platform Becomes Multi-Vendor Mid-Build

A client changed an education platform from single-vendor to multi-vendor mid-build: a structural change, not a feature request. As Associate Project Manager, I asked for more time; the client declined but increased the budget instead. I split the added work so the developers already on the project handled what needed context, and new developers handled what did not. The original delivery held, and Phase 2 shipped after it.

  • Multi-Vendor Pivot
  • Budget Negotiation
  • Work Splitting
View Details
Critical
Live Incident

A Production Outage During a Traffic Spike: Moved to a Pre-Built AWS Backup in 30 Minutes

A Facebook campaign on a quiz and olympiad platform drove concurrent users to 20,000 against an estimated 8,000, and login and registration started timing out: the database sat in Tokyo while users and the server were in Bangladesh, and the cross-region latency broke down under the load. I moved the platform to an AWS ECS and EC2 backup already built for exactly this situation; the move itself took about 30 minutes, at three in the morning, slowed by an SSL verification issue.

  • AWS Backup
  • Cross-Region Latency
  • 20K Concurrent Users
View Details
Critical
Implementation Gap

A Launch Delayed an Hour by an Undisclosed Implementation Choice

Ahead of a SaaS launch for an international client, I asked a developer to build one API route pulling the admin dashboard's data from multiple tables at once. He built several separate calls on the same page instead, and didn't tell me. I fixed the query layer myself with indexing and a rewritten query, and told the client the launch would be about an hour late, rather than let a slow admin page surface on its own.

  • SQL Indexing
  • Query Rewrite
  • Undisclosed Change
View Details
High
Architecture Decision

When the Platform Could Not Carry the Product: Moving to Headless WordPress with Next.js

Close to launch, a large-scale US civic information platform's client brought behaviors plain WordPress could not run: an address-based sample ballot live from the Google Civic API, an interactive polling-place map, and a 700+ candidate directory with real-time filters. I told the client directly the platform could not carry it, negotiated $1,000 and one more month, and split the architecture: WordPress for content and commerce, Next.js for everything interactive. The platform launched, and the client is now building mobile and iOS apps on it.

  • Google Civic API
  • JetEngine + WooCommerce
  • API-First Backend
View Details
Critical
Cloud Infrastructure

AWS ECS Deployment Failure Recovery: A Direct Fix in 35 Minutes, Not a Rollback

Migrating a client's Shopify site to a custom Next.js build on AWS, the engineer running the migration got sick mid-project and a second developer took over. A missed environment variable crashed the live site while the client's campaign was already driving traffic. An alert I have on every deployment woke me at two in the morning; I put visitors on a maintenance page first, then fixed the missing variable directly and redeployed. Service was back in 35 minutes.

  • Vercel to AWS
  • Env Variable Bug
  • Layered Monitoring
View Details
Medium
Workflow Automation

Automating Project Intake: From a Marketing Form to an Assigned Developer

Marketing's project-intake process ran through a form, a hand-typed ClickUp task, a hand-written project document, and a manual developer handoff. I automated the mechanical transcription and put an LLM on the one step that actually needed interpretation: reading the form input and drafting the project documentation, including the developer-facing instructions. That draft comes to me first, not a developer: I review it, record a video brief, and only then assign the work.

  • ClickUp
  • Telegram
  • LLM-Drafted Docs
View Details

Not an incident

This one is here because it is the evidence behind the 7 to 20 figure on the homepage. It is a scaling story, not a recovery, so it is not counted among the incidents above.

Prevention, not recovery

This one is here because it heads off an incident rather than resolves one: reading a client's own technical proposal closely enough to catch expensive assumptions before a sprint started. It is not counted among the incidents above, and, unlike the rest, the work it describes is still in progress.