Delivery incidents resolved
8 documentedPayment Gateway Failure During Launch: Recovered by a System Built Before the Outage
SSLCommerz, the payment gateway behind an education platform, went down mid-launch and stayed down for about four hours, blocking checkout for roughly 300 students. A bKash fallback was live within about an hour, but what actually kept no one lost was a manual recovery path built into the admin panel before this outage ever happened: support used it to process every stuck payment by hand.
- SSLCommerz
- bKash Fallback
- Manual Recovery
Sprint Velocity Dropped 40% Mid-Sprint: What To Do About It
A 4-person Exprovia team's output dropped 40% mid-sprint, with 15 days left against roughly 25 days of remaining work. The cause was not effort: junior developers were taking complicated paths through problems on the codebase's critical parts, where simpler ones existed. I sat with each of them individually, worked through simpler approaches together, and spent the intervening weekend getting ahead of it. Velocity recovered in about four days.
- 1:1 Coaching
- Rate-Limit Bug
- Sprint Recovery
Managing Scope Creep with Clients: A Single-Vendor Platform Becomes Multi-Vendor Mid-Build
A client changed an education platform from single-vendor to multi-vendor mid-build: a structural change, not a feature request. As Associate Project Manager, I asked for more time; the client declined but increased the budget instead. I split the added work so the developers already on the project handled what needed context, and new developers handled what did not. The original delivery held, and Phase 2 shipped after it.
- Multi-Vendor Pivot
- Budget Negotiation
- Work Splitting
A Production Outage During a Traffic Spike: Moved to a Pre-Built AWS Backup in 30 Minutes
A Facebook campaign on a quiz and olympiad platform drove concurrent users to 20,000 against an estimated 8,000, and login and registration started timing out: the database sat in Tokyo while users and the server were in Bangladesh, and the cross-region latency broke down under the load. I moved the platform to an AWS ECS and EC2 backup already built for exactly this situation; the move itself took about 30 minutes, at three in the morning, slowed by an SSL verification issue.
- AWS Backup
- Cross-Region Latency
- 20K Concurrent Users
A Launch Delayed an Hour by an Undisclosed Implementation Choice
Ahead of a SaaS launch for an international client, I asked a developer to build one API route pulling the admin dashboard's data from multiple tables at once. He built several separate calls on the same page instead, and didn't tell me. I fixed the query layer myself with indexing and a rewritten query, and told the client the launch would be about an hour late, rather than let a slow admin page surface on its own.
- SQL Indexing
- Query Rewrite
- Undisclosed Change
When the Platform Could Not Carry the Product: Moving to Headless WordPress with Next.js
Close to launch, a large-scale US civic information platform's client brought behaviors plain WordPress could not run: an address-based sample ballot live from the Google Civic API, an interactive polling-place map, and a 700+ candidate directory with real-time filters. I told the client directly the platform could not carry it, negotiated $1,000 and one more month, and split the architecture: WordPress for content and commerce, Next.js for everything interactive. The platform launched, and the client is now building mobile and iOS apps on it.
- Google Civic API
- JetEngine + WooCommerce
- API-First Backend
AWS ECS Deployment Failure Recovery: A Direct Fix in 35 Minutes, Not a Rollback
Migrating a client's Shopify site to a custom Next.js build on AWS, the engineer running the migration got sick mid-project and a second developer took over. A missed environment variable crashed the live site while the client's campaign was already driving traffic. An alert I have on every deployment woke me at two in the morning; I put visitors on a maintenance page first, then fixed the missing variable directly and redeployed. Service was back in 35 minutes.
- Vercel to AWS
- Env Variable Bug
- Layered Monitoring
Automating Project Intake: From a Marketing Form to an Assigned Developer
Marketing's project-intake process ran through a form, a hand-typed ClickUp task, a hand-written project document, and a manual developer handoff. I automated the mechanical transcription and put an LLM on the one step that actually needed interpretation: reading the form input and drafting the project documentation, including the developer-facing instructions. That draft comes to me first, not a developer: I review it, record a video brief, and only then assign the work.
- ClickUp
- Telegram
- LLM-Drafted Docs
Not an incident
This one is here because it is the evidence behind the 7 to 20 figure on the homepage. It is a scaling story, not a recovery, so it is not counted among the incidents above.
Prevention, not recovery
This one is here because it heads off an incident rather than resolves one: reading a client's own technical proposal closely enough to catch expensive assumptions before a sprint started. It is not counted among the incidents above, and, unlike the rest, the work it describes is still in progress.