Skip to content

Revenue Risk

Payment Gateway Failure During Launch: Recovered by a System Built Before the Outage

SSLCommerz, the payment gateway behind an education platform, went down mid-launch and stayed down for about four hours, blocking checkout for roughly 300 students. A bKash fallback was live within about an hour, but what actually kept no one lost was a manual recovery path built into the admin panel before this outage ever happened: support used it to process every stuck payment by hand.

HighPublished

At a glance

SSLCommerz, the platform's payment gateway, went down mid-launch and stayed down for about four hours.

  • Switched checkout to a bKash fallback within about an hour
  • Used a manual recovery path already built into the admin panel to process each stuck payment by hand
  • Told the client from the first error alert, before any user had to report it
Result

About 300 transactions were stuck while SSLCommerz was down. None were lost: every one was recovered by hand through a system built before this outage happened.

The situation

SSLCommerz, the payment gateway behind an education platform's checkout, stopped processing payments mid-launch and stayed down for about four hours. I found out from an email error alert rather than from a user, which meant the client heard about the outage from me before they heard it from anyone else.

How it surfaced

The alert was specific enough to act on immediately: payment calls to SSLCommerz were failing consistently, not intermittently. A gateway that fails every call rather than some of them points at the gateway itself rather than at our own checkout code, so that is not where I spent the first hour.

What I ruled out, and why

SSLCommerz's own side was not mine to investigate during the outage, and it was still down when I needed an answer for the client, so I did not spend the incident chasing why it had failed. What I could control was whether payment kept working while it stayed down, and that is where the time went instead. The cause, confirmed afterward: the client's API secret key had changed during a developer's copy-paste while working on the integration, so calls were being rejected rather than failing on SSLCommerz's own infrastructure.

The decision and what it cost

The decision that mattered here was not made during the outage. Months earlier, the team had built a bKash fallback path and a manual recovery tool into the admin panel, specifically for a payment gateway going down mid-launch. During the outage itself, the only real decision was to use both: switch checkout to the bKash fallback, and route every payment SSLCommerz had swallowed through the manual admin path instead of waiting for SSLCommerz to come back and hoping nothing had been lost in the meantime.

What I did

The bKash fallback was live within about an hour, which kept new checkouts moving while SSLCommerz stayed down. SSLCommerz itself was down for about four hours in total, and roughly 300 transactions were stuck in that window: payments that had been attempted but never confirmed. Support worked through those by hand, one at a time, through the admin panel's manual entry path, rather than leaving them for users to notice and report.

The outcome

About 300 transactions were stuck while SSLCommerz was down for roughly four hours. None were lost. That is not a claim about SSLCommerz failing gracefully. It did not. It is a claim about what happens after a gateway fails when the recovery path already exists: support processed every stuck payment by hand, and no student who paid walked away without the course they paid for.

What stayed changed

The system that mattered was not built during the incident, it was built before it. A fallback payment method and a manual recovery path into the admin panel both existed because someone had already asked what happens if the gateway goes down mid-launch, before it did. I now ask that question about every dependency a launch depends on, not after the first one breaks.

Related incidents