How Idempotent Retries Amplified GitHub's August Outage to 7 Hours and 47 Minutes

Idempotent Retries
Published on
Written byMayank Singh
How Idempotent Retries Amplified GitHub's August Outage to 7 Hours and 47 Minutes

TL;DR: GitHub's August outage started with a brief infrastructure failure. It ended 7 hours and 47 minutes later. The original component wasn't the problem. The retry layer built to absorb failures made them worse. The right lesson isn't smarter retries. It's deciding whether a request should be retried at all. That decision must come from idempotency classification, not HTTP status code.

Key Takeaways: - Retry amplification turns a transient failure into a cascading outage when clients send duplicate load to a recovering server. - Status codes alone can't tell you whether a retry is safe. Only an idempotency key can. - The AWS Builders Library's "idempotency first, status code second" rule is the architectural decision that prevents retry storms.

The Brief Bug That Lasted 7 Hours and 47 Minutes

Illustration for The Brief Bug That Lasted 7 Hours and 47 Minutes

GitHub's Central US data center experienced a brief infrastructure failure in August. The triggering event was familiar. Traffic hit a new peak and a critical component failed to scale with it.

That part of the story is ordinary. What happened next is what made the incident a case study.

Errors from the failing service triggered a client-side retry loop across dependent services and SDKs. Each retry added load to a system that was already struggling. GitHub's own postmortem states the company had to mitigate the retry behavior before it could safely restore traffic.

The safety net had become the source of the damage. The brief failure became a 7-hour-47-minute incident.

That ratio is the signal. When an original failure time is dwarfed by total outage time, the root cause is rarely the component that failed first. It's the system designed to respond to that failure.

Most teams read the GitHub postmortem and conclude they need smarter retries, faster backoff, or a better client library. That's the wrong lesson.

Why 'Add More Retries' Is the Exact Wrong Lesson

Here's the mechanism nobody talks about. A retry multiplier on a system already at peak capacity means the failing service receives several times its normal load. It gets this load at the worst possible moment.

The server isn't just trying to recover. It's trying to recover while being hammered by clients who think they're being helpful.

The hidden assumption baked into every retry library is that the failing service will recover quickly. When it does, the retry storm subsides and everything looks fine. When it doesn't, retries pile up faster than the service can drain its queue.

The system is drowning in good intentions.

Then there's the status code problem. A client that times out has no way to distinguish between two very different outcomes.

Case one: the request never reached the server, so retrying is safe. Case two: the server processed the operation, but the response was lost on the way back. Retrying here means executing the operation twice. Same status code. Wildly different consequences.

This is the same class of distributed failure that turns a single outage into a system-wide cascade. The client can't know, so the system amplifies its own uncertainty.

The AWS Builders Library and Stripe's engineering blog have warned about this for years. Most SDKs still default to "retry on 5xx." That's not resilience. That's a multiplier on your next incident.

The fix isn't smarter retries. It's deciding whether you should retry at all. That decision has to be made before the request leaves the client, not after the timeout fires.

The Decision Rule: Idempotency First, Status Code Second

Idempotency-first means: before checking the HTTP status code, the client must know whether the operation is safe to repeat. A timed-out POST to a payments API is not safe to retry without an idempotency key, even if the server returned a 504.

The status code tells you the server is in trouble. The idempotency classification tells you whether retrying will make the client's situation better or worse.

The AWS Builders Library on idempotent APIs frames it cleanly: "An idempotent operation is one where a request can be retransmitted or retried with no additional side effects." That property is what makes retries safe. Without it, every retry is a coin flip between "fixed the problem" and "created a duplicate."

The idempotency key is a client-generated UUID attached to the logical operation, not the network request. The same key with the same payload returns the same result. The same key with a different payload is a conflict. The key survives across retries, processes, and regions.

This is why Stripe, AWS, and other infrastructure providers make idempotency keys mandatory for write operations, not optional best practice. They have to. Their systems cannot safely accept a retry without it.

The pattern shows up everywhere a safety mechanism is applied without understanding the operation. Stale Kubernetes upgrade paths create the same amplifier effect. There, the thing meant to protect becomes the thing that fails.

Knowing the rule is one thing. Building it correctly under production load is another. The same client retries across multiple instances. Keys collide. Storage costs compound.

Building Retries That Don't Make the Next Outage Worse

Illustration for Building Retries That Don't Make the Next Outage Worse

Five patterns separate a retry layer that absorbs faults from one that amplifies them.

Generate the idempotency key at the logical operation, not the HTTP request. The key must be created the moment the user action begins, not the moment the SDK builds the call. If the key is tied to the request, you lose it on retry. The key must survive across retries, processes, and active-active failover scenarios.

Store the key and its response server-side for a bounded window. 24 hours is the standard. It's long enough to cover any realistic client retry window, short enough to avoid unbounded storage growth. A matching key returns the cached response, not a re-execution. After the window expires, the key is fresh.

Use a strict status code taxonomy: - Retry candidates: 408, 429, 500, 502, 503, 504. - Never retry: 400, 401, 403, 404, 409, 422. - Terminal: 200, 201. No retry, even on transport failure.

Backoff with jitter, always. Exponential backoff prevents thundering herds. Jitter (full or equal) prevents synchronized retry waves. Fixed-interval retries in production are a guarantee of cascading failure.

Client-side circuit breakers and retry budgets. When failure rate crosses a threshold, stop retrying entirely. A retry budget (a hard cap of 10 retries per minute per client) prevents any single client from becoming the load source during a server-side incident.

These patterns work because they account for the outage scenario, not just the happy path. The failure mode most common in production is exactly this: a component meant to provide resilience becomes the new source of load during recovery.

But the harder problem isn't technical. It's organizational. The biggest resilience failures come from teams that build retry logic without ever testing what happens when the thing being retried is the problem.

What Most CTOs Miss About Resilience Patterns

Most resilience failures look like over-engineering. A team adds retries, fallbacks, and circuit breakers to every endpoint without classifying which operations are actually safe to repeat.

This is the architectural equivalent of a seatbelt that strangulates you in a crash. The mechanism was supposed to protect you. It was applied to operations it doesn't understand.

In regulated environments, the blast radius is worse. In healthcare systems handling patient records, a duplicate write isn't just a bug. It's a patient-safety event. The same idempotency rules apply, but the consequence of getting them wrong is clinical, not financial.

The "we have retries" trap is the most common version. Leaders assume the system is resilient because retries exist. They rarely test what happens when retries themselves are the load source. The GitHub outage is the canonical example. So is the failure pattern in high-availability setups where a single mechanism becomes the next cascade.

Resilience patterns are not a checklist. They're a decision tree that begins with one question: "If this request is processed twice, what breaks?"

Getting this right changes your incident trajectory from hours to minutes. It changes the conversation with your board from "what went wrong" to "what held."

What Changes When You Get This Right

Outage duration compresses. A brief infrastructure blip stays a brief blip because the retry layer cannot amplify the failure. Your incident postmortems change character. Instead of "retries made it worse," they read "retries absorbed the transient fault as designed."

Customer trust compounds. Clients who never see a 7-hour outage because your system handled the retry storm correctly are clients who renew. The resilience layer stops being a source of cascading failures and starts being invisible.

Engineering focus shifts. Teams stop firefighting retry storms and start improving the actual product. The resilience layer, when built on the idempotency-first decision rule, becomes a foundation rather than a liability.

The pattern is the same every time: classify operations by idempotency, not by status code. The rest of the architecture falls into place.

Frequently Asked Questions

How do I know if an API endpoint is safe to retry?

Check whether the operation is idempotent. Repeated execution must produce the same result as a single execution. GET, PUT, and DELETE are idempotent by HTTP convention. POST is not, unless the server supports an idempotency key header such as Stripe's `Idempotency-Key` or AWS's `Idempotency-Token`. If you're calling a POST endpoint without idempotency support, you cannot safely retry it on timeout.

What is retry amplification and how do I prevent it?

Retry amplification occurs when client-side retries, triggered by a server-side failure, increase load on the failing server. This makes recovery slower or impossible. The GitHub August outage lasted 7 hours and 47 minutes partly because of this exact pattern. Prevention requires three things: idempotency keys (so retries don't create duplicate work), backoff with jitter (so retries don't synchronize), and a client-side retry budget (so retries have a hard cap).

How long should I store idempotency keys on the server?

Most implementations store keys for 24 hours. That's long enough to cover any realistic client retry window, short enough to avoid unbounded storage growth. Stripe documents this window explicitly. The key and the response must be stored together. A key without its cached response is useless for deduplication.

Should I retry on 5xx errors?

Only if the operation is idempotent. Retrying a non-idempotent POST on a 500 error risks executing the operation twice. Retrying an idempotent GET or a POST with an idempotency key is safe. The status code tells you the server is in trouble. The idempotency classification tells you whether retrying will make the client's situation better or worse.

What's the difference between an idempotency key and a request ID?

A request ID is for observability. It correlates logs and traces across a single network call. An idempotency key is for correctness. It tells the server that two HTTP requests represent the same logical operation and should produce the same result. Many systems use one UUID for both, but they solve different problems and should be designed independently.

The patterns above work as a starting audit for any team that ships write operations at scale.

About the author

MS
Mayank Singh
Software Developer, Levitation Infotech

Mayank Singh is a software developer at Levitation Infotech, where he builds web and AI-powered applications across the company’s fintech, healthcare, and enterprise projects.

Supercharge Your Success with Our Expertise

Amplify Your Business with Our Expertise. Explore Services Tailored for Your Success.

Get In Touch