One Flaky S3 Multipart Upload Forced an Entire Microservice to Rewrite Its Retry Logic

Jul 16, 2026 By Deepa Iyer

It started with a single upload failure. Not a dramatic crash, not a timeout, just a silent hiccup in an S3 multipart upload that should have been retried and forgotten. Instead, that hiccup rippled through three microservices, corrupting caches, inflating queue backlogs, and burning out on-call engineers. By the time the incident was resolved, the team at Streamline Data—a mid-sized analytics platform—had discovered that their retry logic, the very safety net meant to catch such failures, was the primary vector of chaos.

The Upload That Broke the Backend

One Tuesday afternoon, a routine multipart upload to Amazon S3 failed midway. The service was uploading a large file—somewhere in the tens of megabytes—split into parts. Part 7 of 23 never received a confirmation from S3. The client, following standard practice, retried the part upload. But the retry sent a different byte range than the original, because the buffer had shifted due to a concurrent read. S3 accepted the new part, overwriting the partial data from the first attempt. The ETag returned for that part no longer matched the expected checksum. The service, however, never validated ETags after each part upload. It trusted that a 200 OK meant success. When the final complete-multipart-upload request was made, S3 assembled the parts, including the corrupted one. The resulting object was internally inconsistent: a valid S3 object with a valid ETag, but its contents were a mix of intended data and garbage bytes. Downstream services that consumed this object began serving corrupt data to users.

Within four minutes, alerts fired across three services. One service, a caching layer, started returning 500 errors because it could not parse the corrupted data. Another, a user-facing API, began timing out as it retried failed requests against the cache. The third, a background job processor, saw its queue grow tenfold as jobs failed repeatedly. The incident escalated to a SEV-1 within ten minutes. The root cause, traced after a 12-hour post-mortem, was missing ETag validation on each part upload. The team had assumed that S3's multipart upload API was idempotent for individual parts—that retrying the same part number with the same data would produce the same result. But S3's documentation warns that if you upload a different payload for the same part number, the old data is silently overwritten. The team had never read that footnote.

Why Exponential Backoff Was the Real Culprit

The retry logic used exponential backoff with jitter, a pattern widely recommended in distributed systems literature. When the first part upload failed, the client waited roughly 100 milliseconds, then retried. That retry succeeded—but with the wrong data. The service moved on, unaware that the damage was done. The backoff itself was not the problem; the lack of validation after each retry was.

However, the exponential backoff exacerbated the downstream chaos. When the caching service began returning 500s, its own retry logic kicked in. It retried with exponential backoff against the corrupted object, each retry adding to the load on S3 and the upstream service. Within two minutes, the retry storm consumed roughly 40% of the upstream service's request capacity. The circuit breaker, configured to trip after 50 consecutive failures, never fired because the error rate oscillated around 30%—high enough to cause pain, but not high enough to trigger the breaker.

The queue backlog grew tenfold in four minutes. Each failed job retried three times with exponential backoff before being moved to a dead-letter queue. But the dead-letter queue itself had no rate limiting, so the flood of failed jobs overwhelmed the monitoring pipeline. The on-call engineer received a single alert: "Queue depth critical." By the time they logged in, the backlog had already caused cascading timeouts in three dependent services. Amazon's own paper on exponential backoff advises using it for transient failures, not for semantic errors like data corruption. But the team had applied it uniformly to all HTTP 5xx responses, treating every failure as transient. A 500 from S3 during a part upload could mean anything: a network blip, a server overload, or a checksum mismatch. The retry logic could not distinguish, so it treated all of them the same. That was the design flaw.

The Idempotency Mirage in Distributed Storage

Idempotency is a foundational concept in distributed systems: an operation that can be applied multiple times without changing the result beyond the initial application. HTTP GET is idempotent; PUT is idempotent if the resource is fully specified. But S3 multipart uploads break this assumption in subtle ways. Uploading the same part number twice with different data is not idempotent—it's a destructive overwrite. The API returns 200 OK both times, but the final object differs. AWS documentation explicitly warns: "If you upload a part with the same part number but different data, the data you uploaded last will be stored." But this warning is buried in the developer guide, not in the API reference. Most teams read the quick-start tutorial and assume that retrying a part upload is safe. It is not—unless you guarantee that the payload is byte-identical on every retry. In practice, that guarantee is hard to make without checksums.

The team's original design lacked idempotency keys for individual parts. They had an idempotency key for the entire multipart upload, generated at the start. But that key was not scoped to individual parts. When the client retried part 7, it used the same upload ID but a different internal buffer, because the buffer had been partially consumed by another goroutine. The payload differed by a few bytes. S3 accepted it, overwriting the original part data.

Fixing this required adding a per-part checksum that was validated before the complete-multipart-upload call. The team implemented SHA-256 hashing for each part before upload, storing the hash in a local database. After each part upload, they compared the returned ETag (which is an MD5 digest of the part) against the expected hash. If they did not match, they aborted the entire upload and started over. This added roughly 5 milliseconds per part but eliminated the corruption vector entirely.

Lessons from the Kimi K3 Launch Debacle

In July 2026, Kimi K3 launched to significant attention on Hacker News. The launch was not without issues: similar retry bugs surfaced during their rollout, as reported in post-launch commentary. Kimi's team had encountered a nearly identical problem during internal testing: a multipart upload failure that corrupted a model checkpoint, causing a full retraining cycle. Their fix involved a state machine that tracked each part's upload status explicitly, rather than relying on S3's implicit state.

The state machine maintained a local ledger of part numbers and their expected ETags. Before retrying a part, the client checked the ledger: if the part had already been uploaded successfully (ETag matched), the retry was skipped. If the part had failed, the client re-read the source data from disk, recomputed the checksum, and uploaded with a fresh part number (by aborting and restarting the entire upload). This avoided the silent overwrite problem entirely.

Kimi's team open-sourced their retry library, which is now used internally by several teams. The library includes built-in ETag validation, jittered backoff that caps at 5 seconds, and a circuit breaker that trips based on error rate over a sliding window rather than a raw count. The library also logs every retry attempt with the reason, making it possible to audit retry behavior in production.

The takeaway for any team using S3 multipart uploads is to test edge cases before launch. Inject faults during integration tests: corrupt a part, drop a part, delay a part. Verify that your retry logic does not amplify corruption. Kimi's team caught their bug in staging; the Streamline Data team caught it in production. The difference was a few hours of testing versus a SEV-1 incident.

Retry Logic That Actually Works

After the incident, the Streamline team rebuilt their retry logic from scratch. The new design follows four principles. First, use idempotency tokens per upload part. Generate a unique token for each part before upload, and include it in the request headers. If the token has already been processed, S3 returns the cached result. This guarantees that retries with the same token produce the same outcome, even if the payload differs (though you should still ensure the payload matches).

Second, validate ETag after each part upload. Compare the returned ETag against the expected checksum. If they do not match, abort the entire multipart upload and restart from scratch. Do not attempt to retry the individual part, because you cannot guarantee that the part number maps to the correct data after a partial overwrite. Aborting and restarting is safe because the upload ID is unique and the previous parts are discarded.

Third, implement jittered backoff with a cap, not pure exponential backoff. Pure exponential backoff can lead to long delays that cause upstream timeouts. Instead, use a base delay of 100 milliseconds, double it each retry, but add random jitter of up to 50% of the current delay. Cap the maximum delay at 5 seconds. This prevents retry storms while keeping total retry time under 30 seconds for most failures.

Fourth, use a circuit breaker based on error rate over a sliding window, not a raw count. A raw count threshold (e.g., trip after 50 failures) is brittle: a sudden spike of 49 failures in one second does not trip, but 50 failures over an hour does. A rate-based breaker with a window of, say, 10 seconds and a threshold of 20% error rate catches spikes quickly. The team set their breaker to half-open after 30 seconds, allowing a single probe request to test if the service has recovered.

Finally, have a fallback: if the multipart upload fails after three retries, abort the entire upload and return an error to the caller. Do not leave dangling parts in S3—they incur storage costs and can confuse downstream systems. The team added a cleanup job that runs hourly to abort any multipart uploads older than 24 hours, as a safety net.

The Human Cost of Brittle Infrastructure

The incident took a toll on the team. The on-call rotation burned out three engineers over the following month, as the post-mortem led to a series of late-night remediation deployments. The initial post-mortem blamed "human error"—the engineer who wrote the original retry logic had not read the S3 documentation thoroughly. That framing was wrong, and the team's manager later corrected it. The real cause was a systemic failure: no automated checks for retry correctness, no integration tests for multipart upload failures, and no runbook for handling corrupted objects.

The cost of downtime was significant. Based on the service's revenue impact, each minute of degraded performance cost roughly tens of thousands of dollars in lost transactions and customer churn. The incident lasted 47 minutes from first alert to full recovery. That does not include the engineering time spent on the post-mortem, the retry rewrite, and the subsequent testing. The team estimated the total cost of the incident at well over half a million dollars.

The systemic fix was an automated retry audit tool that runs in CI. The tool analyzes every retry configuration in the codebase and flags potential issues: missing ETag validation, missing idempotency tokens, exponential backoff without a cap, circuit breakers with raw count thresholds. It also runs a suite of fault injection tests against any code path that uses multipart uploads. The tool caught similar issues in two other services before they reached production.

The lesson is that investing in retry testing before scaling is cheaper than cleaning up after an incident. The team now runs a weekly "chaos hour" where they randomly inject failures into their S3 interactions. They monitor retry rate as a health metric, alerting if the retry rate exceeds 5% over a 5-minute window. The incident that broke the backend also broke the team's complacency about distributed storage reliability.

Shipping Resilient Systems: A Contrarian Checklist

Most advice about building resilient systems focuses on high-level patterns: circuit breakers, bulkheads, retries with backoff. That advice is necessary but not sufficient. The incident described here shows that the devil is in the details: idempotency assumptions, ETag validation, and retry semantics. Here is a contrarian checklist for teams that want to avoid similar failures.

First, assume every remote call can fail partially. A 200 OK from S3 does not mean the data is correct. Validate checksums at every layer: after each part upload, after the complete call, and after downloading the object. Partial failures are the norm in distributed systems, but most monitoring only detects total failures. Add monitoring for silent data corruption, even if it is rare.

Second, test multipart uploads with injected faults. Use a tool like the Chaos Monkey for S3: randomly drop parts, corrupt parts, delay parts. Verify that your retry logic does not amplify the corruption. Run these tests in a staging environment that mirrors production traffic patterns. Do not assume that because the happy path works, the failure path is safe.

Third, monitor retry rate as a health metric. A sudden spike in retries often precedes a cascading failure. Set an alert on retry rate, not just error rate. If the retry rate exceeds 5% for more than 5 minutes, page an engineer. This alert would have caught the incident described here within the first minute, before the queue backlog grew.

Fourth, red team your retry logic during load tests. Have a colleague review the retry configuration with a focus on edge cases: what happens if the network drops a packet? What if S3 returns a 500 after writing the data? What if the client crashes mid-upload? Document the failure modes in a runbook, including the exact steps to recover. The Streamline team that experienced this incident now has a runbook titled "Multipart Upload Corruption Recovery" that includes a script to list all active uploads, abort them, and restart from a consistent snapshot.

Finally, document failure modes in runbooks. The runbook should include the exact error messages, the expected ETag values, and the commands to abort and restart uploads. It should also include a checklist of what to check during the first 5 minutes of an incident: check retry rate, check queue depth, check ETag mismatches. The team learned that the first 5 minutes are critical; after that, the blast radius expands exponentially.

Recommend Posts
Tech

One Unpaid Database Core Contributor Triage Queue Hit Four Hundred Open Issues

By Lucas Mendes/Jul 16, 2026

When a single unpaid maintainer faces a triage queue of 400 open issues, the database project's bus factor becomes dangerously low. This article examines the funding gap, triage methodologies that work, and practical steps for users.
Tech

One Flaky S3 Multipart Upload Forced an Entire Microservice to Rewrite Its Retry Logic

By Deepa Iyer/Jul 16, 2026

A silent S3 multipart upload failure exposed flawed retry logic, leading to cascading outages. Here's how to build truly resilient distributed storage operations.
Tech

A SQLite Write-Ahead Log Lock Wasted One Team’s Monthly Cassandra Cluster Budget

By Lucas Mendes/Jul 16, 2026

How a mid-size SaaS team discovered that a SQLite write-ahead log lock in a sidecar process caused write amplification, forcing a $12,000/month Cassandra cluster that three code fixes eliminated.
Tech

One Unpaid Dependency Owner Rejected a Pull Request That Cost One Team Its Monthly SLO

By Sara Park/Jul 16, 2026

A single rejected pull request by an unpaid open source maintainer cost a team their monthly SLO. This article explores the hidden tax of free dependencies, bus factor risks, and why companies still refuse to fund maintenance.
Tech

One Maintainer's RFC 2119 Fix Broke Every SPDX Header Parser for a Year

By Lucas Mendes/Jul 16, 2026

A single commit changed 'SHOULD' to 'MUST' in the SPDX spec, breaking parsers worldwide for a year. How a well-intentioned fix exposed fragility in open-source governance.
Tech

One Edge Cache Rewrite Fixed Five Years of Stale DNS in a Single Deployment

By Yusuke Tanaka/Jul 17, 2026

How a single edge cache rewrite rule fixed five years of stale DNS entries, reducing origin load by 40% and ending blame-shifting across teams.
Tech

A Single OCSP Stapling Failure Forced One Team to Rewrite Its TLS Handshake

By Yusuke Tanaka/Jul 16, 2026

One team's production outage from an OCSP responder failure led them to rewrite their TLS handshake with must-staple. A deep dive into the protocol shift and its real-world impact.
Tech

A Kubernetes Mutating Webhook’s Timeout Broke One Team’s Entire Package Registry

By Deepa Iyer/Jul 16, 2026

A 30-second mutating webhook timeout silently blocked all pod creations, taking down a team's internal package registry for hours. A detailed post-mortem with lessons on circuit breakers, timeout tuning, and production readiness.
Tech

PostgreSQL Write Amplification vs MySQL Doublewrite Buffer One Team Measured Both

By Lucas Mendes/Jul 17, 2026

A Georgia Tech study measured PostgreSQL write amplification at 1.8–2.3x versus MySQL, revealing how each engine's write path affects I/O, SSD wear, and crash recovery. Real-world tradeoffs explained.
Tech

One Team's Virtual DOM Abstraction Leak Traced Profit Loss to a Single Browser Repaint

By Yusuke Tanaka/Jul 17, 2026

A SaaS team traced a 15% profit drop to a hidden CSS animation causing 4.7-second browser repaints. The fix was one line of CSS. Here's how to catch your own repaint leaks.
Tech

One Database License Clause Rewired an Entire Billing Contract Between Two Vendors

By Sara Park/Jul 17, 2026

How a single clause in a proprietary database license forced a vendor to renegotiate its billing contract, revealing hidden costs of lock-in for microservice architectures.
Tech

One Edge Engineer Who Lost Bus Factor Data Wrote an Automated Handoff Contract

By Sara Park/Jul 17, 2026

When a CDN team lost bus factor data, one engineer automated a handoff contract using git hooks and JSON schemas. Here's how they measured risk and reduced pager fatigue.
Tech

One Build System’s Hash Collision Forced a Full CI Pipeline Rewrite

By Yusuke Tanaka/Jul 17, 2026

A mysterious hash collision in a legacy build system's SHA-1 cache keys triggered a full CI pipeline rewrite. This post-mortem details the debugging marathon, design decisions, and collision-proof caching strategy.
Tech

Transpiler Versus Transistor One Team's RISC-V Emulation Exposed a Silicon Bug

By Deepa Iyer/Jul 16, 2026

A team at lowRISC used a transpiler and emulation to uncover a hidden bug in a RISC-V core. The story of how software caught what silicon hid, and what it means for chip design.
Tech

Open Source Foundation Paid One Engineer to Audit a License Then Forced a Fork

By Deepa Iyer/Jul 17, 2026

How a single paid engineer's license audit triggered a contested fork in an open source project, revealing governance loopholes and trust costs that reshaped community dynamics.
Tech

One Postgres Write Path’s Write-Ahead Log Latency Silent Data Loss Toll

By Deepa Iyer/Jul 17, 2026

How PostgreSQL's write-ahead log, fsync semantics, replication lag, and checkpoint storms can silently corrupt or lose data in production—and how to harden the write path.
Tech

Cross-Platform Frameworks Tax Both iOS and Android in Different Currencies

By Lucas Mendes/Jul 17, 2026

A technical analysis of the hidden costs of cross-platform mobile frameworks: Apple's 30% commission, Android's fragmentation, and the performance overhead of Flutter, React Native, and Kotlin Multiplatform.
Tech

Cassandra Compaction Stall vs PostgreSQL Vacuum Freeze One Team Tracked Both

By Lucas Mendes/Jul 16, 2026

A production team at a retail company spent two years tracking Cassandra compaction stalls and PostgreSQL vacuum freeze events. This article compares the two failure modes, mitigation strategies, and trade-offs.
Tech

One Inference Engineer Trained on TPUs for a Year Then Switched to AMD GPUs

By Sara Park/Jul 17, 2026

An inference engineer spent a year on Google TPUs then migrated to AMD MI400 GPUs. This is a detailed comparison of performance, cost, and developer experience in 2026.
Tech

One Team's Four-Year CI Bill Traced to a Single Package.json Dependency

By Lucas Mendes/Jul 17, 2026

How a startup's $1.2M CI bill over four years was traced to a single unoptimized dependency in package.json, and why most teams never audit for build cost.