One Edge Cache Rewrite Fixed Five Years of Stale DNS in a Single Deployment

Jul 17, 2026 By Yusuke Tanaka

For five years, a mid-sized content delivery network suffered from stale DNS entries that pointed to decommissioned origin IPs. The root cause was a single edge cache rewrite rule that had been misconfigured during a routine deployment. When a junior engineer finally fixed it in one afternoon, the team saw an immediate 40% reduction in origin query load and the end of a blame-shifting cycle that had spanned multiple teams. This is the story of how a protocol-level fix—using EDNS0 option codes to force cache eviction—solved a problem that standard DNS practices had failed to address for half a decade.

The Five-Year Stale DNS Bug That Nobody Owned

DNS entries for several production domains pointed to origin IPs that had been decommissioned in 2021. The TTL was set to 86400 seconds—24 hours—so any cache miss propagated slowly. Worse, the edge cache rewrite rule that should have updated these records was itself stale: it had been written for an older version of the CDN software and never updated after a major platform migration.

Multiple teams blamed each other. The networking team said the DNS configuration was correct at the authoritative level. The platform team pointed at the edge cache layer. The SRE team had no monitoring for DNS resolution failures because nobody thought to check. The result was a slow bleed of traffic to dead IPs, causing intermittent 502 errors that were dismissed as "transient network issues."

One engineer, while cleaning up old Terraform state files, noticed the rewrite rule had a hardcoded reference to an origin pool that no longer existed. The rule was supposed to map a set of legacy hostnames to new A and AAAA records, but the mapping had never been applied because the cache key logic didn't honor the new TTL values. The fix was a single line change in a configuration file.

The deployment took 45 minutes, including canary testing. Within an hour, error rates dropped to zero. The team celebrated, but the real lesson was organizational: a five-year-old bug survived because nobody owned the full path from DNS resolution to cache eviction.

How the Edge Cache Rewrite Actually Works

Edge cache rewrites intercept DNS responses at the CDN layer, before the response reaches the client. When a resolver queries for a domain, the edge cache checks its local store. If the record is stale or missing, it fetches from the origin, but the rewrite rule can modify the response in-flight—replacing old IPs with new ones, or adjusting TTL values.

In this case, the rewrite rule was implemented as a Lua script running on the CDN's edge nodes. The script inspected the EDNS0 option codes in the incoming query and, if the domain matched a predefined list, replaced the answer section with fresh A and AAAA mappings. No changes were made to the authoritative nameserver configuration; the rewrite operated entirely at the edge.

The cache key logic was updated to include a version hash derived from the rewrite rule itself. When the rule changed, the cache key changed, forcing all cached entries to be evicted. This was the critical piece that had been missing: the old rule had a static cache key that never invalidated, so even after the origin was updated, the edge kept serving the stale records for the full TTL duration.

Deploying the fix required a Terraform plan that updated the edge configuration across 12 regions. The canary tested one region for 10 minutes, then rolled out globally. The team used a synthetic monitoring tool to verify that DNS resolution returned the correct IPs from every edge location. The entire process was automated, but the key insight was the cache key invalidation—a detail that had been overlooked in the original implementation.

Why Standard DNS Practices Failed Here

Standard DNS practices rely on TTLs and authoritative server updates. But in complex CDN deployments, the path from authoritative server to client is long and layered. The CDN's edge cache sits between the resolver and the origin, and it can cache DNS responses independently. If the edge cache ignores TTL updates—or if the TTL is set too high—stale records persist.

Vendor lock-in prevented the team from using custom resolver hooks. The CDN provider's API allowed rewriting HTTP responses but not DNS responses directly. The team had to work around this by embedding the rewrite logic in the edge compute layer, which was not designed for DNS manipulation. This created a fragile coupling between HTTP and DNS configuration.

Legacy glue records were hardcoded in configuration files that were checked into version control but never reviewed. These records pointed to IPs that had been decommissioned during a data center migration. No monitoring existed for DNS resolution failures because the team assumed the authoritative server was the source of truth. In reality, the edge cache was serving stale data that masked the problem.

RFC 2181 specifies that DNS resolvers should honor TTL values, but many middleware implementations—including the CDN's edge cache—had bugs that caused them to ignore TTL updates under certain conditions. The team discovered that the edge cache was using a default TTL of 86400 seconds even when the authoritative server returned a lower value. This was a known issue in the CDN software, but the vendor had not prioritized a fix.

The Protocol-Level Fix Nobody Talks About

The rewrite uses EDNS0 option codes for cache hinting. EDNS0, defined in RFC 6891, allows DNS messages to carry additional options beyond the standard header. By inserting a custom option code in the response, the edge cache can signal to downstream resolvers that the record should be evicted immediately.

In this implementation, the rewrite rule adds an EDNS0 option with a vendor-specific code that the CDN's edge nodes recognize. When the edge node sees this option, it bypasses its normal cache logic and forces a fresh fetch from the origin. This effectively overrides the TTL for that specific response, without requiring changes to the authoritative server.

The result is a 40% reduction in origin query load in early tests. Because stale entries are evicted sooner, the edge cache serves more fresh responses, reducing the number of queries that reach the origin. The fix works with any DNS-over-HTTPS or DNS-over-TLS setup because the EDNS0 option is preserved across transport layers.

Open-source tooling is now available for replication. A GitHub repository provides a Lua script and Terraform module that can be adapted to any CDN that supports edge compute. The script includes a test suite that validates EDNS0 option handling across multiple resolvers. This makes the fix accessible to teams that lack the in-house expertise to develop it from scratch.

Trade-Offs and Alternatives: Why Not Just Lower the TTL?

A common first reaction is to simply lower the TTL on the authoritative records. But that approach has its own costs. A TTL of, say, 300 seconds means resolvers will query the origin more frequently, increasing load on the authoritative nameserver and potentially increasing latency for end users. In a CDN with millions of queries per day, a low TTL can multiply origin traffic by a factor of 10 or more. The rewrite rule, by contrast, only triggers cache eviction when the record actually changes—most of the time, the edge serves cached data with a longer TTL, keeping origin load low.

Another alternative is to use DNS-based load balancing with health checks, such as AWS Route 53 latency-based routing or Azure Traffic Manager. These services can detect origin failures and redirect traffic automatically. However, they operate at the authoritative level and cannot override a misconfigured edge cache. In this case, the authoritative server was returning correct records, but the edge cache was ignoring them. Health checks would not have helped because the authoritative path was healthy.

There is also the option of using a custom DNS resolver within the CDN, such as CoreDNS or Unbound, configured with aggressive caching policies. This would give the team more control over TTL handling and cache invalidation. But it adds operational complexity—another service to deploy, monitor, and patch. The rewrite rule, once written, requires minimal maintenance and integrates with existing edge compute infrastructure.

Each approach has its place. For teams with low traffic and simple architectures, lowering the TTL might be sufficient. For those with high traffic and complex multi-vendor setups, the EDNS0 rewrite provides a surgical fix without disrupting the rest of the stack. The key is to understand where the staleness originates—in the edge cache, not the authoritative server—and apply the fix at that layer.

Lessons from OnePlus's Exit and Starlink's Fragility

OnePlus's recent exit from phone releases in the US and Europe, as reported by Ars Technica, highlights the cost of ignoring edge infrastructure. OnePlus's decision to pull out was driven by market pressures, but the underlying lesson is that companies that neglect their infrastructure—including DNS and caching—eventually pay the price. A stale DNS entry might seem minor, but it erodes user trust and increases support costs.

The Starlink debate, also covered by Ars Technica, underscores DNS as an attack surface. The article discusses how China and Russia could potentially disrupt Starlink's satellite network. While the focus is on physical attacks, the DNS layer is equally vulnerable. Starlink's routing tables depend on fast, accurate DNS resolution; a cache poisoning attack could redirect traffic or degrade performance. The rewrite pattern described here could be applied to satellite routing tables to enforce cache invalidation at the edge.

Both cases demand resilient cache strategies. OnePlus's exit shows that infrastructure neglect can contribute to business decline. Starlink's fragility demonstrates that even cutting-edge systems rely on basic DNS hygiene. The rewrite pattern is not a silver bullet, but it provides a practical mechanism for enforcing cache consistency in environments where standard DNS practices fall short.

The industry must standardize cache invalidation hooks. Today, each CDN vendor implements its own proprietary mechanism. A standardized EDNS0 option code for cache invalidation would allow portable configurations and reduce the risk of vendor lock-in. Until then, teams must build their own solutions—and the open-source tooling from this case study is a starting point.

Counter-Argument: When Edge Cache Rewrites Are Not the Answer

Not every stale DNS problem benefits from an edge cache rewrite. If the authoritative server itself returns incorrect records, the rewrite will only propagate those errors faster. The team must first verify that the authoritative DNS is correct—otherwise, they risk amplifying a mistake.

Edge cache rewrites also add a layer of complexity. The Lua script must be maintained, tested, and deployed alongside other edge compute logic. If the script contains a bug, it could corrupt DNS responses for all domains, not just the targeted ones. The team mitigated this by using canary deployments and synthetic monitoring, but smaller teams without those resources might find the risk unacceptable.

Moreover, the rewrite relies on vendor-specific EDNS0 option codes. If the CDN provider changes its edge compute platform or stops supporting custom Lua scripts, the rewrite breaks. The team would then need to migrate to a different mechanism, potentially repeating the whole exercise. This is a form of vendor lock-in that the team accepted because the alternative—years of stale DNS—was worse.

For some organizations, the simplest fix is to eliminate the edge cache layer for DNS entirely. If the CDN's DNS caching is unreliable, they can configure resolvers to bypass the edge and query the authoritative server directly. This trades higher latency for consistency. In this case, the team chose the rewrite because the latency impact of bypassing the edge was unacceptable for their real-time applications.

Three Practical Takeaways for Your Infrastructure

First, audit all DNS cache layers with quarterly sweeps. Use synthetic monitoring to verify that DNS resolution returns the expected IPs from every edge location. Automate the sweep with a script that checks against a known baseline and alerts on discrepancies. This catches stale entries before they cause user-facing errors.

Second, implement an EDNS0-based rewrite as a safety net. Even if your authoritative DNS is correct, the edge cache may serve stale data. A rewrite rule that forces cache eviction for critical domains ensures that updates propagate quickly. Pair this with canary deployments to test the rewrite before rolling out globally.

Third, document ownership per cache entry to avoid blame-shifting. Each DNS record should have a designated owner who is responsible for keeping it current. This owner should be listed in the configuration file and notified of any changes to the underlying infrastructure. Without clear ownership, stale entries can persist for years.

Finally, test TTL overrides in staging before production. The rewrite rule should be validated in a staging environment that mirrors production traffic patterns. Use a tool like this S3 multipart upload retry rewrite as a reference for how to design robust retry logic in edge compute. Similarly, this OCSP stapling failure case shows how a single protocol oversight can cascade into a widespread issue. The rewrite pattern is not a cure-all, but it is a practical tool for maintaining DNS hygiene in complex deployments.

Recommend Posts
Tech

One Unpaid Database Core Contributor Triage Queue Hit Four Hundred Open Issues

By Lucas Mendes/Jul 16, 2026

When a single unpaid maintainer faces a triage queue of 400 open issues, the database project's bus factor becomes dangerously low. This article examines the funding gap, triage methodologies that work, and practical steps for users.
Tech

One Flaky S3 Multipart Upload Forced an Entire Microservice to Rewrite Its Retry Logic

By Deepa Iyer/Jul 16, 2026

A silent S3 multipart upload failure exposed flawed retry logic, leading to cascading outages. Here's how to build truly resilient distributed storage operations.
Tech

A SQLite Write-Ahead Log Lock Wasted One Team’s Monthly Cassandra Cluster Budget

By Lucas Mendes/Jul 16, 2026

How a mid-size SaaS team discovered that a SQLite write-ahead log lock in a sidecar process caused write amplification, forcing a $12,000/month Cassandra cluster that three code fixes eliminated.
Tech

One Unpaid Dependency Owner Rejected a Pull Request That Cost One Team Its Monthly SLO

By Sara Park/Jul 16, 2026

A single rejected pull request by an unpaid open source maintainer cost a team their monthly SLO. This article explores the hidden tax of free dependencies, bus factor risks, and why companies still refuse to fund maintenance.
Tech

One Maintainer's RFC 2119 Fix Broke Every SPDX Header Parser for a Year

By Lucas Mendes/Jul 16, 2026

A single commit changed 'SHOULD' to 'MUST' in the SPDX spec, breaking parsers worldwide for a year. How a well-intentioned fix exposed fragility in open-source governance.
Tech

One Edge Cache Rewrite Fixed Five Years of Stale DNS in a Single Deployment

By Yusuke Tanaka/Jul 17, 2026

How a single edge cache rewrite rule fixed five years of stale DNS entries, reducing origin load by 40% and ending blame-shifting across teams.
Tech

A Single OCSP Stapling Failure Forced One Team to Rewrite Its TLS Handshake

By Yusuke Tanaka/Jul 16, 2026

One team's production outage from an OCSP responder failure led them to rewrite their TLS handshake with must-staple. A deep dive into the protocol shift and its real-world impact.
Tech

A Kubernetes Mutating Webhook’s Timeout Broke One Team’s Entire Package Registry

By Deepa Iyer/Jul 16, 2026

A 30-second mutating webhook timeout silently blocked all pod creations, taking down a team's internal package registry for hours. A detailed post-mortem with lessons on circuit breakers, timeout tuning, and production readiness.
Tech

PostgreSQL Write Amplification vs MySQL Doublewrite Buffer One Team Measured Both

By Lucas Mendes/Jul 17, 2026

A Georgia Tech study measured PostgreSQL write amplification at 1.8–2.3x versus MySQL, revealing how each engine's write path affects I/O, SSD wear, and crash recovery. Real-world tradeoffs explained.
Tech

One Team's Virtual DOM Abstraction Leak Traced Profit Loss to a Single Browser Repaint

By Yusuke Tanaka/Jul 17, 2026

A SaaS team traced a 15% profit drop to a hidden CSS animation causing 4.7-second browser repaints. The fix was one line of CSS. Here's how to catch your own repaint leaks.
Tech

One Database License Clause Rewired an Entire Billing Contract Between Two Vendors

By Sara Park/Jul 17, 2026

How a single clause in a proprietary database license forced a vendor to renegotiate its billing contract, revealing hidden costs of lock-in for microservice architectures.
Tech

One Edge Engineer Who Lost Bus Factor Data Wrote an Automated Handoff Contract

By Sara Park/Jul 17, 2026

When a CDN team lost bus factor data, one engineer automated a handoff contract using git hooks and JSON schemas. Here's how they measured risk and reduced pager fatigue.
Tech

One Build System’s Hash Collision Forced a Full CI Pipeline Rewrite

By Yusuke Tanaka/Jul 17, 2026

A mysterious hash collision in a legacy build system's SHA-1 cache keys triggered a full CI pipeline rewrite. This post-mortem details the debugging marathon, design decisions, and collision-proof caching strategy.
Tech

Transpiler Versus Transistor One Team's RISC-V Emulation Exposed a Silicon Bug

By Deepa Iyer/Jul 16, 2026

A team at lowRISC used a transpiler and emulation to uncover a hidden bug in a RISC-V core. The story of how software caught what silicon hid, and what it means for chip design.
Tech

Open Source Foundation Paid One Engineer to Audit a License Then Forced a Fork

By Deepa Iyer/Jul 17, 2026

How a single paid engineer's license audit triggered a contested fork in an open source project, revealing governance loopholes and trust costs that reshaped community dynamics.
Tech

One Postgres Write Path’s Write-Ahead Log Latency Silent Data Loss Toll

By Deepa Iyer/Jul 17, 2026

How PostgreSQL's write-ahead log, fsync semantics, replication lag, and checkpoint storms can silently corrupt or lose data in production—and how to harden the write path.
Tech

Cross-Platform Frameworks Tax Both iOS and Android in Different Currencies

By Lucas Mendes/Jul 17, 2026

A technical analysis of the hidden costs of cross-platform mobile frameworks: Apple's 30% commission, Android's fragmentation, and the performance overhead of Flutter, React Native, and Kotlin Multiplatform.
Tech

Cassandra Compaction Stall vs PostgreSQL Vacuum Freeze One Team Tracked Both

By Lucas Mendes/Jul 16, 2026

A production team at a retail company spent two years tracking Cassandra compaction stalls and PostgreSQL vacuum freeze events. This article compares the two failure modes, mitigation strategies, and trade-offs.
Tech

One Inference Engineer Trained on TPUs for a Year Then Switched to AMD GPUs

By Sara Park/Jul 17, 2026

An inference engineer spent a year on Google TPUs then migrated to AMD MI400 GPUs. This is a detailed comparison of performance, cost, and developer experience in 2026.
Tech

One Team's Four-Year CI Bill Traced to a Single Package.json Dependency

By Lucas Mendes/Jul 17, 2026

How a startup's $1.2M CI bill over four years was traced to a single unoptimized dependency in package.json, and why most teams never audit for build cost.