PostgreSQL Write Amplification vs MySQL Doublewrite Buffer One Team Measured Both

Jul 17, 2026 By Lucas Mendes

Every database write path makes a bet. PostgreSQL bets that logging the entire modified page is the safest way to guarantee crash recovery. MySQL bets that a small, dedicated buffer for double-writing pages is cheaper than writing full pages to the write-ahead log. Both bets protect data integrity, but they have very different costs in I/O, SSD wear, and recovery speed. A 2025 study by Professor Ada Chen's group at Georgia Tech put both engines on identical hardware and measured exactly what those costs are.

The Write Path Divergence

PostgreSQL's write-ahead log (WAL) records every change to the database. When a transaction modifies a row, PostgreSQL writes the entire 8 KB page containing that row to the WAL—a mechanism called full-page writes. This ensures that if a crash occurs mid-write, the WAL contains a complete copy of the page, allowing recovery to reconstruct it without depending on a partially written data file. The tradeoff is that a single row update can cause 8 KB of WAL traffic, even if the change is only a few bytes.

MySQL's InnoDB storage engine takes a different approach with the doublewrite buffer. Before writing a page to its final location in the data file, InnoDB writes a copy to a 128 KB buffer in the shared tablespace. If a crash happens during the write, the doublewrite buffer provides a clean copy to replay. This avoids writing full pages to the redo log, keeping redo log writes smaller. However, it adds an extra write step—every page is written twice: once to the doublewrite buffer and once to the data file.

The Georgia Tech team ran sysbench OLTP read-write workloads on identical hardware with PostgreSQL 16 and MySQL 8.0. PostgreSQL showed 1.8 to 2.3 times more I/O per transaction, primarily due to full-page writes inflating WAL volume. MySQL's doublewrite buffer added roughly 15% write latency per transaction, but the total I/O was consistently lower. Crash recovery tests favored MySQL, which recovered in about half the time of PostgreSQL under the same workload.

To understand these differences in practice, consider a typical e-commerce transaction: a user adds an item to their cart, updates inventory, and creates an order record. In PostgreSQL, each of these operations touches different pages (the cart page, the inventory page, the order page), and the first modification to each page after a checkpoint triggers a full-page write. That means three full 8 KB pages written to the WAL, even though the actual data changed might be only a few hundred bytes. In MySQL, the same transaction writes the three changed pages to the doublewrite buffer sequentially (a single 24 KB sequential write), then flushes them to their random locations. The sequential write is fast, and the random writes are coalesced, so total I/O is lower. A mid-sized online retailer that migrated from PostgreSQL to MySQL reported a 40% reduction in storage I/O wait time after switching, with no change in application code.

Full-Page Writes vs Doublewrite Mechanics

PostgreSQL's full-page write behavior is tied to checkpoints. Between checkpoints, the first modification to a page after a checkpoint triggers a full-page write to the WAL. Subsequent changes to the same page before the next checkpoint write only the difference (the WAL record). This means write amplification is highest for workloads that touch many distinct pages—typical in OLTP with random updates. For append-heavy workloads, amplification is lower because pages are new and written once.

MySQL's doublewrite buffer batches writes. InnoDB writes a batch of pages (up to 128 KB) sequentially to the doublewrite buffer, then flushes those pages to their random locations in the data file. The sequential write to the buffer is fast, and the subsequent random writes benefit from the buffer's ordering. On storage with a fast sequential write path (like SSDs), this overhead is modest. On spinning disks, the doublewrite can actually improve throughput by coalescing writes.

Both mechanisms prevent torn pages—pages that are partially written during a crash. Without protection, a torn page can contain a mix of old and new data, making recovery impossible without a full backup. PostgreSQL's full-page writes embed the entire page in the WAL, so recovery simply re-applies the WAL. MySQL's doublewrite buffer holds a pristine copy, so recovery reads the buffer and overwrites the corrupted page. The integrity guarantee is equivalent, but the resource cost differs.

A deeper look at the checkpoint interaction reveals an interesting nuance. In PostgreSQL, the checkpoint_completion_target parameter controls how quickly checkpoint writes are spread out. Setting it to 0.9 (the default is 0.5) can smooth I/O spikes, but it does not reduce total WAL volume. A team at a large social media company found that increasing this parameter from 0.5 to 0.9 reduced peak I/O latency by 30%, but the overall write amplification remained the same. They also experimented with full_page_writes = off on a subset of servers with ZFS storage, which eliminated full-page writes entirely. However, they had to add ZFS checksums to detect torn pages, adding CPU overhead. The tradeoff was acceptable for their workload, but they cautioned that it requires deep storage expertise.

Real-World Operational Tradeoffs

Higher write amplification directly translates to more SSD wear. For PostgreSQL, a workload that generates 100 GB of logical writes can produce 180–230 GB of WAL writes. Over a year, a team at a mid-sized e-commerce company reported 30% faster SSD depletion compared to MySQL on the same hardware. For deployments with limited SSD endurance (e.g., consumer-grade NVMe), this can force earlier replacement cycles.

MySQL's doublewrite buffer doubles the write load on storage for every page written to the data files. However, because the buffer writes are sequential, the actual I/O cost is often lower than PostgreSQL's random WAL writes. On systems with battery-backed write cache (e.g., RAID controllers with BBU), MySQL can disable the doublewrite buffer entirely, eliminating its overhead. This is a common tuning for high-performance MySQL deployments, but it sacrifices crash safety if the battery fails.

PostgreSQL benefits from fast SSDs with power-loss protection (PLP). PLP capacitors ensure that in-flight writes complete during a power failure, reducing the need for full-page writes. Some PostgreSQL teams have experimented with reducing full-page writes by increasing checkpoint intervals, but this extends recovery time after a crash. The tradeoff between write amplification and recovery speed is a constant tension in PostgreSQL tuning.

Let's examine a concrete scenario: a financial services company running a transaction processing system on PostgreSQL. Their workload is heavily update-oriented, with frequent updates to account balances and transaction records. Each update touches a different page, so full-page writes are triggered often. They measured WAL generation at roughly 2.1 times the logical write volume. Over a month, this translated to about 1.5 TB of WAL data, requiring frequent archiving and consuming significant storage. After migrating to MySQL, they observed a write amplification factor of roughly 1.6 (due to doublewrite), reducing monthly WAL+redo data to about 1.1 TB. The reduction in storage costs alone justified the migration, even accounting for the application changes needed.

Another example comes from a logistics company that uses PostgreSQL for tracking shipments. Their workload is insert-heavy—each new shipment creates a new row, and pages are appended rather than updated. They found that full-page writes were rare because most pages were written only once. Their write amplification was only about 1.2x, close to MySQL's doublewrite overhead. In this case, the difference between the two engines was negligible, and they chose to stay with PostgreSQL due to its advanced geospatial features for route optimization.

Tuning to Reduce Amplification

PostgreSQL offers several knobs to reduce write amplification. Increasing checkpoint_completion_target spreads checkpoint writes over a longer period, smoothing I/O but not reducing total WAL volume. The full_page_writes parameter can be turned off, but only on systems with reliable storage (e.g., ZFS with checksums). Some teams use WAL compression, which can reduce write volume by roughly 30% on compressible data, at the cost of additional CPU.

MySQL's main tuning lever is innodb_flush_log_at_trx_commit. Setting it to 2 (instead of the default 1) flushes the redo log once per second instead of per transaction, reducing write I/O by up to 50% for write-heavy workloads. The tradeoff is that up to one second of transactions can be lost on a crash. For many applications, this is acceptable. The doublewrite buffer itself can be disabled with innodb_doublewrite=0, but this is only recommended when using storage with atomic page writes (e.g., Fusion-io or certain NVMe drives).

Partitioning large tables spreads writes across multiple tablespaces or disks, reducing contention. For PostgreSQL, partitioning can also limit full-page writes to smaller segments. A team at a large ad-tech company reported 10–20% I/O reduction after partitioning their largest table by date. These gains are workload-specific and require careful testing.

Beyond these standard knobs, there are more advanced techniques. For PostgreSQL, using a separate WAL volume on faster storage (e.g., NVMe) can reduce the impact of write amplification on the main data volume. Some teams also use WAL archiving to a separate location to avoid I/O contention. For MySQL, tuning the size of the doublewrite buffer (via innodb_doublewrite_batch_size) can improve throughput for workloads with many small writes. The default batch size is 128 KB, but increasing it to 256 KB can reduce the number of I/O operations at the cost of slightly higher latency per batch. A database reliability engineer at a major cloud provider shared that they increased the batch size to 512 KB for a write-heavy workload and saw a 15% improvement in throughput, with no degradation in crash recovery time.

When One System Beats the Other

For write-heavy logging applications where each transaction is small and independent, MySQL's doublewrite overhead (roughly 15% per transaction) is often acceptable, especially if the storage can handle sequential writes efficiently. PostgreSQL's full-page writes can cause significantly higher WAL volume, leading to faster log rotation and more frequent archiving. In benchmarks from Percona, MySQL sustained 30% higher throughput on a sysbench write-only workload with default settings.

For mixed OLTP/OLAP workloads where some queries scan large portions of data, PostgreSQL's write amplification can hurt overall performance because the WAL competes with read I/O. A 2024 Amazon RDS performance analysis showed that PostgreSQL instances with high write amplification experienced 20–40% higher read latency during peak hours compared to MySQL instances with similar write loads. The doublewrite buffer's sequential writes interfere less with random reads.

Replica lag is another area where the difference matters. PostgreSQL full-page writes on the primary generate larger WAL segments that must be shipped to replicas. On networks with limited bandwidth, this can cause replicas to fall behind. MySQL's smaller redo log records are cheaper to ship. In a test by the Georgia Tech team, PostgreSQL replicas lagged by an average of 2.3 seconds under a 10,000 write-per-second load, while MySQL replicas stayed within 0.5 seconds.

Consider a scenario where an application requires strong consistency and low write latency. A gaming company running a leaderboard service chose PostgreSQL because its full-page writes ensure that no data is lost even in the event of a power failure. They accepted the higher I/O cost because the integrity guarantee was critical for their scoring system. In contrast, a social media analytics platform chose MySQL because their write volume was extremely high (millions of events per second) and they could tolerate a small window of data loss. They disabled the doublewrite buffer and used battery-backed cache, achieving write amplification close to 1.0x. The choice depended on their tolerance for data loss and their storage infrastructure.

Counter-Arguments and Edge Cases

Some PostgreSQL advocates argue that full-page writes are a feature, not a bug. Because the WAL contains complete page images, point-in-time recovery (PITR) is more robust—you can restore to any transaction boundary without worrying about torn pages in the data files. MySQL's redo log, by contrast, only contains changes, so recovery from a torn page requires the doublewrite buffer to be intact. If the doublewrite buffer itself is corrupted (rare but possible), recovery may fail. PostgreSQL's approach trades higher I/O for simpler recovery logic.

Another counter-argument: modern storage with atomic write sizes of 4 KB or 8 KB can eliminate torn pages entirely. Some NVMe drives guarantee that writes up to 8 KB are atomic. In that case, both full-page writes and the doublewrite buffer become redundant. However, most cloud storage (EBS, Azure managed disks) does not expose atomic write guarantees, so these mechanisms remain necessary for safety.

There is also the question of CPU overhead. PostgreSQL's WAL compression reduces I/O but consumes CPU cycles—roughly 5–10% more CPU usage on a typical server. MySQL's doublewrite buffer is I/O-bound, not CPU-bound. For CPU-constrained systems, MySQL may be preferable despite higher I/O. The Georgia Tech study noted that PostgreSQL's CPU usage was 12% higher on average for the same throughput, partly due to compression and full-page write overhead.

Workloads with very large rows (e.g., blobs) can invert the comparison. PostgreSQL's full-page writes become more expensive per row because the 8 KB page limit means many rows per page—but if a row exceeds 8 KB (TOAST storage), the overhead changes. MySQL's doublewrite buffer treats all pages equally, so large rows do not amplify the doublewrite cost. A team at a media company found that PostgreSQL's write amplification was 3x for tables with large text columns, while MySQL remained at roughly 2x due to the doublewrite buffer.

Another edge case is the use of compression at the storage layer. If the underlying filesystem or storage device compresses data (e.g., ZFS with compression, or certain SSDs with inline compression), the effective write amplification can be reduced. PostgreSQL's full-page writes may compress better than MySQL's doublewrite buffer because the pages are written sequentially in the WAL, allowing compression algorithms to find patterns. In a test with ZFS compression enabled, PostgreSQL's WAL volume shrank by 40%, while MySQL's doublewrite buffer only compressed by 20% because the pages were already partially compressed by InnoDB's page compression. This narrowed the gap between the two engines significantly.

Key Takeaways for Your Next Database Choice

Before choosing between PostgreSQL and MySQL, measure your own write patterns. If your workload is update-heavy with random access patterns, MySQL's doublewrite buffer will likely produce less I/O. If your workload is insert-heavy or append-only, PostgreSQL's full-page writes are triggered less often, narrowing the gap. Use a tool like pg_test_fsync or sysbench with your actual schema and data sizes.

Do not ignore write amplification in total cost of ownership estimates. Higher I/O means more expensive storage, faster SSD replacement, and potentially higher cloud costs (e.g., provisioned IOPS on AWS). A team at a fintech startup calculated that PostgreSQL's write amplification added roughly 15% to their monthly RDS bill compared to MySQL for the same workload. Over a year, that difference paid for an additional junior engineer.

Test crash recovery time with both engines under your workload. The Georgia Tech study found that MySQL recovered from a simulated crash about twice as fast as PostgreSQL for the same dataset. For applications with strict recovery time objectives (RTO), this could be a deciding factor. But if your team is more experienced with PostgreSQL's tooling, the operational comfort may outweigh the performance difference.

Match the protection mechanism to your storage layer. If you use enterprise SSDs with power-loss protection, PostgreSQL's full-page writes can be reduced or disabled. If you use cloud storage with built-in durability (e.g., EBS with crash-consistent snapshots), MySQL's doublewrite buffer might be redundant. There is no universal winner—only the right tradeoff for your specific hardware and durability requirements.

Finally, consider the ecosystem. PostgreSQL's logical replication and extension ecosystem (e.g., pg_stat_statements for monitoring) may justify the higher I/O cost for some teams. MySQL's managed services (Aurora, RDS) offer automated tuning that can mask the doublewrite overhead. The Georgia Tech team emphasized that their results are a starting point, not a verdict. Every workload is different, and the best way to decide is to run your own benchmarks on your own hardware.

For more on operational database surprises, see this story of a microservice retry rewrite and the database core contributor triage queue.

Recommend Posts
Tech

One Unpaid Database Core Contributor Triage Queue Hit Four Hundred Open Issues

By Lucas Mendes/Jul 16, 2026

When a single unpaid maintainer faces a triage queue of 400 open issues, the database project's bus factor becomes dangerously low. This article examines the funding gap, triage methodologies that work, and practical steps for users.
Tech

One Flaky S3 Multipart Upload Forced an Entire Microservice to Rewrite Its Retry Logic

By Deepa Iyer/Jul 16, 2026

A silent S3 multipart upload failure exposed flawed retry logic, leading to cascading outages. Here's how to build truly resilient distributed storage operations.
Tech

A SQLite Write-Ahead Log Lock Wasted One Team’s Monthly Cassandra Cluster Budget

By Lucas Mendes/Jul 16, 2026

How a mid-size SaaS team discovered that a SQLite write-ahead log lock in a sidecar process caused write amplification, forcing a $12,000/month Cassandra cluster that three code fixes eliminated.
Tech

One Unpaid Dependency Owner Rejected a Pull Request That Cost One Team Its Monthly SLO

By Sara Park/Jul 16, 2026

A single rejected pull request by an unpaid open source maintainer cost a team their monthly SLO. This article explores the hidden tax of free dependencies, bus factor risks, and why companies still refuse to fund maintenance.
Tech

One Maintainer's RFC 2119 Fix Broke Every SPDX Header Parser for a Year

By Lucas Mendes/Jul 16, 2026

A single commit changed 'SHOULD' to 'MUST' in the SPDX spec, breaking parsers worldwide for a year. How a well-intentioned fix exposed fragility in open-source governance.
Tech

One Edge Cache Rewrite Fixed Five Years of Stale DNS in a Single Deployment

By Yusuke Tanaka/Jul 17, 2026

How a single edge cache rewrite rule fixed five years of stale DNS entries, reducing origin load by 40% and ending blame-shifting across teams.
Tech

A Single OCSP Stapling Failure Forced One Team to Rewrite Its TLS Handshake

By Yusuke Tanaka/Jul 16, 2026

One team's production outage from an OCSP responder failure led them to rewrite their TLS handshake with must-staple. A deep dive into the protocol shift and its real-world impact.
Tech

A Kubernetes Mutating Webhook’s Timeout Broke One Team’s Entire Package Registry

By Deepa Iyer/Jul 16, 2026

A 30-second mutating webhook timeout silently blocked all pod creations, taking down a team's internal package registry for hours. A detailed post-mortem with lessons on circuit breakers, timeout tuning, and production readiness.
Tech

PostgreSQL Write Amplification vs MySQL Doublewrite Buffer One Team Measured Both

By Lucas Mendes/Jul 17, 2026

A Georgia Tech study measured PostgreSQL write amplification at 1.8–2.3x versus MySQL, revealing how each engine's write path affects I/O, SSD wear, and crash recovery. Real-world tradeoffs explained.
Tech

One Team's Virtual DOM Abstraction Leak Traced Profit Loss to a Single Browser Repaint

By Yusuke Tanaka/Jul 17, 2026

A SaaS team traced a 15% profit drop to a hidden CSS animation causing 4.7-second browser repaints. The fix was one line of CSS. Here's how to catch your own repaint leaks.
Tech

One Database License Clause Rewired an Entire Billing Contract Between Two Vendors

By Sara Park/Jul 17, 2026

How a single clause in a proprietary database license forced a vendor to renegotiate its billing contract, revealing hidden costs of lock-in for microservice architectures.
Tech

One Edge Engineer Who Lost Bus Factor Data Wrote an Automated Handoff Contract

By Sara Park/Jul 17, 2026

When a CDN team lost bus factor data, one engineer automated a handoff contract using git hooks and JSON schemas. Here's how they measured risk and reduced pager fatigue.
Tech

One Build System’s Hash Collision Forced a Full CI Pipeline Rewrite

By Yusuke Tanaka/Jul 17, 2026

A mysterious hash collision in a legacy build system's SHA-1 cache keys triggered a full CI pipeline rewrite. This post-mortem details the debugging marathon, design decisions, and collision-proof caching strategy.
Tech

Transpiler Versus Transistor One Team's RISC-V Emulation Exposed a Silicon Bug

By Deepa Iyer/Jul 16, 2026

A team at lowRISC used a transpiler and emulation to uncover a hidden bug in a RISC-V core. The story of how software caught what silicon hid, and what it means for chip design.
Tech

Open Source Foundation Paid One Engineer to Audit a License Then Forced a Fork

By Deepa Iyer/Jul 17, 2026

How a single paid engineer's license audit triggered a contested fork in an open source project, revealing governance loopholes and trust costs that reshaped community dynamics.
Tech

One Postgres Write Path’s Write-Ahead Log Latency Silent Data Loss Toll

By Deepa Iyer/Jul 17, 2026

How PostgreSQL's write-ahead log, fsync semantics, replication lag, and checkpoint storms can silently corrupt or lose data in production—and how to harden the write path.
Tech

Cross-Platform Frameworks Tax Both iOS and Android in Different Currencies

By Lucas Mendes/Jul 17, 2026

A technical analysis of the hidden costs of cross-platform mobile frameworks: Apple's 30% commission, Android's fragmentation, and the performance overhead of Flutter, React Native, and Kotlin Multiplatform.
Tech

Cassandra Compaction Stall vs PostgreSQL Vacuum Freeze One Team Tracked Both

By Lucas Mendes/Jul 16, 2026

A production team at a retail company spent two years tracking Cassandra compaction stalls and PostgreSQL vacuum freeze events. This article compares the two failure modes, mitigation strategies, and trade-offs.
Tech

One Inference Engineer Trained on TPUs for a Year Then Switched to AMD GPUs

By Sara Park/Jul 17, 2026

An inference engineer spent a year on Google TPUs then migrated to AMD MI400 GPUs. This is a detailed comparison of performance, cost, and developer experience in 2026.
Tech

One Team's Four-Year CI Bill Traced to a Single Package.json Dependency

By Lucas Mendes/Jul 17, 2026

How a startup's $1.2M CI bill over four years was traced to a single unoptimized dependency in package.json, and why most teams never audit for build cost.