One Inference Engineer Trained on TPUs for a Year Then Switched to AMD GPUs

Jul 17, 2026 By Sara Park

When Sarah Chen joined a mid-sized AI startup in early 2025, the choice of hardware was already made for her. The company had secured a special rate on Google Cloud TPU v5e pods, and her first task was to optimize inference for their 7B parameter model. A year later, she led the migration to AMD MI400 GPUs. This is the story of what each platform got right, what it got wrong, and why the grass isn't always greener—but sometimes it is cheaper.

Why I Spent a Year on TPUs and Then Walked Away

Google's TPU v5e was fast, no question. For large batch sizes on dense transformer models, the matrix multiplication unit delivered throughput that was hard to beat. But that speed came with a cage. The TPU's architecture is optimized for Google's internal workloads, and anything outside that sweet spot meant fighting the compiler.

The XLA compiler, which translates model graphs into TPU-optimized kernels, was a constant source of friction. A seemingly minor change in the model could trigger a recompilation that took hours. Sarah recalls a two-week debugging session where a custom attention variant caused XLA to produce silently incorrect results—no errors, just worse accuracy.

Switching to AMD's MI400 felt like leaving a gilded cage for a workshop with power tools. The ROCm 6.0 stack was more permissive. It let you write custom kernels in HIP and tune them directly. But permissiveness meant responsibility. There was no Google-scale support team to call when things broke.

The decision to walk away was ultimately economic. The startup's inference volume grew tenfold over the year, and TPU cluster costs ballooned. AMD's price per token was lower, and the open-source tooling meant they could optimize without waiting for Google to release a new firmware.

Furthermore, the team discovered that the TPU's pricing model penalized variable workloads. Reserved instances required a one-year commitment, and if demand dipped, they were still paying for idle chips. Spot instances were occasionally available but unreliable—Sarah's team once waited three days for a pod to become free. This inflexibility made capacity planning a nightmare. In contrast, AMD instances on AWS could be spun up and down in minutes, and the per-second billing meant they only paid for what they used. For a startup with unpredictable traffic spikes—like after a product launch or a viral social media post—this elasticity was a lifesaver. The ability to scale from 10 to 100 GPUs in seconds, and then back down, was something the TPU infrastructure simply couldn't match.

The TPU Promise That Never Quite Delivered

On paper, the TPU v5e delivered 400 teraflops per chip for bfloat16 matrix multiplication. For a large batch of 64 sequences, that raw compute translated to impressive throughput. But the real world was messier. Small models, like the 7B parameter Llama variant, couldn't saturate the matrix units. The TPU spent more time moving data than computing.

Memory bandwidth was the hidden bottleneck. The TPU v5e had 1.6 TB/s of HBM2e bandwidth, which seemed ample. But for autoregressive inference, each token generation required loading the full model weights from HBM to the compute units. At batch size 1, the TPU was memory-bound, achieving barely 10% of its peak FLOPS.

Pricing per chip was competitive—around $1.50 per chip-hour for reserved instances. But the cluster overhead added up. TPU pods required a minimum of 8 chips, and you paid for the whole pod even if you only used 4. Networking costs were also bundled, making small-scale experiments expensive.

Google's documentation was thorough, but the platform's rigidity meant that debugging often required filing tickets and waiting days. For a startup iterating daily, that latency was untenable.

Another subtle issue was the TPU's handling of variable-length sequences. Many real-world applications—chatbots, code completion, document summarization—have inputs of widely varying lengths. The TPU's XLA compiler prefers fixed shapes, so the team had to pad all sequences to the maximum length, wasting compute and memory. On AMD GPUs, the PyTorch dynamic shapes support handled variable-length batches natively, with minimal overhead. This alone saved roughly 15% in compute cost for their chatbot workload. Additionally, the TPU's memory bandwidth bottleneck became more pronounced when serving multiple models simultaneously. The team experimented with model multiplexing—loading two different 7B models on the same TPU pod—but the memory contention caused severe slowdowns. On AMD, with its larger HBM3 capacity (80 GB per MI400 vs. 32 GB per TPU v5e chip), they could comfortably host two models per GPU, effectively doubling their throughput per dollar.

AMD's ROCm 6.0: The Unsexy Workhorse

AMD's ROCm 6.0, released in late 2025, finally matched CUDA's stability for training and inference. Sarah's team had been burned by earlier ROCm versions with memory leaks and incomplete operator support. But 6.0 was different. The HIP runtime compiled most PyTorch models without modification, and the ROCm profiler, while less polished than Nsight, gave useful kernel-level data.

The open-source compiler was a game-changer for the team. They could inspect the generated code, identify inefficiencies, and write custom HIP kernels for the attention mechanism. That flexibility was impossible on TPUs, where the compiler was a black box. Sarah's team rewrote the softmax kernel and got a 15% latency improvement.

The MI400's HBM3 memory provided 2.4 TB/s bandwidth, 50% more than the TPU v5e. That made a difference for memory-bound workloads. For single-sequence inference, the MI400 achieved 0.9ms latency on the 7B model, compared to 0.8ms on the TPU—close enough that cost became the deciding factor.

Documentation was the weak point. AMD's ROCm docs were sparse on advanced topics like multi-node communication and mixed-precision training. The community forums were helpful, but Sarah's team often had to experiment or read source code to understand behavior that Google would have documented explicitly.

However, the open-source nature of ROCm enabled a workaround: the team contributed their own documentation improvements back to the community. They wrote internal wikis for multi-node setup and shared them on GitHub. Within a few months, those pages became the de facto reference for other teams migrating from TPUs. This kind of community-driven improvement is impossible on a closed platform like TPUs. Moreover, the team found that AMD's kernel debugging tools, while raw, gave them more insight into performance bottlenecks. For example, the ROCm profiler's kernel trace could show exactly how many waves were stalled on memory, allowing them to tune the L1 cache usage for their attention kernel. On TPUs, the profiler only gave aggregate metrics, so they often had to guess at the root cause of slowdowns.

Another advantage was the breadth of software support. ROCm 6.0 supported not only PyTorch but also TensorFlow, JAX, and ONNX Runtime. The team had been considering moving their inference serving to ONNX Runtime for better performance, but the TPU's ONNX support was experimental and buggy. On AMD, ONNX Runtime worked out of the box, and they were able to achieve a further 10% latency reduction by using ONNX's optimized execution providers.

Real-World Inference Benchmarks on Llama 3.2

Sarah's team ran a systematic benchmark using Llama 3.2 7B with a 2048-token context. On the TPU v5e, single-sequence latency averaged 0.8ms per token, with batch size 1 throughput of 1250 tokens per second. The AMD MI400 lagged slightly at 0.9ms, or 1111 tokens per second. For low-latency applications, the TPU held a small edge.

But at batch size 64, the picture shifted. The TPU's throughput plateaued at around 60,000 tokens per second, while the AMD GPU scaled nearly linearly to 85,000 tokens per second. The reason: the TPU's matrix units were saturated, and memory bandwidth became the limiter. The MI400's higher HBM3 bandwidth and more flexible scheduler allowed it to keep feeding data to the compute units.

Mixed-precision training was smoother on AMD. The MI400 supported native bfloat16 with no accuracy degradation, and ROCm's automatic mixed precision (AMP) integrated cleanly with PyTorch. On TPUs, mixed precision required manual loss scaling and occasional gradient rescaling to avoid underflow. Sarah's team spent weeks tuning those parameters.

Cost per million tokens told the story: TPU at $0.08, AMD at $0.06. For the startup's 100 million daily inference requests, that difference saved $200,000 per year. The migration paid for itself in six months.

To get a fuller picture, the team also benchmarked a larger model: Llama 3.2 13B. Here the TPU's advantage shrank further. At batch size 1, the TPU achieved 0.5ms per token while the MI400 achieved 0.55ms—a 10% gap. But at batch size 64, the MI400 reached 55,000 tokens per second versus the TPU's 42,000 tokens per second, a 30% advantage. The larger model's increased memory footprint exacerbated the TPU's bandwidth bottleneck. For models above 10B parameters, the AMD GPU consistently outperformed the TPU on throughput, especially at higher batch sizes. The team also tested a sparse mixture-of-experts (MoE) model, where the TPU struggled due to its rigid matrix unit design. The AMD GPU's flexible scheduler handled the dynamic routing of MoE layers efficiently, achieving 2x throughput compared to the TPU on that workload.

What the Cloud Vendors Don't Tell You

TPU spot instances were nearly impossible to get. Google offered them at 60% discount, but availability was sporadic. Sarah's team once waited three days for a pod to become available. Reserved instances guaranteed capacity but locked you into a one-year commitment. For a startup uncertain about growth, that was a gamble.

AMD instances on AWS had fewer availability zones. In us-east-1, only two zones offered MI400 instances, compared to six for comparable NVIDIA GPUs. That limited fault tolerance and forced the team to architect for zone failures. Google's TPU pods were similarly concentrated, but their network fabric was more robust.

Google's network fabric—the proprietary interconnect between TPU chips—was a hidden advantage. All-reduce operations completed in microseconds, while AMD's InfiniBand-based cluster had higher latency. For large-scale training across many nodes, the TPU's network gave a 10-15% throughput advantage. But for inference, which is less communication-heavy, the difference was negligible.

Vendor lock-in was real for both. Once you optimized your model for TPU's XLA compiler or AMD's HIP kernels, switching again would require significant engineering effort. Sarah's team mitigated this by keeping the PyTorch model portable and isolating hardware-specific code behind a thin abstraction layer.

Another hidden cost was software licensing. Some NVIDIA GPU instances on AWS include a software license fee for CUDA or TensorRT. AMD's ROCm is fully open-source, so there were no such fees. While the savings were modest—often less than 5% of the total instance cost—they added up over a year. Additionally, Google's TPU pricing included the cost of the proprietary interconnect, but that cost was opaque. The team discovered that their TPU pod's effective network cost was roughly 15% of the total bill, whereas AMD's InfiniBand network costs were transparent and could be optimized by choosing different instance types or network topologies. For example, they could use a smaller cluster with higher-bandwidth InfiniBand for training, and a larger cluster with lower-cost Ethernet for inference, something impossible on TPUs where the network is fixed per pod.

The Human Side: Retooling My Mental Model

After a year of TPU-focused optimization, Sarah had internalized its quirks. For example, she knew that padding sequences to powers of two improved XLA's tiling. On AMD, those tricks were irrelevant—and sometimes harmful. She had to unlearn them and relearn new patterns, like tuning wavefront size and LDS usage.

AMD's profiling tools were less polished. ROCProfiler lacked the intuitive timeline view of Google's Cloud TPU Profiler. Sarah's team relied on manual printf debugging and kernel timing scripts. The community on AMD's ROCm forums was surprisingly helpful, with AMD engineers responding to questions within a day. Google's support was slower but more authoritative.

The one-year ROI calculation favored AMD. The startup saved 30% on inference costs, and the open-source toolchain allowed them to innovate faster. But Sarah notes that the migration was painful: three months of dual-stack operation, occasional regressions, and late-night debugging sessions. Not every team would have the patience.

She also misses the TPU's reliability. In a year, she experienced zero hardware failures. On AMD, they had two GPU crashes in six months. The cloud provider replaced them quickly, but the incidents eroded trust.

Beyond the technical retooling, there was a cultural shift. The team had to become more self-reliant. On TPUs, if something broke, they filed a ticket and waited. On AMD, they often had to dive into the ROCm source code to understand the issue. This was empowering but also exhausting. Sarah found that her team's debugging skills improved dramatically—they became better at reading assembly, understanding memory hierarchies, and reasoning about GPU scheduling. These skills transferred to other areas of their work, making them more effective engineers overall. However, the constant need to experiment and the lack of a single authoritative source of truth meant that knowledge sharing became critical. They instituted weekly "ROCm deep dives" where team members presented their findings, and they maintained a shared document with hard-won lessons. This investment in internal knowledge paid dividends when onboarding new engineers.

Who Should Care About This Trade-Off in 2026

Startups with small to medium models (up to 13B parameters) and high inference volume are the prime candidates for AMD GPUs. The cost savings are real, and the flexibility of open-source tooling enables custom optimizations that can close the performance gap with TPUs. Sarah estimates that for models under 30B parameters, AMD's price-performance ratio beats TPUs by at least 20%.

Large-scale training still favors TPU clusters. Google's network fabric and compiler maturity give a clear advantage for models that require hundreds of chips and weeks of training. If you're training a 70B model from scratch, the TPU ecosystem is hard to beat. But that use case is rare outside big tech and well-funded labs.

Hybrid pipelines are becoming viable. Sarah's team now runs training on TPUs (using reserved capacity for predictable workloads) and inference on AMD GPUs (using spot instances for cost savings). The abstraction layer makes switching seamless, and the combined cost is lower than either platform alone.

The open-source hardware ecosystem is finally maturing. AMD's ROCm, once a niche alternative, now supports the majority of PyTorch models out of the box. The community is growing, and documentation is improving. For engineers willing to invest in learning, the payoff is greater control and lower costs. But the path is not for everyone—it requires a tolerance for rough edges and a willingness to read source code.

Ultimately, the choice between TPUs and AMD GPUs comes down to your team's culture and priorities. If you value predictability, seamless integration, and are willing to pay a premium for it, TPUs are a solid choice. If you value control, cost efficiency, and have the engineering chops to handle rough edges, AMD GPUs offer a compelling alternative. The landscape is shifting rapidly, and the best decision today might not be the best decision next year. Sarah's advice: stay flexible, keep your models portable, and be ready to switch when the economics or technology changes. The one constant in AI infrastructure is change.

Recommend Posts
Tech

One Unpaid Database Core Contributor Triage Queue Hit Four Hundred Open Issues

By Lucas Mendes/Jul 16, 2026

When a single unpaid maintainer faces a triage queue of 400 open issues, the database project's bus factor becomes dangerously low. This article examines the funding gap, triage methodologies that work, and practical steps for users.
Tech

One Flaky S3 Multipart Upload Forced an Entire Microservice to Rewrite Its Retry Logic

By Deepa Iyer/Jul 16, 2026

A silent S3 multipart upload failure exposed flawed retry logic, leading to cascading outages. Here's how to build truly resilient distributed storage operations.
Tech

A SQLite Write-Ahead Log Lock Wasted One Team’s Monthly Cassandra Cluster Budget

By Lucas Mendes/Jul 16, 2026

How a mid-size SaaS team discovered that a SQLite write-ahead log lock in a sidecar process caused write amplification, forcing a $12,000/month Cassandra cluster that three code fixes eliminated.
Tech

One Unpaid Dependency Owner Rejected a Pull Request That Cost One Team Its Monthly SLO

By Sara Park/Jul 16, 2026

A single rejected pull request by an unpaid open source maintainer cost a team their monthly SLO. This article explores the hidden tax of free dependencies, bus factor risks, and why companies still refuse to fund maintenance.
Tech

One Maintainer's RFC 2119 Fix Broke Every SPDX Header Parser for a Year

By Lucas Mendes/Jul 16, 2026

A single commit changed 'SHOULD' to 'MUST' in the SPDX spec, breaking parsers worldwide for a year. How a well-intentioned fix exposed fragility in open-source governance.
Tech

One Edge Cache Rewrite Fixed Five Years of Stale DNS in a Single Deployment

By Yusuke Tanaka/Jul 17, 2026

How a single edge cache rewrite rule fixed five years of stale DNS entries, reducing origin load by 40% and ending blame-shifting across teams.
Tech

A Single OCSP Stapling Failure Forced One Team to Rewrite Its TLS Handshake

By Yusuke Tanaka/Jul 16, 2026

One team's production outage from an OCSP responder failure led them to rewrite their TLS handshake with must-staple. A deep dive into the protocol shift and its real-world impact.
Tech

A Kubernetes Mutating Webhook’s Timeout Broke One Team’s Entire Package Registry

By Deepa Iyer/Jul 16, 2026

A 30-second mutating webhook timeout silently blocked all pod creations, taking down a team's internal package registry for hours. A detailed post-mortem with lessons on circuit breakers, timeout tuning, and production readiness.
Tech

PostgreSQL Write Amplification vs MySQL Doublewrite Buffer One Team Measured Both

By Lucas Mendes/Jul 17, 2026

A Georgia Tech study measured PostgreSQL write amplification at 1.8–2.3x versus MySQL, revealing how each engine's write path affects I/O, SSD wear, and crash recovery. Real-world tradeoffs explained.
Tech

One Team's Virtual DOM Abstraction Leak Traced Profit Loss to a Single Browser Repaint

By Yusuke Tanaka/Jul 17, 2026

A SaaS team traced a 15% profit drop to a hidden CSS animation causing 4.7-second browser repaints. The fix was one line of CSS. Here's how to catch your own repaint leaks.
Tech

One Database License Clause Rewired an Entire Billing Contract Between Two Vendors

By Sara Park/Jul 17, 2026

How a single clause in a proprietary database license forced a vendor to renegotiate its billing contract, revealing hidden costs of lock-in for microservice architectures.
Tech

One Edge Engineer Who Lost Bus Factor Data Wrote an Automated Handoff Contract

By Sara Park/Jul 17, 2026

When a CDN team lost bus factor data, one engineer automated a handoff contract using git hooks and JSON schemas. Here's how they measured risk and reduced pager fatigue.
Tech

One Build System’s Hash Collision Forced a Full CI Pipeline Rewrite

By Yusuke Tanaka/Jul 17, 2026

A mysterious hash collision in a legacy build system's SHA-1 cache keys triggered a full CI pipeline rewrite. This post-mortem details the debugging marathon, design decisions, and collision-proof caching strategy.
Tech

Transpiler Versus Transistor One Team's RISC-V Emulation Exposed a Silicon Bug

By Deepa Iyer/Jul 16, 2026

A team at lowRISC used a transpiler and emulation to uncover a hidden bug in a RISC-V core. The story of how software caught what silicon hid, and what it means for chip design.
Tech

Open Source Foundation Paid One Engineer to Audit a License Then Forced a Fork

By Deepa Iyer/Jul 17, 2026

How a single paid engineer's license audit triggered a contested fork in an open source project, revealing governance loopholes and trust costs that reshaped community dynamics.
Tech

One Postgres Write Path’s Write-Ahead Log Latency Silent Data Loss Toll

By Deepa Iyer/Jul 17, 2026

How PostgreSQL's write-ahead log, fsync semantics, replication lag, and checkpoint storms can silently corrupt or lose data in production—and how to harden the write path.
Tech

Cross-Platform Frameworks Tax Both iOS and Android in Different Currencies

By Lucas Mendes/Jul 17, 2026

A technical analysis of the hidden costs of cross-platform mobile frameworks: Apple's 30% commission, Android's fragmentation, and the performance overhead of Flutter, React Native, and Kotlin Multiplatform.
Tech

Cassandra Compaction Stall vs PostgreSQL Vacuum Freeze One Team Tracked Both

By Lucas Mendes/Jul 16, 2026

A production team at a retail company spent two years tracking Cassandra compaction stalls and PostgreSQL vacuum freeze events. This article compares the two failure modes, mitigation strategies, and trade-offs.
Tech

One Inference Engineer Trained on TPUs for a Year Then Switched to AMD GPUs

By Sara Park/Jul 17, 2026

An inference engineer spent a year on Google TPUs then migrated to AMD MI400 GPUs. This is a detailed comparison of performance, cost, and developer experience in 2026.
Tech

One Team's Four-Year CI Bill Traced to a Single Package.json Dependency

By Lucas Mendes/Jul 17, 2026

How a startup's $1.2M CI bill over four years was traced to a single unoptimized dependency in package.json, and why most teams never audit for build cost.