Table of Contents
Google gets a Marvell stock warrant worth up to $12.2 billion
In sum – what we know:
- A warrant, not a chip swap – Marvell issued Google a stock warrant on August 18, 2026, for up to 58,970,907 shares at $206.58, worth roughly $12.2 billion if fully exercised.
- Near-memory compute is the point – The deal covers inference accelerators, storage controllers, memory interface controllers, and NICs designed to cut the data movement that bottlenecks LLM inference.
- Custom silicon keeps expanding – Google joins Amazon’s Trainium and Microsoft’s Maia in deepening bespoke chip efforts, pressuring Nvidia to lean on CUDA and its Mellanox networking edge.
Marvell Technology and Google have agreed to co-design custom semiconductor products for Google’s tensor processing unit (TPU) ecosystem. Alongside the deal, Marvell issued Google a stock warrant on August 18, 2026, allowing it to purchase up to 58,970,907 shares at $206.58 apiece — a potential investment of roughly $12.2 billion if fully exercised.
To be clear, this isn’t Marvell building Google’s next TPU. The Marvell-designed chips will sit alongside or “attach to the TPU ecosystem” rather than replace them. The scope includes AI inference accelerators that could be built offload workloads from the primary TPUs — specialized silicon for the parts of the stack where general-purpose compute is the wrong tool.
Pushing near-memory compute
The less glamorous chips are the most interesting part of the announcement. Beyond inference accelerators, the collaboration covers custom storage controllers, memory interface controllers, and specialized network interface controllers (NICs) for cluster-level AI workloads. That list points toward near-memory compute — architectures that move computation closer to the memory modules themselves, cutting down on the data movement that dominates both latency and energy consumption in modern AI systems.
For LLM inference specifically, that matters a lot. Attention mechanisms require heavy, repeated access to large parameter matrices, which means one of the bottlenecks is often how fast data can shuttle between memory and compute rather than how fast the compute itself runs. Placing logic near the memory attacks that constraint directly, improving throughput where it actually counts. A dedicated inference chip optimized for transformer models should also deliver better performance-per-watt and lower latency than a general-purpose accelerator — the same logic that drove Google to build TPUs in the first place, now applied a layer deeper in the stack.
This is also where custom silicon diverges from the merchant model. Nvidia’s GPUs have to serve everyone, from research labs to hyperscalers to enterprises running mixed workloads, and that breadth carries a cost. The Marvell chips can be tailored exclusively to Google’s workloads and Google’s data center topology, with no compromises for anyone else’s use case. And because the agreement spans storage controllers through NICs, the two companies can co-design the entire data path — from SSDs to memory to network fabric — around TPU clusters, rather than optimizing each component in isolation. That system-level approach is hard to replicate with off-the-shelf parts, regardless of how fast any individual chip is.
Shifts in the custom silicon market
Google, of course, is also working with Broadcom, and that’s unlikely to change. Again, the new chips aren’t actual Google TPUs, which is what Google and Broadcom have been working on together.
Memory interfaces, storage controllers, and NICs have long been commodity components — necessary, interchangeable, and largely invisible. This agreement treats them as strategic, deeply customized tools for AI performance, worthy of multiyear co-design and billions in committed procurement. That is becoming more of a trend amongst hyperscalers, who are increasingly relying less on generic designs.
Nvidia should pay attention too, even if nothing here displaces its GPUs outright. Purpose-built inference accelerators reduce Google’s dependence on merchant silicon for specific internal workloads, and every workload that migrates to custom chips is GPU growth that doesn’t happen. Amazon has Trainium, Microsoft has Maia, and Google now has both its own TPUs and a deepening bench of external co-design partners. The vertically integrated AI stack is becoming the default posture for hyperscalers, and that puts the burden on Nvidia to lean harder on the things custom silicon can’t easily match — the CUDA software ecosystem and its networking advantages through the Mellanox lineage.
None of this resolves overnight, and the warrant’s fiscal 2033 horizon says as much. But the direction is hard to miss. The chips that define AI infrastructure over the next several years won’t only be the headline accelerators — they’ll be the controllers and interfaces around them, designed for one customer, one workload, and one data center at a time.