Plutux Logo
Plutux
Actualizando traducción
Google Develops 'Frozen v2' Chip With Gemini Baked Into Silicon — A 6-10x Efficiency Play for 2028 insight cover
Industry NewsGOOGL11 min de lectura

Google Develops 'Frozen v2' Chip With Gemini Baked Into Silicon — A 6-10x Efficiency Play for 2028

Reuters reported on July 20, 2026 that Google is developing a new server chip codenamed 'Frozen v2' that embeds elements of its Gemini model directly into the hardware. The chip is projected to be 6–10x more efficient than current custom Google AI silicon (measured by tokens served per watt) and is targeted for deployment as early as 2028. The 'Frozen' program runs alongside but does not replace Google's existing TPU roadmap (TPU 8t/8i announced at Cloud Next '26) and signals an architectural shift toward model-silicon co-design, putting further pressure on the GPU-centric AI compute stack.

Publicado 21 jul 2026Actualizado 21 jul 2026

Event

Frozen v2

Server chip project informally dubbed “Frozen v2” integrating Gemini elements into hardware (Reuters, 2026-07-20)

Efficiency target

6–10x

More tokens served per unit of power vs. Google’s latest custom AI chips (Reuters, 2026-07-20)

Deployment window

2028

Google expects to deploy as soon as 2028 (Reuters, 2026-07-20)

Google’s reported “Frozen v2” chip is a signal that hyperscalers are moving from “accelerate the model” to “bake the model into the silicon.” If the rumored 6–10x token-per-watt jump holds and the 2028 timing is right, the inference cost curve (and therefore capex + power demand planning) shifts again—this time targeting Gemini specifically rather than general-purpose AI compute.

Event

Frozen v2

Server chip project informally dubbed “Frozen v2” integrating Gemini elements into hardware (Reuters, 2026-07-20)

Efficiency target

6–10x

More tokens served per unit of power vs. Google’s latest custom AI chips (Reuters, 2026-07-20)

Deployment window

2028

Google expects to deploy as soon as 2028 (Reuters, 2026-07-20)

What happened

Frozen v2 reframes Google’s roadmap from “TPU upgrades” to “Gemini co-designed with hardware”

Reuters reports that Google is developing a new server chip codenamed “Frozen v2,” designed to incorporate elements of its Gemini model directly into the hardware. The key headline isn’t just a better accelerator—it’s model-specific silicon specialization aimed at inference efficiency.

Load-bearing Reuters details (the claims the market will trade on)

Codename / concept

Frozen v2

A new server chip integrating Gemini elements into silicon

Metric

Tokens per watt

Efficiency measured as tokens served per unit of power

Expected gain

6–10x

Against Google’s latest custom AI chips

Timing

As early as 2028

Targeted deployment window

Relationship to TPU

Separate track

Reuters characterizes it as a new set of homegrown chips rather than replacing the TPU roadmap

This analysis depends on Reuters’ 6–10x and 2028 claims; they are not independently verifiable from primary Google technical specs yet. Treat the magnitude as a high-volatility input to valuation and supply-chain conclusions.

Baseline & comparison

TPU 8t/8i raises the bar—but Frozen v2 is about a different lever: tokens per watt via deeper model embedding

Google’s own Next ’26 messaging distinguishes TPU 8t (training) and TPU 8i (reasoning/inference). TPU 8i is already positioned as an inference-specialized system (including MoE-oriented design choices), but Frozen v2—if it exists as reported—targets further efficiency by pushing Gemini-specific structure into hardware.

What Google publicly says TPU 8i is optimized for (inference/agentic workloads)
SystemStated purposePublicly cited inference-related details
TPU 8iInference and reinforcement learning / reasoning384 MB on-chip SRAM; 288 GB HBM3; 19.2 Tb/s ICI bandwidth; CAE reduces on-chip latency by up to 5x; “up to 80% better performance per dollar” vs prior generation; “soon” availability to Cloud customers
TPU 8tTraining powerhouseNearly 3x higher compute performance vs previous generations; 9,600 chips per superpod; 121 exaflops compute per superpod; 2 petabytes shared memory
  • Google is already productizing inference specialization with TPU 8i (performance-per-dollar improvements and low-latency engines are explicitly marketed).
  • Frozen v2, per Reuters, is a further step: the efficiency lever shifts from system-level optimization to potentially architecture-level model embedding (Gemini elements baked into silicon).
  • If token-per-watt truly jumps 6–10x, the economics can swing even if TPU 8i is “only” meaningfully better than Ironwood; improvements compound over rising inference demand.

Causal chain

The 6–10x claim likely implies Frozen v2 tackles inference bottlenecks beyond raw compute

A jump like 6–10x in tokens per watt usually can’t come from “more FLOPS” alone. It implies a co-design strategy that reduces wasted compute, cuts data movement, and matches silicon execution units to the dominant shapes in Gemini inference.

Where a “tokens per watt” uplift can come from (hypothesis map tied to what Google is already signaling with TPU 8i)

Not a measurement—this frames which efficiency levers are consistent with Google’s public TPU 8i design emphasis and Reuters’ Frozen v2 premise.

Unidad: relative importance (heuristic)

Lower on-chip latency / fewer cycles per token

Consistent with TPU 8i marketing for latency reduction via CAE

60

Reduced memory traffic per token (HBM/SRAM strategy)

Consistent with TPU 8i SRAM + bandwidth emphasis

50

Better fit for MoE/routing patterns

TPU 8i marketed for agentic workflows and MoE models; co-design could specialize routing

45

Hardwiring Gemini architecture elements

Directly aligned with Reuters: embed elements of Gemini into hardware

40

This is inference: the article hasn’t disclosed microarchitecture. But the causal story matches Google’s prior direction (TPU 8i is inference-specialized; Frozen v2 is reported as model-embedded).

Impact

Frozen v2 pressures the GPU-centric stack most through inference TCO planning, not training headline benchmarks

Even if GPUs remain strong for broad training and fast iteration, an inference-specific, token-per-watt breakthrough hits the part of the P&L that matters day-to-day: serving cost at scale. That tends to shift buying from “general accelerators” toward bespoke inference silicon and the systems that wrap it.

Multi-dimensional impact map (direct + upstream + downstream)
LayerLikely effect of Frozen v2 (if targets are real)Who wins / who loses (listed comparables)
Direct (Google/Alphabet)Lower inference cost per token; potential to expand Gemini usage without proportional power/capacity increasesAlphabet benefits; margin and cash flow sensitivity via capex + opex efficiency
Downstream channelCloud inference pricing pressure + improved unit economics for agentic workflowsCloud customers get lower effective cost; competition for merchant-GPU-priced inference intensifies for alternatives to Google’s stack
Upstream: custom silicon co-design & packagingMore reliance on ASIC/XPU co-design ecosystems and advanced packaging supply chainsBroadcom (custom AI chip co-design / infrastructure stack) likely benefits; others depend on whether they supply packaging/interconnect components
Upstream: mobile/edge silicon demand spilloversModel-embedded design can create trickle-down expectations for inference efficiency across device classesUncertain for MediaTek without direct verification of partnership scope in this specific event; treat as a scenario, not a fact
Competitive compute stack (GPU-centric)If token-per-watt rises 6–10x, GPUs face tougher economics in the “inference at hyperscaler scale” segmentNVIDIA at risk on inference share—though training and software moat may cushion

Supply-chain view

The most important supply-chain variable is advanced co-design + packaging, not just the compute die

Google’s reported direction (model embedded in hardware) increases the probability that the differentiator sits across the stack: interconnect topology, high-bandwidth memory strategy, and high-yield packaging capable of shipping at hyperscaler volumes.

  • Advanced packaging and packaging-enabled memory bandwidth are already central in Google’s TPU story (TPU 8i publicly emphasizes large HBM capacity and very high ICI bandwidth).
  • If Frozen v2 is “separate from TPU” rather than “a smaller TPU upgrade,” it likely expands the set of platform-level constraints (new board/topology, new memory/per-token cycle budget), increasing spend with custom-XPU ecosystems.
  • This environment tends to be favorable to companies that sit in the middle of custom accelerator co-design and the infrastructure needed to integrate dies into systems; Broadcom is the most direct listed beneficiary among your topic set.
We did not obtain a primary Google technical document specifying Frozen v2’s packaging or supplier list in this session; supply-chain linkage is therefore scenario-based on the custom-ASIC ecosystem context, not event-specific BOM proof.

Fundamental lens (listed companies)

Alphabet is positioned to self-fund model-silicon bets; Broadcom is levered to custom-XPU ramps; NVIDIA’s exposure depends on inference-share durability

Use the financials to map capacity to invest and where the market may re-rate exposure. Frozen v2 is a tech story, but hyperscaler silicon moves are capital- and supply-chain-intensive—so balance-sheet and operating cash flow matter.

Core financial scale check (latest figures from the data tools used in this session)
CompanyFY revenueFY operating cash flowFY free cash flowWhat matters for Frozen v2 read-through
Alphabet$402.963B (FY2025)$164.713B (FY2025)$73.266B (FY2025)Can absorb multi-year infrastructure R&D and capex; Frozen v2 aims at inference economics that scale with Gemini usage
Broadcom$63.887B (FY2025)$27.537B (FY2025)$26.914B (FY2025)Custom silicon ecosystem leverage; if hyperscalers increase custom inference ASIC share, Broadcom’s co-design + infrastructure positioning is structurally advantaged
NVIDIA$130.497B (FY2025)$64.089B (FY2025)$60.853B (FY2025)More exposed to merchant GPU inference economics; however, NVIDIA’s training breadth and software moat can cushion—effects hinge on whether inference share shifts away

Alphabet vs Broadcom vs NVIDIA: operating cash flow as a funding indicator for long-cycle hardware bets

FY operating cash flow from the data tools used in this session.

Unidad: USD

Alphabet

Operating cash flow FY2025

164,713,000,000

Broadcom

Operating cash flow FY2025

27,537,000,000

NVIDIA

Operating cash flow FY2025

64,089,000,000

The “self-funding advantage” is real at the hyperscaler level: Alphabet’s FY2025 operating cash flow of about $164.7B provides room to pursue model-specific chips like Frozen v2 without relying on external accelerator procurement economics.

Valuation/pricing risk to the GPU stack

If Frozen v2 delivers, GPUs face a reset in “inference per watt” competitive comparisons—especially for MoE-style workloads

The market tends to price inference silicon on throughput-per-watt and throughputs-per-dollar. Google already market-tested this with TPU 8i (up to 80% better performance per dollar than Ironwood). Frozen v2’s reported 6–10x token-per-watt target implies a potential step-change in the next round of competitive comparisons.

Why this matters specifically for inference economics (not just total compute capacity)
MechanismWhy it changes TCOMost likely arena impacted
Higher tokens per wattLower power cost per token; reduces required facility capacity for a given demandHigh-volume inference (Gemini serving, agentic workflows, MoE routing)
Lower cycles/token via model embeddingFewer compute cycles per generated token; lowers both power and accelerator-hours demandLatency-sensitive and high-throughput token generation
System-level specialization (as with TPU 8i)Improves memory bandwidth and reduces on-chip latency; raises effective utilizationClusters optimized for the model’s inference shape
The risk to NVIDIA is not that GPUs stop working—it’s that the “infer at minimum cost” portion of demand shifts toward more bespoke model-silicon from hyperscalers.

What to watch next (1–3 years)

The biggest milestone isn’t Frozen v2’s name—it’s any public signal that model-embedded chips are shipping into production systems

  • Google Cloud availability timing: any “soon”/GA-like language for Frozen v2-class inference beyond internal pilots (Reuters says 2028 target).
  • Comparative inference metrics: whether Google publishes token-per-watt, latency, or cost-per-token for Gemini-serving scenarios against existing TPU generations.
  • Supply-chain triangulation: any additional reporting or supplier mentions around packaging/interconnect and custom co-design partners as Frozen v2 transitions from prototype to ramp.
  • GPU allocation signals: hyperscaler capex mix commentary (e.g., whether NVIDIA orders are affected at the inference rack/system level).

Synthesis / stance

Frozen v2 is a “model-silicon co-design accelerant”—bullish for custom-ASIC ecosystems, bearish for merchant inference dominance

My base-case interpretation is that Reuters’ 6–10x token-per-watt target, if realized, pushes the industry further toward hardware that is specialized to the model family and its inference shapes. That tends to increase the value of custom co-design and advanced system integration, while forcing GPU-centric inference economics to defend their position at scale.

Thesis mapping: what must be true for each listed company exposure
CompanyPrimary “if true” driverKey variable to verify
AlphabetFrozen v2 reduces Gemini serving cost enough to expand usage and protect margins despite rising demandWhether the token-per-watt uplift shows up in measurable production economics by/after 2028
BroadcomHyperscalers increase custom inference ASIC share and keep expanding co-design/integration needsAny event-specific supplier role expansion (not yet proven here) tied to Google/others’ model-embedded chips
NVIDIAGPU inference share declines at the margin, but training breadth and software ecosystem slow the impactWhether NVIDIA’s customer mix shifts away from inference-dominant deployments as token-per-watt becomes the dominant buyer metric
© Plutux Technology Limited 2026