TPUs are specialized ASICs for tensor processing in machine-learning systems . That specialization is the reason they can be attractive for large, regular tensor workloads: when the compiler path, tensor shapes, batching, and sharding are TPU-friendly, more of the silicon can be kept busy.
H100 takes a broader route. It is heavily optimized for AI through Tensor Cores, but NVIDIA's H100 SXM spec table also includes conventional FP64 and FP32 performance as well as multiple lower-precision Tensor Core modes . That breadth matters when the same accelerator pool must support experiments with different precision requirements or workloads that are not all identical deep-learning jobs.
Raw specifications show the shape of the trade-off, but they are not an apples-to-apples benchmark. TPU and GPU tables often report different precision modes, different system assumptions, and different scaling paths.
| Accelerator | Public memory figure | Public bandwidth figure | Public compute figures | Best read as |
|---|---|---|---|---|
| TPU v5e | 16GB HBM per chip | 8.1e11 bytes/s per chip | 1.97e14 BF16 FLOPs/s per chip; 3.94e14 INT8 FLOPs/s per chip | A TPU option with less per-chip HBM than v5p or v6e in the JAX table; check memory fit carefully . |
| TPU v5p | 96GB HBM per chip | 2.8e12 bytes/s per chip | 4.59e14 BF16 FLOPs/s per chip; 9.18e14 INT8 FLOPs/s per chip | The highest HBM-per-chip TPU row among v5e, v5p, and v6e in the JAX table . |
| TPU v6e | 32GB HBM per chip | 1.6e12 bytes/s per chip | 9.20e14 BF16 FLOPs/s per chip; 1.84e15 INT8 FLOPs/s per chip | The highest listed BF16 and INT8 per-chip throughput among these TPU rows . |
| NVIDIA H100 SXM | 80GB HBM3 | 3.35TB/s | 67 TFLOPS FP32; 989 TFLOPS TF32 Tensor Core; 1,979 TFLOPS BF16/FP16 Tensor Core; 3,958 TFLOPS FP8 Tensor Core; 3,958 TOPS INT8 Tensor Core | Broad precision coverage, high memory bandwidth, and a more general accelerator profile . |
Google Cloud also documents H100-backed A3 machine types with 1, 2, 4, or 8 attached H100 GPUs and 80GB HBM3 per GPU . Google Cloud's AI Hypercomputer material also describes TPUs and A3 VMs running H100 GPUs as part of the same AI infrastructure portfolio . In practice, the choice is not always TPU on Google Cloud versus GPU somewhere else.
A TPU is the stronger candidate when specialization is an advantage rather than a constraint. Put it high on the shortlist if:
TPUs can be compelling when the workload keeps the chips busy and avoids costly rewrites. But that is a workload result, not a universal property. Google has published performance-per-dollar material for GPUs and TPUs in AI inference, which reinforces that serving economics depend on the model and setup rather than a single universal accelerator ranking .
NVIDIA H100 is the stronger candidate when flexibility is worth more than specialization. It is especially attractive when:
The strongest argument for H100 is not always that one GPU beats one TPU chip in every benchmark. It is that the GPU is the more flexible accelerator when requirements change.
Pricing comparisons are tempting, but they can be fragile. One third-party comparison listed Google Cloud TPU v5e at about $1.20 per chip-hour and an Azure ND H100 v5 example at about $12.84 per 80GB H100 GPU-hour . That is cross-cloud and unofficial, so it should be treated as directional rather than a universal TPU-is-cheaper conclusion.
A better cost comparison measures the whole system:
The practical metric is cost per useful output: per training step, per converged model, per inference token, or per latency target.
| Priority | Better default | Why |
|---|---|---|
| TPU-friendly deep learning on Google Cloud | Google TPU | Public TPU docs emphasize pod scale, HBM, bandwidth, and BF16/INT8 throughput for model scaling . |
| Broad precision support | NVIDIA H100 GPU | H100 SXM lists FP64, FP32, TF32 Tensor Core, BF16/FP16 Tensor Core, FP8 Tensor Core, and INT8 Tensor Core modes . |
| Existing Google Cloud deployment with optionality | Benchmark both | Google Cloud documents A3 H100 machine types and also positions TPUs and H100 A3 VMs in its AI infrastructure portfolio . |
| Lowest inference cost | Benchmark both | Google has published performance-per-dollar analysis for AI inference, while third-party chip-hour examples are directional and cross-cloud . |
| Existing GPU-first production stack | NVIDIA H100 GPU | Avoiding migration risk can matter more than a theoretical accelerator-efficiency gain. |
Treat TPU as the more specialized AI accelerator and H100 as the more flexible accelerator platform. If your model is TPU-friendly, deep-learning-heavy, and already headed for Google Cloud, a TPU can be the better cost-performance bet. If you need broad numeric modes, mixed workloads, GPU-oriented operational continuity, or lower migration risk, NVIDIA H100 GPUs are usually the safer default .
The only reliable final answer is a workload-specific benchmark that measures throughput, memory behavior, utilization, total cost, and engineering effort on the exact model you plan to train or serve.