Back to tokcalc

H100 vs H200 for LLM Inference: Throughput, Latency, VRAM & Cost

Side-by-side comparison of NVIDIA H100 SXM5 80GB vs H200 SXM5 141GB for serving Llama 3.3 70B (FP8). See exactly how context length, VRAM, and bandwidth affect your deployment capacity — with transparent formulas, not marketing claims.

Specification comparison
SpecH100 SXM5H200 SXM5Winner
VRAM80 GB HBM3141 GB HBM3eH200
Memory bandwidth3,350 GB/s4,800 GB/sH200
FP16/BF16 dense TFLOPS990 TF990 TFTie
FP8 supportNativeNativeTie
NVLink bandwidth900 GB/s900 GB/sTie
Typical cloud price~$2.50/hr~$4.00/hrH100
Release year20222024H200
TDP (power)700W700WTie
Live comparison — adjust context length & GPU count
All numbers computed by tokcalc's open-source formulas. Click a config to open the full calculator.
H100 SXM580GB
Decode tok/s0.000
Aggregate tok/s0.000
TTFT7.13 s
VRAM needed201.67 GB / 160GB
Cost/M tokens—
Fits?✗ No — over budget
Open in calculator →
H200 SXM5141GB
Decode tok/s112.7
Aggregate tok/s901.5
TTFT7.13 s
VRAM needed201.67 GB / 282GB
Cost/M tokens$2.46
Fits?✓ Yes
Open in calculator →
Throughput by context length
Llama 3.3 70B FP8, 2× GPU, batch 8, continuous batching 2.0×

Key findings

1.

H200 doubles H100 concurrency at every context length due to 76% more VRAM (141 vs 80 GB). The free VRAM after loading model weights is the bottleneck — H200 has 70 GB free vs H100's 9 GB for Llama 70B FP8.

2.

H100 can't serve 32K+ context for Llama 70B FP8 on a single GPU — the KV cache alone needs 10+ GB, exceeding the 9 GB free after loading 70.6 GB of weights. H200 handles 32K with room for 6 concurrent users.

3.

H200 is 43% faster per-token (4,800 vs 3,350 GB/s HBM bandwidth) — decode throughput scales linearly with memory bandwidth for memory-bound LLM inference.

4.

H100 is 60% cheaper per GPU-hour (~$2.50 vs ~$4.00) but the cost-per-token is often LOWER on H200 because the higher throughput more than compensates for the higher hourly price.

FAQ

Should I use H100 or H200 for Llama 3.3 70B?

For context lengths under 8K, H100 is sufficient and cheaper. For 8K-32K, H200 is strongly preferred (H100 can't fit 32K on a single GPU). For 128K+, H200 is the minimum — H100 can't fit it even with TP×2.

How much faster is H200 than H100 for LLM inference?

H200 is ~43% faster per-token decode throughput (4,800 vs 3,350 GB/s HBM bandwidth). For aggregate batched throughput, H200 is 2-8× better than H100 at long context because it has 7× more free VRAM for KV cache.

Is H200 worth the extra cost over H100?

At 32K+ context, yes — H100 can't even fit the workload, so H200 is the only option. At short context (4K), H100 may be more cost-effective. Use the calculator above to find your break-even point.

Can H100 serve 128K context for Llama 70B?

No — a single H100 (80 GB) can't fit Llama 70B FP8 (70.6 GB weights) + 128K KV cache (~42 GB) = 112 GB total. You'd need at least 2× H100 (160 GB) or 1× H200 (141 GB).