Compute Exchange
Reserved GPU Rental/RESERVED H100 NVL

RESERVED H100 NVL

NVIDIA H100 NVL 94GB

The H100 NVL brings 94GB of HBM3 and SXM-class throughput to PCIe servers — built for LLM inference, with paired-card NVLink bridging for models that spill past a single GPU.

AVAILABLE TERM LENGTH
1MO3MO6MO12MO24MO36MO

All term lengths available. Deploys in standard PCIe hosts, so provisioning is typically faster than SXM parts.

TECHNICAL SPECIFICATIONS
RESERVED
VRAM
94 GB HBM3
MEMORY BANDWIDTH
3.9 TB/s
FP 16 TENSOR
1,979 TFLOPS (sparse)
FP 8 TENSOR
3,958 TFLOPS (sparse)
TDP
400W
FORM FACTOR
PCIe Gen5 (dual-slot)
INTERCONNECT
NVLink bridge (600 GB/s) / PCIe 5.0
ARCHITECTURE
Hopper
Partner Network

AGGREGATED ACROSS
LEADING NEOCLOUDS

Compute Exchange aggregates reserved capacity from a verified network of leading AI-native cloud providers and hyperscalers. All partners undergo identity, capacity, SLA, and operational verification before quotes surface on the network.

You receive a normalized comparison across providers in a single quote response — rather than evaluating each neocloud's contract structure, billing model, and SLA terms in isolation.

WORKLOAD FIT

RESERVED H100 NVL
USE CASES

LLM inference serving
KV-cache-heavy workloads
Paired-GPU model sharding
PCIe cluster upgrades
WHY RESERVE

RESERVED H100 NVL
VS ON-DEMAND

H100 NVL sits in the sweet spot for inference fleets: more memory than the SXM5 H100 with simple PCIe deployment. Reserved terms lock pricing that undercuts on-demand meaningfully at steady utilization.

Frequently Asked Questions

RESERVED H100 NVL
KEY QUESTIONS

What makes H100 NVL different from H100 PCIe and SXM5?+
H100 NVL is the inference-tuned PCIe variant: 94GB of HBM3 (versus 80GB), 3.9TB/s of bandwidth approaching SXM5 levels, and NVLink bridging between card pairs. It deploys in standard PCIe Gen5 hosts — no SXM baseboard required — which widens the provider pool and shortens provisioning.
What reservation terms are available for H100 NVL?+
The full ladder, 1 through 36 months. Supply is healthy across mid-tier and AI-native providers, so short pilots and long production commits both quote quickly. The discount curve steepens meaningfully at 12 months and beyond.
When should I choose H100 NVL over H200 for inference?+
H100 NVL wins on price-per-token for models that fit comfortably in 94GB or shard cleanly across bridged pairs. Step up to H200 when context lengths push KV-cache past what 94GB sustains at your target batch size — the 141GB of HBM3e is the deciding factor, not raw compute.
Ready to Reserve?

LIVE QUOTE FOR
RESERVED H100 NVL

Compute Exchange returns indicative pricing within 24 hours, anchored to your specific quantity, region, and condition. We do not publish active counterparty listings.

Request a Quote