Compute Exchange
Hardware market/Used GPUs/USED A100 40GB

USED A100 40GB

NVIDIA A100 40GB SXM4

Used A100 40GB units are increasingly attractive for cost-conscious teams. While memory-limited for the largest models, they deliver excellent throughput for inference, fine-tuning, and dev or test environments at the lowest pricing point of any 312 FP16 TFLOPS data center GPU.

USED A100 40GB — Indicative Range (Q3 2026)
$4,000 – $6,000

Based on aggregated supplier quotes and broker market data. Final pricing depends on quantity, region, condition, and warranty terms. Compute Exchange does not publish active counterparty listings.

Based on aggregated supplier quotes and broker market data. Final pricing depends on quantity, region, condition, and warranty terms. Compute Exchange does not publish active counterparty listings.

TECHNICAL SPECIFICATIONS
USED
VRAM
40 GB HBM2e
MEMORY BANDWIDTH
1.56 TB/s
FP 16 TENSOR
312 TFLOPS (sparse)
FP 8 TENSOR
Not supported
TDP
400W
FORM FACTOR
SXM4
INTERCONNECT
NVLink 3.0 / NVSwitch (600 GB/s)
ARCHITECTURE
Ampere
Frequently Asked Questions

USED A100 40GB — KEY QUESTIONS

What is the price range for used A100 40GB?+
Indicative pricing for used A100 40GB SXM4 is $4,000 to $6,000 per unit as of Q3 2026, around half of the $11,000 original MSRP. The 40GB variant trades at a steeper discount than the 80GB variant because memory is the binding constraint for many modern inference and fine-tuning workloads.
When does 40GB of memory become limiting?+
Inference for models above roughly 30B parameters typically overflows 40GB once activation memory and KV-cache are accounted for. Fine-tuning with full optimizer state for 13B-plus models also requires more than 40GB without aggressive sharding. For workloads that fit comfortably under those thresholds, A100 40GB delivers identical compute throughput to the 80GB variant.
What workloads suit a budget used A100 40GB cluster?+
Inference for 7B to 30B parameter models, parameter-efficient fine-tuning with LoRA or QLoRA, classical computer vision and tabular ML training, dev and test environments mirroring Ampere production, and academic research workloads. Stay under the 40GB ceiling and the price-per-throughput economics are unmatched.
Used A100 40GB versus used L40S — which delivers more for inference?+
Used L40S at $5,500 to $7,500 delivers 733 FP16 sparse TFLOPS plus FP8 acceleration with 48GB GDDR6X. Used A100 40GB at $4,000 to $6,000 delivers 312 FP16 sparse TFLOPS without FP8. For FP8-quantized LLM serving, L40S wins on tokens-per-second-per-dollar. For Ampere-style production workloads still on FP16/BF16, A100 40GB is the cheaper option.
Ready to Source or List?

LIVE QUOTE FOR
USED A100 40GB

Compute Exchange returns indicative pricing within 24 hours, anchored to your specific quantity, region, and condition. We do not publish active counterparty listings.

Request a Quote