Nvidia A40

Nvidia A40

Data center GPU for large-batch inference and professional visualization.

Last update 8 minutes ago
Compare vs other GPUs →
Aggregating historical prices...

Key Specifications

Architecture
Ampere
Memory
48GB GDDR6
Memory Bandwidth
696 GB/s
Release date
Q4 2020

A40 Pricing and Availability

For smaller projects, the A40 is available around $0.34/hr per GPU (on-demand). This tier offers a 83% discount compared to the higher end of the market ($2.05/hr). Spot instances start lower, at $0.32/hr per GPU.

6 providers

How this list works

The two views

By provider shows one card per company, placed where its best offer ranked. By configuration lists every offer.

Order

What you can rent comes first: in stock, then waitlist, then not reported, then out of stock. Priced offers rank ahead of quote-only, and on-demand ahead of other billing types.

Within a group, five factors set the order, largest first:

  1. Location: datacenter proximity, blended with provider HQ.
  2. Price: hourly price, per GPU and in total.
  3. Billing type: reserved ahead of spot, spot ahead of quote-only.
  4. Specs: more VRAM, vCPUs and RAM.
  5. Provider diversity: in By configuration, a provider's repeat rows rank slightly lower, so one company can't take all the top spots.

Search

We match the provider, GPU model, form factor, billing type, availability and instance name. Matching is partial and case-insensitive. Everyday words work too, so "interruptible" finds spot and "sold out" finds out of stock.

Transparency and funding

  • Ads and sponsors: Paid placements sit at the top of the list and are always labeled as sponsored content. Sponsorship never influences the organic ranking itself.
  • Affiliates: Affiliate links are marked. We may earn a commission if you click them, but commissions never affect the order.
  • Prices: Shown in USD, converted at daily reference rates where a provider publishes in another currency. A month means 720 hours. How we estimate costs.
Vast.ai logo

Vast.ai Our sponsor

USA 2 configs 1x

On-Demand from $0.34 Spot from $0.32
From $0.34 / GPU / hr On-Demand Visit website
Runpod logo

Runpod In stock

USA 5 configs 1x-8x

On-Demand from $0.44 Reserved on request
From $0.44 / GPU / hr On-Demand Visit website
GPU.ai logo

GPU.ai In stock

UAE 1 config 1x

On-Demand from $0.44
From $0.44 / GPU / hr On-Demand Visit website
Sesterce logo

Sesterce In stock

France 1 config 1x

On-Demand from $2.05
From $2.05 / GPU / hr On-Demand Visit website
Exoscale logo

Exoscale

Switzerland 4 configs 1x-8x

On-Demand from $1.05
From $1.05 / GPU / hr On-Demand Visit website
Database Mart logo

Database Mart

USA 4 configs 1x

Reserved from $0.41
From $0.41 / GPU / hr Reserved (24mo) Visit website

Heads up: We do our best to keep these specs & prices accurate. However, cloud costs may fluctuate based on region, usage, and other factors not listed here. These are estimates based on common setups and are for informational purposes only. Always verify current rates & exact specs with the provider before provisioning.

Frequently Asked Questions

Why choose the A40?

48GB GDDR6 with Ampere Tensor Cores. Large VRAM for inference on bigger models. Supports both AI and professional visualization workloads.

When is the A40 not a good fit?

No FP8 support. PCIe-only. For pure AI inference, the L40S offers better performance with Ada Lovelace architecture at similar or lower cost.

Are A40 prices going up or down?

The median on-demand price across providers has risen about 28% since August 2025, from $0.58 to $0.74/hr per GPU. This reflects the market-wide median, which moves both when providers change prices and when lower or higher priced offerings enter the market.

What size AI models can the A40 run?

With 48GB of VRAM, the A40 can typically run models up to about 30B parameters in FP16, or 70B-class models in 4-bit quantized form for inference.

How much VRAM does the A40 have?

The A40 has 48GB of VRAM. Multi-GPU setups increase total memory, but that memory is not automatically pooled across GPUs.

What is the A40's memory bandwidth?

The A40 has 696 GB/s of memory bandwidth. Higher bandwidth helps with faster data transfer between GPU memory and compute cores.

What data types does the A40 support?

The A40 supports 6 precision formats. Training: BF16, FP16, TF32, FP32. Inference: INT4, INT8.

Does the A40 support NVLink?

No. The A40 is a PCIe-only GPU with no NVLink, so it is better suited to single-GPU inference and smaller-scale workloads than large distributed training jobs.

How much does the A40 cost per hour?

A40 pricing currently ranges from $0.32/hr to $2.05/hr per GPU, depending on the provider, instance type, and billing model.

How much does the A40 cost per month?

At 720 hours per month, one A40 can cost between $231.98 to $1,473.12 per month, depending on the provider. Reserved and spot pricing can lower that further.

Which cloud providers offer the A40?

The A40 is available from 6 cloud providers, including Runpod, Exoscale, Sesterce. Pricing and availability vary by region and billing model.

Can I rent the A40 in the cloud?

Yes. We currently track 17 A40 listings across 6 cloud providers:

Billing type Listings Avg $/GPU/hr
On-demand 11 $0.80/hr
Reserved 5 $0.47/hr
Spot 1 $0.32/hr

Technical Specifications

GPU architecture NVIDIA Ampere architecture
GPU memory 48 GB GDDR6
Memory bandwidth 696 GB/s
Interconnect interface NVIDIA® NVLink® 112.5 GB/s (bidirectional), PCIe Gen4: 64GB/s
NVIDIA Ampere architecture-based CUDA Cores 10,752
NVIDIA second-generation RT Cores 84
NVIDIA third-generation Tensor Cores 336
Peak FP32 TFLOPS (non-Tensor) 37.4
Peak FP16 Tensor TFLOPS with FP16 Accumulate 149.7, 299.4*
Peak TF32 Tensor TFLOPS 74.8, 149.6*
RT Core performance TFLOPS 73.1
Peak BF16 Tensor TFLOPS with FP32 Accumulate 149.7, 299.4*
Peak INT8 Tensor TOPS, Peak INT4 Tensor TOPS 299.3, 598.6, 598.7, 1,197.4
Form factor 4.4” (H) x 10.5” (L) dual slot
Display ports 3x DisplayPort 1.4**, Supports NVIDIA Mosaic and Quadro® Sync
Max power consumption 300 W
Power connector 8-pin CPU
Thermal solution Passive
Virtual GPU (vGPU) software support NVIDIA vPC/vApps, NVIDIA RTX Virtual Workstation, NVIDIA Virtual Compute Server
vGPU profiles supported See the Virtual GPU Licensing Guide
NVENC, NVDEC 1x, 2x (includes AV1 decode)
Secure and measured boot with hardware root of trust Yes (optional)
NEBS ready Level 3
Compute APIs CUDA, DirectCompute, OpenCL™, OpenACC®
Graphics APIs DirectX 12.07, Shader Model 5.17, OpenGL 4.68, Vulkan 1.18
MIG support No

Source: official Nvidia A40 datasheet.

Alternatives to Nvidia A40

Last updated