The NVIDIA A100X is an enterprise-class GPU built on the Ampere architecture, featuring 80 GB of memory with 1,935 GB/s of memory bandwidth and a 300 W TDP. It delivers 19.5 TFLOPS of FP32 performance and 312 TFLOPS of FP16 tensor performance.
Hardware specifications
Same across all providers
No price history available
Price tracking data will appear here once available
0 offerings from 0 providers
Sorted by price ascending. Compare configurations side-by-side.
| Provider | Count | vCPU | RAM | Region | Per GPU hour | Total/hr | Action |
|---|---|---|---|---|---|---|---|
| Try broadening your filters or clear all filters to start over. | |||||||
LLMs that fit on the NVIDIA A100X
Estimated VRAM for popular open-weight LLMs at 16-bit (FP16/BF16), 8-bit (FP8/INT8) and 4-bit (INT4/MXFP4) precision, against 80 GB per GPU. Showing the 12 largest that fit.
| Model | Parameters | 16-bit | 8-bit | 4-bit |
|---|---|---|---|---|
| Laguna-S 2.1 Poolside | 118B 8B active | 282 GB Too large | 155 GB Too large | 90 GB Tight on 1× · 6% spare |
| gpt-oss-120b OpenAI | 117B 5.1B active | — Not released | — Not released | 78 GB 1× · 1.7 GB free |
| Qwen3.6 35B-A3B Alibaba | 35B 3B active | 86 GB Tight on 1× · 11% spare | 47 GB 1× · 33 GB free | 28 GB 1× · 52 GB free |
| Qwen3 32B Alibaba | 32.8B | 79 GB 1× · 1.4 GB free | 43 GB 1× · 37 GB free | 25 GB 1× · 55 GB free |
| Gemma 4 31B Google | 31.3B | 75 GB 1× · 4.9 GB free | 41 GB 1× · 39 GB free | 24 GB 1× · 56 GB free |
| GLM-4.7 Flash Z.ai | 31.2B | 75 GB 1× · 5.1 GB free | 41 GB 1× · 39 GB free | 24 GB 1× · 56 GB free |
| Qwen3.6 27B Alibaba | 27.8B | 67 GB 1× · 13 GB free | 37 GB 1× · 43 GB free | 21 GB 1× · 59 GB free |
| Mistral Small 3.2 24B Mistral AI | 24B | 58 GB 1× · 22 GB free | 32 GB 1× · 48 GB free | 18 GB 1× · 62 GB free |
| gpt-oss-20b OpenAI | 21B 3.6B active | — Not released | — Not released | 17 GB 1× · 63 GB free |
| Gemma 3 12B Google | 12.2B | 29 GB 1× · 51 GB free | 16 GB 1× · 64 GB free | 9.4 GB 1× · 71 GB free |
| Granite 4.1 8B IBM | 8.8B | 21 GB 1× · 59 GB free | 12 GB 1× · 68 GB free | 6.8 GB 1× · 73 GB free |
| Qwen3 8B Alibaba | 8.2B | 20 GB 1× · 60 GB free | 11 GB 1× · 69 GB free | 6.3 GB 1× · 74 GB free |
Estimates: model weights plus 20% for KV cache and runtime overhead, at short context lengths. Long contexts and large batches need more. Rows marked “tight” hold the weights but not that full margin. A ×N figure is the total VRAM across N of these GPUs and makes no claim about interconnect throughput. How we estimate this · Start from a model instead
Want more from this page?
See incorrect data?