What runs this LLM?
Pick the model you want to serve. We size it from the released checkpoint and list the cheapest configurations you can rent today that hold it β with the on-demand and serverless rate for each.
1 Β· Choose a model
Context-length sizing isnβt published for this architecture.
2 Β· What it needs
Solar Open2 250B at 16-bit
601 GB of VRAM
at short context β measured from the released checkpoint, not estimated from parameter count.
- Model weights
- 501 GB
- KV cachefolded into overhead
- β
- Runtime overhead20% of weights
- 100 GB
About this model
- Parameters
- 250B
- Active per token
- 15B
- Released at
- 16-bit
- Measured weights
- 501 GB
3 Β· Where to run it
Bookable configurations
Cheapest first, checked against 98 GPU types with live bookable capacity.
| Configuration | Total VRAM | On-demand | Serverless | GPU details |
|---|---|---|---|---|
| 8Γ NVIDIA A100 80GB | 640 GB 39 GB free | $5.68/hr $0.71/GPU-hr | $15.75/hr $1.97/GPU-hr | See all prices |
| 8Γ NVIDIA RTX PRO 6000 | 768 GB 167 GB free | $8.00/hr $1.00/GPU-hr Lium | $16.63/hr $2.08/GPU-hr | See all prices |
| 8Γ Intel Gaudi 2 | 768 GB 167 GB free | $8.31/hr $1.04/GPU-hr | β | See all prices |
| 4Γ AMD Instinct MI300X | 768 GB 167 GB free | $11.36/hr $2.84/GPU-hr | β | See all prices |
| 4Γ AMD Instinct MI325X | 1,024 GB 423 GB free | $12.35/hr $3.09/GPU-hr | β | See all prices |
| 16Γ NVIDIA RTX 5090Tight fit | 512 GB 89 GB short of the margin | $14.08/hr $0.88/GPU-hr | β | See all prices |
| 8Γ NVIDIA H100 | 640 GB 39 GB free | $14.96/hr $1.87/GPU-hr | $28.60/hr $3.58/GPU-hr | See all prices |
| 4Γ NVIDIA B200 | 720 GB 119 GB free | $23.52/hr $5.88/GPU-hr | $26.88/hr $6.72/GPU-hr | See all prices |
Showing the 8 cheapest. Model weights plus 20% for runtime overhead, from measured checkpoint sizes; a stated context length adds its KV cache on top. Each price is the cheapest bookable rate for that exact instance size, for the whole configuration β the cheaper of the two is highlighted and orders the table. A ΓN figure is total VRAM across N cards and makes no claim about interconnect throughput.
How this is worked out
Weight sizes are measured from the actual checkpoint files on Hugging Face β never derived from parameter count, which gets quantized releases wrong. On top of the weights we assume 20% for activations, CUDA context and allocator slack. Stating a context length adds its KV cache explicitly, computed only for architectures whose cache shape the config states unambiguously.
A configuration that holds the weights and cache but not that full margin is marked as a tight fit rather than hidden β it will load, and may well run at batch size one, but expect trouble under load.
Prices are the cheapest bookable rate for that exact instance size, refreshed with the rest of the catalog.
