What runs this LLM?

Pick the model you want to serve. We size it from the released checkpoint and list the cheapest configurations you can rent today that hold it β€” with the on-demand and serverless rate for each.

1 Β· Choose a model

Model
Precision
Context

Context-length sizing isn’t published for this architecture.

2 Β· What it needs

Qwen3.6 27B at 16-bit

67 GB of VRAM

at short context β€” measured from the released checkpoint, not estimated from parameter count.

Model weights
56 GB
KV cachefolded into overhead
β€”
Runtime overhead20% of weights
11 GB

About this model

Qwen3.6 27BAlibaba
Parameters
27.8B
Released at
16-bit
Measured weights
56 GB
Source checkpointQwen/Qwen3.6-27B

3 Β· Where to run it

Bookable configurations

Cheapest first, checked against 98 GPU types with live bookable capacity.

ConfigurationTotal VRAMOn-demandServerlessGPU details
8Γ— NVIDIA Tesla V100 16GB
128 GB
61 GB free
$0.24/hr
$0.03/GPU-hr
Vast.ai logoVast.ai
β€”See all prices
6Γ— NVIDIA RTX 3060
72 GB
5.3 GB free
$0.30/hr
$0.05/GPU-hr
Vast.ai logoVast.ai
β€”See all prices
7Γ— NVIDIA GTX 1070Tight fit
56 GB
11 GB short of the margin
$0.35/hr
$0.05/GPU-hr
Vast.ai logoVast.ai
β€”See all prices
4Γ— NVIDIA Tesla P100Tight fit
64 GB
2.7 GB short of the margin
$0.36/hr
$0.09/GPU-hr
Vast.ai logoVast.ai
β€”See all prices
7Γ— NVIDIA Titan Xp
84 GB
17 GB free
$0.42/hr
$0.06/GPU-hr
Vast.ai logoVast.ai
β€”See all prices
7Γ— NVIDIA GTX 1080 Ti
77 GB
10 GB free
$0.42/hr
$0.06/GPU-hr
Vast.ai logoVast.ai
β€”See all prices
6Γ— NVIDIA RTX A4000
96 GB
29 GB free
$0.42/hr
$0.07/GPU-hr
Vast.ai logoVast.ai
β€”See all prices
1Γ— NVIDIA A100 80GB
80 GB
13 GB free
$0.43/hr
Vast.ai logoVast.ai
$1.97/hr
Verda logoVerda
See all prices

Showing the 8 cheapest. Model weights plus 20% for runtime overhead, from measured checkpoint sizes; a stated context length adds its KV cache on top. Each price is the cheapest bookable rate for that exact instance size, for the whole configuration β€” the cheaper of the two is highlighted and orders the table. A Γ—N figure is total VRAM across N cards and makes no claim about interconnect throughput.

How this is worked out

Weight sizes are measured from the actual checkpoint files on Hugging Face β€” never derived from parameter count, which gets quantized releases wrong. On top of the weights we assume 20% for activations, CUDA context and allocator slack. Stating a context length adds its KV cache explicitly, computed only for architectures whose cache shape the config states unambiguously.

A configuration that holds the weights and cache but not that full margin is marked as a tight fit rather than hidden β€” it will load, and may well run at batch size one, but expect trouble under load.

Prices are the cheapest bookable rate for that exact instance size, refreshed with the rest of the catalog.

Keep looking