Which open-weight LLM should you run?

Compare 23 models by size, license and the VRAM they actually need, and see the cheapest cloud GPU that serves each one right now.

23 models98 GPU types with bookable capacity
Size
Showing 23 of 23 modelsTick up to 4 models to see them side by side.
CompareModelLicenseVRAM to serveCheapest to serveMax contextActions
Qwen3 8BAlibaba8.2Breleased at 16-bitNot declared
16-bit
20 GB
8-bit
11 GB
4-bit
6.3 GB
1× NVIDIA Tesla P4at 4-bit$0.02/hr on-demandVast.ai logoat Vast.ai
40KWhat runs it?
Gemma 4 31BGoogle31.3Breleased at 16-bitNot declared
16-bit
75 GB
8-bit
41 GB
4-bit
24 GB
2× NVIDIA Tesla V100 16GBat 4-bit$0.06/hr on-demandVast.ai logoat Vast.ai
256KWhat runs it?
Qwen3 4BAlibaba4Breleased at 16-bitNot declared
16-bit
9.7 GB
8-bit
5.3 GB
4-bit
3.1 GB
1× NVIDIA GTX 1050 Tiat 4-bit$0.01/hr on-demandVast.ai logoat Vast.ai
40KWhat runs it?
gpt-oss-20bOpenAI21B3.6B activereleased at 4-bitNot declared
16-bit
8-bit
4-bit
17 GB
128KWhat runs it?
Llama 3.1 8B InstructMeta8Breleased at 16-bitNot declared
16-bit
19 GB
8-bit
11 GB
4-bit
6.2 GB
1× NVIDIA Tesla P4at 4-bit$0.02/hr on-demandVast.ai logoat Vast.ai
What runs it?
Qwen3.6 27BAlibaba27.8Breleased at 16-bitNot declared
16-bit
67 GB
8-bit
37 GB
4-bit
21 GB
2× NVIDIA Tesla V100 16GBat 4-bit$0.06/hr on-demandVast.ai logoat Vast.ai
256KWhat runs it?
gpt-oss-120bOpenAI117B5.1B activereleased at 4-bitNot declared
16-bit
8-bit
4-bit
78 GB
128KWhat runs it?
Qwen3 32BAlibaba32.8Breleased at 16-bitNot declared
16-bit
79 GB
8-bit
43 GB
4-bit
25 GB
2× NVIDIA Tesla V100 16GBat 4-bit$0.06/hr on-demandVast.ai logoat Vast.ai
40KWhat runs it?
Qwen3.6 35B-A3BAlibaba35B3B activereleased at 16-bitNot declared
16-bit
86 GB
8-bit
47 GB
4-bit
28 GB
2× NVIDIA Tesla V100 16GBat 4-bit$0.06/hr on-demandVast.ai logoat Vast.ai
256KWhat runs it?
Mistral 7B Instruct v0.3Mistral AI7.2Breleased at 16-bitNot declared
16-bit
17 GB
8-bit
9.6 GB
4-bit
5.6 GB
1× NVIDIA Tesla P4at 8-bit$0.02/hr on-demandVast.ai logoat Vast.aiTight fit8 GB
32KWhat runs it?
Kimi K3Moonshot AI2.8T104B activereleased at 4-bitNot declared
16-bit
8-bit
4-bit
1,873 GB
1,024KWhat runs it?
GLM-4.7 FlashZ.ai31.2Breleased at 16-bitNot declared
16-bit
75 GB
8-bit
41 GB
4-bit
24 GB
2× NVIDIA Tesla V100 16GBat 4-bit$0.06/hr on-demandVast.ai logoat Vast.ai
198KWhat runs it?
DeepSeek V4 FlashDeepSeek284Breleased at 4-bitNot declared
16-bit
8-bit
4-bit
192 GB
1,024KWhat runs it?
GLM-5.2Z.ai753Breleased at 16-bitNot declared
16-bit
1,808 GB
8-bit
994 GB
4-bit
579 GB
8× NVIDIA A100 80GBat 4-bit$5.68/hr on-demandVast.ai logoat Vast.ai
1,024KWhat runs it?
DeepSeek R1DeepSeek671B37B activereleased at 8-bitNot declared
16-bit
8-bit
826 GB
4-bit
481 GB
8× NVIDIA A100 80GBat 4-bit$5.68/hr on-demandVast.ai logoat Vast.ai
160KWhat runs it?
Granite 4.1 8BIBM8.8Breleased at 16-bitNot declared
16-bit
21 GB
8-bit
12 GB
4-bit
6.8 GB
1× NVIDIA Tesla P4at 4-bit$0.02/hr on-demandVast.ai logoat Vast.ai
128KWhat runs it?
Gemma 3 12BGoogle12.2Breleased at 16-bitNot declared
16-bit
29 GB
8-bit
16 GB
4-bit
9.4 GB
1× NVIDIA Tesla P4at 4-bit$0.02/hr on-demandVast.ai logoat Vast.aiTight fit8 GB
What runs it?
DeepSeek V4 ProDeepSeek1.6T49B activereleased at 4-bitNot declared
16-bit
8-bit
4-bit
1,038 GB
1,024KWhat runs it?
Kimi K2.5Moonshot AI1T32B activereleased at 4-bitNot declared
16-bit
8-bit
4-bit
714 GB
8× NVIDIA A100 80GB$5.68/hr on-demandVast.ai logoat Vast.aiTight fit640 GB
256KWhat runs it?
Qwen3 235B-A22BAlibaba235B22B activereleased at 16-bitNot declared
16-bit
564 GB
8-bit
310 GB
4-bit
181 GB
8× NVIDIA RTX 3090at 4-bit$0.72/hr on-demandVast.ai logoat Vast.ai
40KWhat runs it?
Mistral Small 3.2 24BMistral AI24Breleased at 16-bitNot declared
16-bit
58 GB
8-bit
32 GB
4-bit
18 GB
2× NVIDIA Tesla V100 16GBat 8-bit$0.06/hr on-demandVast.ai logoat Vast.ai
128KWhat runs it?
Laguna-S 2.1Poolside118B8B activereleased at 16-bitNot declared
16-bit
282 GB
8-bit
155 GB
4-bit
90 GB
8× NVIDIA Tesla V100 16GBat 4-bit$0.24/hr on-demandVast.ai logoat Vast.ai
1,024KWhat runs it?
Solar Open2 250BUpstage250B15B activereleased at 16-bitNot declared
16-bit
601 GB
8-bit
330 GB
4-bit
192 GB
8× NVIDIA RTX 3090at 4-bit$0.72/hr on-demandVast.ai logoat Vast.aiTight fit192 GB
1,024KWhat runs it?

How this table is made

Every size is measured from the weight files of the released checkpoint on Hugging Face — never derived from the parameter count, which gets quantized releases wrong. VRAM to serve is those weights plus 20% for activations, CUDA context and allocator slack, at short context. A model is never sized above the precision it was released at.

Cheapest to serve is the lowest bookable rate, on-demand or serverless, for the smallest configuration a provider actually sells that holds the model — always labelled with its rental type, because a scale-to-zero rate is not an hourly one. The price is for the whole configuration.

License is the one field we did not measure: it is what the model card declares, and it is shown as such. There are no benchmark scores here on purpose — quality ratings are curated third-party data, and this catalog only carries numbers we can stand behind.

Keep going

Frequently asked questions

How is the VRAM needed to run each model calculated?
From the released checkpoint, not from the parameter count. We sum the weight files of each model's Hugging Face repo, classify the precision from the measured bytes per parameter, and add 20% for activations, CUDA context and allocator slack. Quantized sizes are only listed at or below the precision a model was actually released at, and a model we cannot measure confidently is held back rather than estimated.
What is the cheapest way to serve Qwen3 8B right now?
At 4-bit, Qwen3 8B needs 6.3 GB of VRAM and fits comfortably in 1× Tesla P4, bookable on demand from $0.02 per hour at Vast.ai. That is the cheapest such configuration among the bookable offerings we track, refreshed with the rest of the catalog; tight fits are left out of this answer.
Which models are listed?
23 open-weight LLMs, curated for popularity and measured from their released checkpoints. For each one the catalog lists the parameter count, the license as declared on the model card, the VRAM needed at every precision it was released at, and the cheapest cloud GPU configuration that holds it today, serverless included.
Why are there no benchmark scores?
Because we did not measure them. Benchmark results and arena ratings are third-party data that change with every evaluation harness, and this catalog only carries numbers we can stand behind: size, VRAM and the price to serve. Use it to shortlist models you can afford to run, then judge quality on your own task.