Which open-weight LLM should you run?
Compare 23 models by size, license and the VRAM they actually need, and see the cheapest cloud GPU that serves each one right now.
| Compare | Model | License | VRAM to serve | Cheapest to serve | Max context | Actions |
|---|---|---|---|---|---|---|
| Qwen3 8BAlibaba8.2Breleased at 16-bit | Not declared |
| 40K | What runs it? | ||
| Gemma 4 31BGoogle31.3Breleased at 16-bit | Not declared |
| 256K | What runs it? | ||
| Qwen3 4BAlibaba4Breleased at 16-bit | Not declared |
| 40K | What runs it? | ||
| gpt-oss-20bOpenAI21B3.6B activereleased at 4-bit | Not declared |
| 128K | What runs it? | ||
| Llama 3.1 8B InstructMeta8Breleased at 16-bit | Not declared |
| — | What runs it? | ||
| Qwen3.6 27BAlibaba27.8Breleased at 16-bit | Not declared |
| 256K | What runs it? | ||
| gpt-oss-120bOpenAI117B5.1B activereleased at 4-bit | Not declared |
| 128K | What runs it? | ||
| Qwen3 32BAlibaba32.8Breleased at 16-bit | Not declared |
| 40K | What runs it? | ||
| Qwen3.6 35B-A3BAlibaba35B3B activereleased at 16-bit | Not declared |
| 256K | What runs it? | ||
| Mistral 7B Instruct v0.3Mistral AI7.2Breleased at 16-bit | Not declared |
| 32K | What runs it? | ||
| Kimi K3Moonshot AI2.8T104B activereleased at 4-bit | Not declared |
| 1,024K | What runs it? | ||
| GLM-4.7 FlashZ.ai31.2Breleased at 16-bit | Not declared |
| 198K | What runs it? | ||
| DeepSeek V4 FlashDeepSeek284Breleased at 4-bit | Not declared |
| 1,024K | What runs it? | ||
| GLM-5.2Z.ai753Breleased at 16-bit | Not declared |
| 1,024K | What runs it? | ||
| DeepSeek R1DeepSeek671B37B activereleased at 8-bit | Not declared |
| 160K | What runs it? | ||
| Granite 4.1 8BIBM8.8Breleased at 16-bit | Not declared |
| 128K | What runs it? | ||
| Gemma 3 12BGoogle12.2Breleased at 16-bit | Not declared |
| — | What runs it? | ||
| DeepSeek V4 ProDeepSeek1.6T49B activereleased at 4-bit | Not declared |
| 1,024K | What runs it? | ||
| Kimi K2.5Moonshot AI1T32B activereleased at 4-bit | Not declared |
| 256K | What runs it? | ||
| Qwen3 235B-A22BAlibaba235B22B activereleased at 16-bit | Not declared |
| 40K | What runs it? | ||
| Mistral Small 3.2 24BMistral AI24Breleased at 16-bit | Not declared |
| 128K | What runs it? | ||
| Laguna-S 2.1Poolside118B8B activereleased at 16-bit | Not declared |
| 1,024K | What runs it? | ||
| Solar Open2 250BUpstage250B15B activereleased at 16-bit | Not declared |
| 1,024K | What runs it? |
How this table is made
Every size is measured from the weight files of the released checkpoint on Hugging Face — never derived from the parameter count, which gets quantized releases wrong. VRAM to serve is those weights plus 20% for activations, CUDA context and allocator slack, at short context. A model is never sized above the precision it was released at.
Cheapest to serve is the lowest bookable rate, on-demand or serverless, for the smallest configuration a provider actually sells that holds the model — always labelled with its rental type, because a scale-to-zero rate is not an hourly one. The price is for the whole configuration.
License is the one field we did not measure: it is what the model card declares, and it is shown as such. There are no benchmark scores here on purpose — quality ratings are curated third-party data, and this catalog only carries numbers we can stand behind.