Chat: a speech bubble DeepSeek R1 Chat MoE Mixture of experts, and the clearest case for two counts: 671B has to fit, 37B runs.
671B / 37B needs ~446 GB at int4 from €14.96/hr
Chat: a speech bubble Qwen3 235B-A22B Chat MoE Mixture of experts. Needs datacentre memory, then generates at a 22B model's pace.
235B / 22B needs ~156 GB at int4 from €2.24/hr
Chat: a speech bubble Llama 3.3 70B Chat Dense and large, with the long window the earlier 70B did not have.
70B needs ~47 GB at int4 from €1.12/hr
Chat: a speech bubble Qwen3 32B Chat Dense, and about the largest that fits one 48 GB card at full precision.
32B needs ~21 GB at int4 from €0.31/hr
Chat: a speech bubble QwQ 32B Chat A reasoning model: it emits far more tokens per answer, so cost per token dominates.
32B needs ~21 GB at int4 from €0.31/hr
Chat: a speech bubble Qwen3 30B-A3B Chat MoE The local sweet spot: 30B of weights to hold, 3.3B of work per token.
30.5B / 3.3B needs ~20 GB at int4 from €0.31/hr
Chat: a speech bubble Gemma 3 27B Chat Dense, long window, and takes images as well as text.
27B needs ~18 GB at int4 from €0.31/hr
Chat: a speech bubble Mistral Small 3.1 24B Chat Sized to fit a single 32 GB card once quantised, which is what it does.
24B needs ~16 GB at int4 from €0.31/hr
Chat: a speech bubble Phi-4 14B Chat Small and dense, with a short window that keeps the cache cheap.
14B needs ~9.3 GB at int4 from €0.31/hr
Chat: a speech bubble Gemma 3 12B Chat The 27B model's smaller sibling. Fits a 16 GB card at full precision.
12B needs ~8.0 GB at int4 from €0.31/hr
Chat: a speech bubble Qwen3 8B Chat Long context at a size a desktop card holds without quantising.
8B needs ~5.3 GB at int4 from €0.31/hr
Chat: a speech bubble Gemma 3 4B Chat Small enough that the context cache, not the weights, is the larger half.
4B needs ~2.7 GB at int4 from €0.31/hr
Code: angle brackets Qwen2.5 Coder 32B Code Code model at the size where a 48 GB card stops being optional.
32B needs ~21 GB at int4 from €0.31/hr
Chat: a speech bubble Qwen2.5 72B Chat Large and long-context together, which is the most demanding pair.
72B needs ~48 GB at int4 from €1.12/hr
Chat: a speech bubble Llama 3 70B Chat The original large one. Multi-card at full precision on anything but an H200.
70B needs ~47 GB at int4 from €1.12/hr
Chat: a speech bubble Mixtral 8x7B Chat MoE Mixture of experts. All weights load even though two of eight run per token.
47B / 13B needs ~31 GB at int4 from €0.78/hr
Code: angle brackets DeepSeek Coder 33B Code Larger code model. Wants a 48 GB card or better at full precision.
33B needs ~22 GB at int4 from €0.31/hr
Chat: a speech bubble Gemma 2 27B Chat Mid-size. The awkward one that does not fit 24 GB at full precision.
27B needs ~18 GB at int4 from €0.31/hr
Code: angle brackets Code Llama 13B Code Code completion at a size that fits one mid-range card.
13B needs ~8.6 GB at int4 from €0.31/hr
Vision: an eye LLaVA 1.6 13B Vision Images in, text out. The vision tower adds to the weights below.
13B needs ~8.6 GB at int4 from €0.31/hr
Chat: a speech bubble Llama 3 8B Chat A small general chat model. The usual first thing anyone serves.
8B needs ~5.3 GB at int4 from €0.31/hr
Chat: a speech bubble Qwen2.5 7B Chat Very long context. The cache, not the weights, is what costs here.
7B needs ~4.7 GB at int4 from €0.31/hr
Chat: a speech bubble Mistral 7B Chat Small, long context for its size.
7B needs ~4.7 GB at int4 from €0.31/hr
Chat: a speech bubble Phi-3 Mini Chat Small enough for a consumer card, with a long window.
3.8B needs ~2.5 GB at int4 from €0.31/hr
Catalogue: a grid of tiles BGE Large Embeddings Embeddings. Tiny, and throughput-bound rather than memory-bound.
0.335B needs ~0.2 GB at int4 from €0.31/hr
Chat: a speech bubble Llama 3.1 405B Chat The largest open Llama. A datacentre of cards even quantised.
405B needs ~269 GB at int4 from €7.48/hr
Chat: a speech bubble Llama 3.1 70B Chat The 70B that gained the long window Llama 3 lacked.
70B needs ~47 GB at int4 from €1.12/hr
Chat: a speech bubble Llama 3.1 8B Chat Long context at a size a desktop card holds. A common default.
8B needs ~5.3 GB at int4 from €0.31/hr
Chat: a speech bubble Llama 3.2 3B Chat Small and long-context, sized for a laptop.
3.2B needs ~2.1 GB at int4 from €0.31/hr
Chat: a speech bubble Llama 3.2 1B Chat About as small as a useful chat model gets.
1.23B needs ~0.8 GB at int4 from €0.31/hr
Vision: an eye Llama 3.2 90B Vision Vision Images in, at the large end. The vision tower adds to the weights.
90B needs ~60 GB at int4 from €1.12/hr
Vision: an eye Llama 3.2 11B Vision Vision The smaller multimodal Llama. Text and images in.
11B needs ~7.3 GB at int4 from €0.31/hr
Chat: a speech bubble Qwen2.5 32B Chat Dense 32B, about the largest that fits one 48 GB card at full precision.
32B needs ~21 GB at int4 from €0.31/hr
Chat: a speech bubble Qwen2.5 14B Chat Mid-size dense chat with a long window.
14B needs ~9.3 GB at int4 from €0.31/hr
Chat: a speech bubble Qwen2.5 3B Chat Small dense chat for a consumer card.
3B needs ~2.0 GB at int4 from €0.31/hr
Chat: a speech bubble Qwen2.5 1.5B Chat Tiny, runs almost anywhere.
1.5B needs ~1.0 GB at int4 from €0.31/hr
Chat: a speech bubble Qwen2.5 0.5B Chat The smallest of the family. The cache outweighs the weights.
0.5B needs ~0.3 GB at int4 from €0.31/hr
Code: angle brackets Qwen2.5 Coder 14B Code Code model at a size a single mid-range card holds.
14B needs ~9.3 GB at int4 from €0.31/hr
Code: angle brackets Qwen2.5 Coder 7B Code Small code model, long context for its size.
7B needs ~4.7 GB at int4 from €0.31/hr
Vision: an eye Qwen2-VL 72B Vision Large vision-language model. Reads images and documents.
72B needs ~48 GB at int4 from €1.12/hr
Chat: a speech bubble Mistral Large 2 123B Chat Mistral's largest open weights, dense.
123B needs ~82 GB at int4 from €3.74/hr
Chat: a speech bubble Mixtral 8x22B Chat MoE Mixture of experts: 141B load, 39B run per token.
141B / 39B needs ~94 GB at int4 from €3.74/hr
Chat: a speech bubble Mistral Nemo 12B Chat A 12B with a long window, built with NVIDIA.
12B needs ~8.0 GB at int4 from €0.31/hr
Code: angle brackets Codestral 22B Code Mistral's code model, sized for one 32 GB card.
22B needs ~15 GB at int4 from €0.31/hr
Vision: an eye Pixtral 12B Vision Mistral's multimodal 12B. Images and text in.
12B needs ~8.0 GB at int4 from €0.31/hr
Chat: a speech bubble Gemma 2 9B Chat Dense 9B with a short, cheap window.
9B needs ~6.0 GB at int4 from €0.31/hr
Chat: a speech bubble Gemma 2 2B Chat Small dense chat, runs on almost anything.
2.6B needs ~1.7 GB at int4 from €0.31/hr
Chat: a speech bubble Phi-3.5 MoE Chat MoE Mixture of experts: 42B load, 6.6B run per token.
42B / 6.6B needs ~28 GB at int4 from €0.78/hr
Chat: a speech bubble Phi-3 Medium 14B Chat The larger dense Phi, long context.
14B needs ~9.3 GB at int4 from €0.31/hr
Chat: a speech bubble DeepSeek V3 Chat MoE Mixture of experts: 671B load, 37B run. Datacentre memory to hold.
671B / 37B needs ~446 GB at int4 from €14.96/hr
Chat: a speech bubble Command R+ 104B Chat Cohere's larger open model, built for retrieval and tools.
104B needs ~69 GB at int4 from €1.12/hr
Chat: a speech bubble Command R 35B Chat The 35B Command R, long context.
35B needs ~23 GB at int4 from €0.31/hr
Chat: a speech bubble Yi 1.5 34B Chat 01.AI's 34B dense chat model.
34B needs ~23 GB at int4 from €0.31/hr
Chat: a speech bubble DBRX Chat MoE Databricks' mixture of experts: 132B load, 36B run per token.
132B / 36B needs ~88 GB at int4 from €3.74/hr
Code: angle brackets StarCoder2 15B Code Open code model trained on permissively-licensed source.
15B needs ~10.0 GB at int4 from €0.31/hr
Catalogue: a grid of tiles BGE-M3 Embeddings Multilingual embeddings with a long input window.
0.567B needs ~0.4 GB at int4 from €0.31/hr
Catalogue: a grid of tiles Nomic Embed Embeddings Small open embeddings, long input for its size.
0.137B needs ~0.1 GB at int4 from €0.31/hr