Skip to content

Model directory

Every open model, one place

The 57 open models we can size, searchable by name and kind. Each card shows what it is, the memory it needs, and the cheapest card to rent it on. Open one to fit it in detail, or start from your own card in the Cookbook.

DeepSeek R1
ChatMoE

Mixture of experts, and the clearest case for two counts: 671B has to fit, 37B runs.

671B / 37Bneeds ~446 GB at int4from €14.96/hr
Qwen3 235B-A22B
ChatMoE

Mixture of experts. Needs datacentre memory, then generates at a 22B model's pace.

235B / 22Bneeds ~156 GB at int4from €2.24/hr
Llama 3.3 70B
Chat

Dense and large, with the long window the earlier 70B did not have.

70Bneeds ~47 GB at int4from €1.12/hr
Qwen3 32B
Chat

Dense, and about the largest that fits one 48 GB card at full precision.

32Bneeds ~21 GB at int4from €0.31/hr
QwQ 32B
Chat

A reasoning model: it emits far more tokens per answer, so cost per token dominates.

32Bneeds ~21 GB at int4from €0.31/hr
Qwen3 30B-A3B
ChatMoE

The local sweet spot: 30B of weights to hold, 3.3B of work per token.

30.5B / 3.3Bneeds ~20 GB at int4from €0.31/hr
Gemma 3 27B
Chat

Dense, long window, and takes images as well as text.

27Bneeds ~18 GB at int4from €0.31/hr
Mistral Small 3.1 24B
Chat

Sized to fit a single 32 GB card once quantised, which is what it does.

24Bneeds ~16 GB at int4from €0.31/hr
Phi-4 14B
Chat

Small and dense, with a short window that keeps the cache cheap.

14Bneeds ~9.3 GB at int4from €0.31/hr
Gemma 3 12B
Chat

The 27B model's smaller sibling. Fits a 16 GB card at full precision.

12Bneeds ~8.0 GB at int4from €0.31/hr
Qwen3 8B
Chat

Long context at a size a desktop card holds without quantising.

8Bneeds ~5.3 GB at int4from €0.31/hr
Gemma 3 4B
Chat

Small enough that the context cache, not the weights, is the larger half.

4Bneeds ~2.7 GB at int4from €0.31/hr
Qwen2.5 Coder 32B
Code

Code model at the size where a 48 GB card stops being optional.

32Bneeds ~21 GB at int4from €0.31/hr
Qwen2.5 72B
Chat

Large and long-context together, which is the most demanding pair.

72Bneeds ~48 GB at int4from €1.12/hr
Llama 3 70B
Chat

The original large one. Multi-card at full precision on anything but an H200.

70Bneeds ~47 GB at int4from €1.12/hr
Mixtral 8x7B
ChatMoE

Mixture of experts. All weights load even though two of eight run per token.

47B / 13Bneeds ~31 GB at int4from €0.78/hr
DeepSeek Coder 33B
Code

Larger code model. Wants a 48 GB card or better at full precision.

33Bneeds ~22 GB at int4from €0.31/hr
Gemma 2 27B
Chat

Mid-size. The awkward one that does not fit 24 GB at full precision.

27Bneeds ~18 GB at int4from €0.31/hr
Code Llama 13B
Code

Code completion at a size that fits one mid-range card.

13Bneeds ~8.6 GB at int4from €0.31/hr
LLaVA 1.6 13B
Vision

Images in, text out. The vision tower adds to the weights below.

13Bneeds ~8.6 GB at int4from €0.31/hr
Llama 3 8B
Chat

A small general chat model. The usual first thing anyone serves.

8Bneeds ~5.3 GB at int4from €0.31/hr
Qwen2.5 7B
Chat

Very long context. The cache, not the weights, is what costs here.

7Bneeds ~4.7 GB at int4from €0.31/hr
Mistral 7B
Chat

Small, long context for its size.

7Bneeds ~4.7 GB at int4from €0.31/hr
Phi-3 Mini
Chat

Small enough for a consumer card, with a long window.

3.8Bneeds ~2.5 GB at int4from €0.31/hr
BGE Large
Embeddings

Embeddings. Tiny, and throughput-bound rather than memory-bound.

0.335Bneeds ~0.2 GB at int4from €0.31/hr
Llama 3.1 405B
Chat

The largest open Llama. A datacentre of cards even quantised.

405Bneeds ~269 GB at int4from €7.48/hr
Llama 3.1 70B
Chat

The 70B that gained the long window Llama 3 lacked.

70Bneeds ~47 GB at int4from €1.12/hr
Llama 3.1 8B
Chat

Long context at a size a desktop card holds. A common default.

8Bneeds ~5.3 GB at int4from €0.31/hr
Llama 3.2 3B
Chat

Small and long-context, sized for a laptop.

3.2Bneeds ~2.1 GB at int4from €0.31/hr
Llama 3.2 1B
Chat

About as small as a useful chat model gets.

1.23Bneeds ~0.8 GB at int4from €0.31/hr
Llama 3.2 90B Vision
Vision

Images in, at the large end. The vision tower adds to the weights.

90Bneeds ~60 GB at int4from €1.12/hr
Llama 3.2 11B Vision
Vision

The smaller multimodal Llama. Text and images in.

11Bneeds ~7.3 GB at int4from €0.31/hr
Qwen2.5 32B
Chat

Dense 32B, about the largest that fits one 48 GB card at full precision.

32Bneeds ~21 GB at int4from €0.31/hr
Qwen2.5 14B
Chat

Mid-size dense chat with a long window.

14Bneeds ~9.3 GB at int4from €0.31/hr
Qwen2.5 3B
Chat

Small dense chat for a consumer card.

3Bneeds ~2.0 GB at int4from €0.31/hr
Qwen2.5 1.5B
Chat

Tiny, runs almost anywhere.

1.5Bneeds ~1.0 GB at int4from €0.31/hr
Qwen2.5 0.5B
Chat

The smallest of the family. The cache outweighs the weights.

0.5Bneeds ~0.3 GB at int4from €0.31/hr
Qwen2.5 Coder 14B
Code

Code model at a size a single mid-range card holds.

14Bneeds ~9.3 GB at int4from €0.31/hr
Qwen2.5 Coder 7B
Code

Small code model, long context for its size.

7Bneeds ~4.7 GB at int4from €0.31/hr
Qwen2-VL 72B
Vision

Large vision-language model. Reads images and documents.

72Bneeds ~48 GB at int4from €1.12/hr
Mistral Large 2 123B
Chat

Mistral's largest open weights, dense.

123Bneeds ~82 GB at int4from €3.74/hr
Mixtral 8x22B
ChatMoE

Mixture of experts: 141B load, 39B run per token.

141B / 39Bneeds ~94 GB at int4from €3.74/hr
Mistral Nemo 12B
Chat

A 12B with a long window, built with NVIDIA.

12Bneeds ~8.0 GB at int4from €0.31/hr
Codestral 22B
Code

Mistral's code model, sized for one 32 GB card.

22Bneeds ~15 GB at int4from €0.31/hr
Pixtral 12B
Vision

Mistral's multimodal 12B. Images and text in.

12Bneeds ~8.0 GB at int4from €0.31/hr
Gemma 2 9B
Chat

Dense 9B with a short, cheap window.

9Bneeds ~6.0 GB at int4from €0.31/hr
Gemma 2 2B
Chat

Small dense chat, runs on almost anything.

2.6Bneeds ~1.7 GB at int4from €0.31/hr
Phi-3.5 MoE
ChatMoE

Mixture of experts: 42B load, 6.6B run per token.

42B / 6.6Bneeds ~28 GB at int4from €0.78/hr
Phi-3 Medium 14B
Chat

The larger dense Phi, long context.

14Bneeds ~9.3 GB at int4from €0.31/hr
DeepSeek V3
ChatMoE

Mixture of experts: 671B load, 37B run. Datacentre memory to hold.

671B / 37Bneeds ~446 GB at int4from €14.96/hr
Command R+ 104B
Chat

Cohere's larger open model, built for retrieval and tools.

104Bneeds ~69 GB at int4from €1.12/hr
Command R 35B
Chat

The 35B Command R, long context.

35Bneeds ~23 GB at int4from €0.31/hr
Yi 1.5 34B
Chat

01.AI's 34B dense chat model.

34Bneeds ~23 GB at int4from €0.31/hr
DBRX
ChatMoE

Databricks' mixture of experts: 132B load, 36B run per token.

132B / 36Bneeds ~88 GB at int4from €3.74/hr
StarCoder2 15B
Code

Open code model trained on permissively-licensed source.

15Bneeds ~10.0 GB at int4from €0.31/hr
BGE-M3
Embeddings

Multilingual embeddings with a long input window.

0.567Bneeds ~0.4 GB at int4from €0.31/hr
Nomic Embed
Embeddings

Small open embeddings, long input for its size.

0.137Bneeds ~0.1 GB at int4from €0.31/hr

57 of 57 models · memory and price are estimates that err upward, not benchmarks