Skip to content

Compare

Preview

Two to four models, side by side

Put a handful of models in the same row and read the trade-off: how fast each runs, the memory it takes, and what a million tokens costs. Speed, memory and cost are real estimates; the answers are a shared sample, because comparing generated text would be a quality claim this page will not make.

Models · pick 2 to 4

Qwen3 8B

8B params

A GPU runs one small calculation across thousands of cores at once, which is the shape of the matrix maths a model is built from. So work that crawls on a CPU streams in real time here.

sample answer

Speed
~176 t/s
Memory
5.3 GB
Cost / 1M
€0.49
First token
~180 ms

Phi-4 14B

14B params

A GPU runs one small calculation across thousands of cores at once, which is the shape of the matrix maths a model is built from. So work that crawls on a CPU streams in real time here.

sample answer

Speed
~101 t/s
Memory
9.3 GB
Cost / 1M
€0.85
First token
~180 ms

Qwen3 32B

32B params

A GPU runs one small calculation across thousands of cores at once, which is the shape of the matrix maths a model is built from. So work that crawls on a CPU streams in real time here.

sample answer

Speed
~44 t/s
Memory
21 GB
Cost / 1M
€1.95
First token
~180 ms

Speed, memory and cost are estimates that err upward, computed from each model and the card; a green tick marks the best value in the row. First-token time is illustrative, and the answers are a shared sample.

A real compare runs the same prompt through each model and puts the actual answers, latencies and costs beside each other: the agent runtime, which is the platform under this. The sizing that makes the speed, memory and cost real today is the Cookbook.

Browse models