Compare
PreviewTwo to four models, side by side
Put a handful of models in the same row and read the trade-off: how fast each runs, the memory it takes, and what a million tokens costs. Speed, memory and cost are real estimates; the answers are a shared sample, because comparing generated text would be a quality claim this page will not make.
Qwen3 8B
8B params
A GPU runs one small calculation across thousands of cores at once, which is the shape of the matrix maths a model is built from. So work that crawls on a CPU streams in real time here.
sample answer
- Speed
- ~176 t/s
- Memory
- 5.3 GB
- Cost / 1M
- €0.49
- First token
- ~180 ms
Phi-4 14B
14B params
A GPU runs one small calculation across thousands of cores at once, which is the shape of the matrix maths a model is built from. So work that crawls on a CPU streams in real time here.
sample answer
- Speed
- ~101 t/s
- Memory
- 9.3 GB
- Cost / 1M
- €0.85
- First token
- ~180 ms
Qwen3 32B
32B params
A GPU runs one small calculation across thousands of cores at once, which is the shape of the matrix maths a model is built from. So work that crawls on a CPU streams in real time here.
sample answer
- Speed
- ~44 t/s
- Memory
- 21 GB
- Cost / 1M
- €1.95
- First token
- ~180 ms
Speed, memory and cost are estimates that err upward, computed from each model and the card; a green tick marks the best value in the row. First-token time is illustrative, and the answers are a shared sample.
A real compare runs the same prompt through each model and puts the actual answers, latencies and costs beside each other: the agent runtime, which is the platform under this. The sizing that makes the speed, memory and cost real today is the Cookbook.
Browse models