Metered per second
From the moment it runs to the moment you stop it, in euro. A stopped instance costs nothing.
Compute, storage, networking and serving. Take one or all of them, keep them for ninety seconds or a month, and pay for exactly that.
The fleet
Pick any card in the catalogue and you launch it the same way: one command, your image, your key, billed per second.
The cookbook
Browse every model against the fleet: what it needs to run, the cheapest card that holds it, its speed, and roughly what a million tokens comes to. No account, no sign-in.
| Model | Params | What it needs to run | VRAM | Fits on | Tok/s | Per 1M tok | Per hour |
|---|---|---|---|---|---|---|---|
Qwen3 235B-A22B Mixture of experts. Needs datacentre memory, then generates at a 22B model's pace. | 235B22B active | at int8 · 8k ctx | 295 GB | A100×44 cards | ~65 | €19.18 | €4.48Find |
Qwen3 32B Dense, and about the largest that fits one 48 GB card at full precision. | 32B | at int8 · 8k ctx | 40 GB | A100fits | ~45 | €6.98 | €1.12Find |
Qwen3 30B-A3B The local sweet spot: 30B of weights to hold, 3.3B of work per token. | 30.5B3.3B active | at int8 · 8k ctx | 38 GB | A100tight | ~330 | €0.66 | €0.78Find |
Qwen3 8B Long context at a size a desktop card holds without quantising. | 8B | at int8 · 8k ctx | 10 GB | RTX 4090fits | ~88 | €0.98 | €0.31Find |
QwQ 32B A reasoning model: it emits far more tokens per answer, so cost per token dominates. | 32B | at int8 · 8k ctx | 40 GB | A100fits | ~45 | €6.98 | €1.12Find |
Gemma 3 27B Dense, long window, and takes images as well as text. | 27B | at int8 · 8k ctx | 34 GB | A100fits | ~40 | €5.37 | €0.78Find |
Mistral Small 3.1 24B Sized to fit a single 32 GB card once quantised, which is what it does. | 24B | at int8 · 8k ctx | 30 GB | A100fits | ~45 | €4.78 | €0.78Find |
Gemma 3 12B The 27B model's smaller sibling. Fits a 16 GB card at full precision. | 12B | at int8 · 8k ctx | 15 GB | RTX 4090fits | ~59 | €1.46 | €0.31Find |
Gemma 3 4B Small enough that the context cache, not the weights, is the larger half. | 4B | at int8 · 8k ctx | 5.0 GB | RTX 4090fits | ~176 | €0.49 | €0.31Find |
DeepSeek R1 Mixture of experts, and the clearest case for two counts: 671B has to fit, 37B runs. | 671B37B active | at int8 · 8k ctx | 841 GB | H200×66 cards | ~91 | €68.64 | €22.44Find |
57 models · 57 run on a card in the catalogue at int8, 8k context. Cheapest serving card, largest saving first. Estimates that err upward, not measurements.
This rental · metered per second · example
Settles as · sample rows
Last 60 seconds · one tick a second
€0.000803 each
Rate
frozen at match
Split
85 / 15
Rent a card by the second, launch it with your own key, and stop paying the moment you stop. Whatever providers have online, at today's price. Follow one rental end to end.
From the moment it runs to the moment you stop it, in euro. A stopped instance costs nothing.
Launch a notebook or a container on somebody's idle card, with your own SSH key.
Whatever providers have online right now. A machine already rented is not offered to you.
The rate is locked onto the rental, so a provider re-pricing never changes what you pay.
Every metered second splits
exact integer, no cent leaks
Metered per second
Metered per second, a job pays for the minutes it ran. Rounded up to the hour, it pays for the rest of the hour too, and the waste is worst just past each one.
illustrative · list rate × job length, from the one price list · 1 square is 10 minutes of the card
The fleet
From flagship to cost-efficient consumer cards, every GPU is metered by the second with marketplace pricing.
Top of the stack for the biggest models.
Built for serious training runs.
The dependable workhorse.
Great value for inference and media.
Fast, flexible, surprisingly capable.
Lean and cheap for steady inference.
Spec sheet
| GPU | Architecture | VRAM | Bandwidth | Interconnect | € / GPU-hour | Source |
|---|---|---|---|---|---|---|
| NVIDIA H200 | Hopper | 141 GB HBM3e | 4.8 TB/s | NVLink 4 | €3.74 | nvidia.com |
| NVIDIA H100 | Hopper | 80 GB HBM3 | 3.35 TB/s | NVLink 4 | €2.89 | nvidia.com |
| NVIDIA A100 | Ampere | 80 GB HBM2e | 2.0 TB/s | NVLink 3 | €1.12 | nvidia.com |
| NVIDIA A100 | Ampere | 40 GB HBM2e | 1.6 TB/s | NVLink 3 | €0.78 | nvidia.com |
| NVIDIA L40S | Ada Lovelace | 48 GB GDDR6 | 864 GB/s | PCIe 4 | €0.64 | nvidia.com |
| NVIDIA RTX 4090 | Ada Lovelace | 24 GB GDDR6X | 1.0 TB/s | PCIe 4 | €0.31 | nvidia.com |
| NVIDIA A10 | Ampere | 24 GB GDDR6 | 600 GB/s | PCIe 4 | €0.22 | nvidia.com |
In practice
Renting a card by the second changes what you can afford to try. Each of these is a job it is built for. Open one to see how it holds up.
Take an H100 for the length of a training run and give it back. No reservation, no minimum, and no card sitting idle between experiments.
One console
Provision GPUs, stream deploy logs, and monitor utilisation. Switch between instances without leaving the page.
A rental
↑ click an instance to switch. Preview data
One of them is probably you. Both run on the same ledger, the same regions and the same per-second meter.
FIG 01 the marketplace
renter cli · kracht launch
The renter's own machine. One command becomes a signed request naming a GPU model and region, never a machine.
Earn from idle
You set the price, you keep 85%, Kracht takes 15% and handles the renter, the metering and the payout. Drag to your setup.
● in the band: priced to rent
How much of the month a card is actually rented.
What you paid for a card, to see payback time.
After the 15% Kracht fee, on 4 cards at 65% utilisation.
≈ €1,166 per card · €55,949 a year across 4 cards.
Bars are illustrative monthly take-home; it moves with demand.
Every primitive, wired together
Instances, storage and keys are one account and one meter, not four products you join up yourself.
One meter
Every primitive bills against the same per-second clock.
One ledger
Append-only entries, unique on an idempotency key.
One account
Instances, volumes and keys hang off the billable account.
Start where it runs
No reservation, no minimum, your own SSH key. This is the real path, not a preview: Compute runs end to end, metered per second, in euro.
Launch a GPU
Resources
Compute is the part of the marketplace that runs today. These are the three places worth going once you have seen how it launches.
Documentation
How it works, endpoint by endpoint
Pricing
Every rate, and what a month of one card costs
Browse live capacity
What is actually available, at today's prices
Changelog
What shipped, dated, including what did not
Status
What is up, and what was not
Help centre
The questions that come up more than once
Per GPU-hour, in euro
It runs. Rentals have been placed, metered and settled end to end on real hardware, in euro, with the split applied. The parts of Kracht that are not finished are Studio and Autopilot, and each of those pages says so at the top.
Continue through the platform
Rent on-demand compute by the second, or list idle hardware and let it pay for itself. Same marketplace, both directions.
No card required