Skip to content
The platform

Rent a whole GPU, by the second.

Whole machines from people who own the cards, launched with your own key and settled in euro to the microeuro.

Compute

  • Whole machines, never a slice of one
  • Metered per second, in euro, to the microeuro
  • The rate is frozen when you match, not when you finish
  • Seven GPU models, from a 24GB 4090 to a 141GB H200

The lineup

Power, work, optimise

Three products on one machine. Compute is the power, Studio is the work, and Autopilot optimises it.

POWERWORKOPTIMISEa renter rents a carda team ships a workflowan owner sets a budgeta cardper stepwhat itcostCOMPUTELIVEcards from every provider, meteredSTUDIOPLANNEDsteps that each draw a cardAUTOPILOTPLANNEDreads the cost, proposes a changereleasedfetchno cardembedL40SanswerH100idle-holda card held while waitingright-sizea card bigger than the jobleft-runninga box nobody releasedback to the power: stop · resize · moveONE FOUNDATION · lit when this step uses itACCOUNTIN USEwho, and what they may doCATALOGUEIN USEevery card providers listMETERIN USEper second, every rentalLEDGERIN USE85 / 15, to the microeuroJOB
01 / 06
Compute · power

A renter asks for a card. Compute matches one from the catalogue, brings it up, and the meter starts counting seconds.

Studio and Autopilot are still being built. This is a worked example of how the three are designed to connect: the finding is one of Autopilot's worked examples, and the cards are catalogue names.

Compute · live

Everything Compute, in one console

Pick a side, pick a card, and the shapes, the fleet, the split and the work it suits all answer together.

ComputeLiveCardH100€2.89/hrto rent

Three shapes, one meter

Serverless for spiky inference, GPU Cloud for a box that's yours, Clusters for multi-node runs, all under the same per-second meter.

Scales with the traffic

You never hold a machine: a call lands on whichever card is warm, runs, and releases. Nothing is billed between calls.

Held between calls
Nothing
Idle cost
€0.00

Good for

Spiky traffic, demos, and anything that sits idle most of the day.

four cards · two workingautoscaling
a requestfrom your appa requestfrom your appa requestfrom your appthe front doorpicks a card that is readycard 1workingcard 2workingcard 3asleepcard 4asleepWHAT IS BILLED34.2 of 120 card-secondscard 1card 2The gaps are free. Nothing is held between requests, so nothing is charged for them.
illustrative
Billing
Per second
Rate
Frozen at match
Tenancy
Whole machine

If you are the one renting

Size it, launch it, stop paying

The Cookbook says which card a model needs and what an hour of it costs. Three commands take that card from nothing to settled.

Not sure which card fits your model?The cookbook works out the memory a model needs and the cheapest card that holds it.Open the cookbook
terminal · a rental in three lines
  1. $curl -fsSL https://kracht.ai/install | sh
    kracht v0.1.1 for linux/amd64
    checksum ok
    installed /usr/local/bin/kracht
  2. $kracht launch --gpu 1 --region eu-west
    Rented 8fa2c1d0 on h100-node-02 (eu-west) at 2.890000 EUR/hr
    Billing has started. Stop it with: kracht stop 8fa2c1d0
  3. $kracht stop 8fa2c1d0
    8fa2c1d0 is stopping. The provider picks this up on its next heartbeat, so `kracht ls` may still show it running for a few seconds. Billing stops at teardown.
Illustrative output. Rates are the catalogue's.Every command

Studio · planned

Studio is a canvas, so we made it one

Load a shape, click a step to inspect the card it holds, then switch to Trace to watch the agent itself: every plan, every tool call, every retry, timed and counted.

StudioPlannedillustrative
Shape
The questionand the session stateToolssearch, lookup, sqlPolicychecked before it leavesAn answeronce it stops re-planningA traceevery span, timedA billper second, per stepThe agentplan · act · observeObserve, then plan againuntil it answers or hits its step limit

each step rents its own card, for exactly as long as it runs

Step

agent

The agent loop: plan, call a tool, read the result, decide again. Runs until it answers or hits its step limit.

Card
H100 80GB
Estimate
2m
Est. cost
€0.08

Holds one card for the whole loop, however many times it goes round.

Estimated timeline0 → 2m
Estimated run cost€0.09
Wall clock2m
Billed while waiting€0.00

Autopilot · planned

Autopilot is a queue you work through

An inbox of findings, each with the metering that proves it, each waiting on you to approve or dismiss. It proposes; you decide.

AutopilotPlannedillustrative
Run spend
€196
six steps, per second
Flagged
€37
held but not used
Open
3
waiting on you
Applied
€0.00
until you approve

recommendation · idle-hold-0417

waiting on you

A card held while waiting on a person

The review step holds an H100 for the whole time a human takes to look at the output. Utilisation is flat zero for that window.

What it costs you today

€37/ run

Reversible · yes, at any time

The metering that proves it

GPU utilisation, one run0% for 2h 10m

0% GPU utilisation for 2h 10m

The changehold no card during review
nothing is applied until you press approve

The gate is the product. Autopilot never edits a running workload on its own. It writes a recommendation, attaches the metering it is based on, and waits for you.

Acts on its own
Never
Reads from
Real metering
Cost to decline
€0
Spend projection · today → end of monthworked example · Autopilot is plannedspent so far €842 · hover a card to see its range
€400€700€1,000€1,300€1,550Sep 8today · Sep 14Sep 23Sep 30€1,490approve nothing€1,371approve the flagged€1,298all + guardrailsdaily
spent so farprojectionlikely range
Approve nothing
€1,490
+€648 from here

The three findings keep costing what they cost.

Approve the flagged
€1,371
−€119 against doing nothing

The idle hold, the oversized card and the box nobody released.

Approve all + guardrailsbest outcome
−€192saved
€1,298 projected

Adds a ceiling per project and auto-release on idle.

A projection is not a promise. The chart starts today, from what you have actually spent. Everything on it is arithmetic on this month’s observed rate, not observation.

Worked example: one small account’s month, with sample daily figures that add up to the €842 the console states. The band is an illustrative range, not a confidence interval we have measured.

Roadmap

Now, next, later. Nothing dressed up.

Three columns, a live count on each, and a status chip on every item. Compute is finished; the other two you can look around today, and we will not pretend otherwise.

Now4 · shipped
  • Compute, Shipped

    Rent a whole GPU by the second, launched with your own key.

    live on real hardware

  • Per-second metering, Shipped

    Metered per second and settled in euro, to the microeuro.

    real rentals settled

  • The 85/15 split, Shipped

    Providers keep 85% of every second their card is rented.

    exact-integer split

  • EU data residency, Shipped

    Hardware and ledger inside the EU, priced in euro.

    euro ledger

Next3 · in progress
  • Studio, In progress

    Wire many rentals into one workflow you can read back.

    demo you can look around

  • Agent tracing, In progress

    Spans, tool calls and tokens recorded for every run.

    observability surface

  • Evaluation, In progress

    Score a prompt set across models in one run.

    designed

Later3 · planned
  • Autopilot, Planned

    Find the spend you are not using, and stop it on approval.

    designed, not built

  • Spend guardrails, Planned

    A ceiling per project, and a warning before a run passes it.

    designed

  • Auto-release on idle, Planned

    Stop a rental when its job exits, once you have approved it.

    designed

Two sides, one marketplace

Spin up a GPU, or earn from your own

Rent on-demand compute by the second, or list idle hardware and let it pay for itself. Same marketplace, both directions.

No card required

85%
revenue to providers
Per-second
billing, live
EUR
metered to the microeuro