Your own replicas, behind one URL
A deployment puts one model behind one OpenAI-compatible endpoint. Replicas come and go; the URL does not. You never pick a GPU — you set a price ceiling, and the cheapest card that fits is assigned and re-assigned for you.

Four controls, and no GPU picking
Every candidate GPU's hourly price is shown before you confirm.
| Control | What it decides |
|---|---|
| Price cap ($/replica-hour) | A hard ceiling — a GPU above it is never assigned, at deploy time or later |
| Min / max replicas | The range autoscaling may move within |
| Autoscale | On: follow load inside the range. Off: hold at min |
| Output storage | The default bucket for generated images and video |
Scaling that reads the work, not the request count
- 01
The cap is a ceiling, not a suggestion
Raise it and the GPU pool widens immediately. Lower it and over-cap replicas are replaced gradually, one at a time, with progress on the card — the endpoint never dips while that happens.
- 02
Load is measured in tokens
The scaler re-evaluates every 30 seconds over a rolling window. A 100k-token document analysis and a two-line chat are weighted by the memory they actually hold, and a replica on a larger GPU counts as more capacity than one on a smaller card.
- 03
Careful on the way down
A replica is only removed when nothing is queued, load has stayed low across the window, and the survivors could absorb its share and still sit below the scale-up point. Two replicas at 40% each stay at two, however idle they look.
- 04
Billed by GPU-hour of running replicas
No per-request and no per-token charge on your own deployments. Pause and the replicas wind down in about 30 seconds — configuration, URL and registration are kept, and billing stops with them.

Pick a card and start the pod
Sign up, choose the GPU, and the pod is yours in under two minutes. It bills by the hour and stops when you stop it.
Already a user? Invite friends and earn 15% of their first top-up.
