A verified catalog, and any repo after that
The official catalog is curated, measured and one click from serving. Past it the engine will run most open architectures, so the model you care about is probably already supported.

What runs where
Serverless bills by the GPU-hour its replicas run. A few image and video families ship as ComfyUI packs on a pod instead.
| Modality | Families | Where it runs |
|---|---|---|
| Text and embeddings | Qwen · DeepSeek · Kimi · GLM · Gemma | Serverless, per GPU-hour |
| Image | Qwen-Image · Krea 2 · SDXL · FLUX.1 | Serverless, plus ComfyUI packs on a pod |
| Video | MiniMax H3 · Wan · LTX-Video | Serverless, plus ComfyUI packs on a pod |
The numbers that decide it
- 01
Measured VRAM, not a guess
Every catalog card carries context length, quantization, parameter count and the measured VRAM requirement — the number that decides which GPUs can serve it. There is no price on the card, because cost only exists as the hourly GPU price you cap at deploy time.
- 02
Quantization picks the GPU, not you
FP16, AWQ, GPTQ, GGUF and INT8/4 run anywhere in the fleet. FP8 needs an Ada, Hopper or Blackwell card; FP4 needs Blackwell. Cards that cannot run the weights are excluded from the deploy panel with the reason shown.
- 03
Your own repo is one field
Register any Hugging Face repo with nothing but its name. Modality, framework, file size and the VRAM estimate are derived from the metadata, and every derived value can be overridden.
- 04
Or call one somebody else already runs
Shared serving lists deployments other teams have opened to guests. Nothing to deploy, billed per token, and your existing workspace key already works.

Pick a card and start the pod
Sign up, choose the GPU, and the pod is yours in under two minutes. It bills by the hour and stops when you stop it.
Already a user? Invite friends and earn 15% of their first top-up.
