
01
vLLM and SGLang ship configured
Both templates come up as an OpenAI-compatible server. Change base_url in your existing client and leave the rest of the code alone.

02
Match the card to the model, not the other way round
An L4 serves embeddings for pennies and an H100 keeps a 70B fed. Moving between them is a redeploy, not a renegotiation.

03
Or run no pod at all
Serverless gives you the same OpenAI-compatible endpoint with models already deployed, billed per call and scaled to zero between them.
Ready to launch
Templates for this work
Official images with the drivers and the stack already in place.
- vLLM
- SGLang
- NVIDIA Triton Inference Server
- NVIDIA NIM
- Text Embeddings Inference
- Ollama
- OpenWebUI
Hardware
Cards that suit this work
Every Meshive pod gets a whole, non-virtualized GPU. These are the ones we would reach for first.
- NVIDIA L40S48GB at a steadier price than H100
- H100 SXM5Highest throughput for the 70B class
- RTX PRO 6000 Blackwell96GB when the context window is long
- NVIDIA L4Small models and embeddings, cheaply

Getting started
Three steps
- 1Start the vLLM or SGLang template and pass the model id you want served.
- 2The server comes up OpenAI-compatible - point your client at it and change nothing else.
- 3Scale by adding pods, or move the model to Serverless and stop managing them.

Pick a card and start the pod
Sign up, choose the GPU, and the pod is yours in under two minutes. It bills by the hour and stops when you stop it.
Already a user? Invite friends and earn 15% of their first top-up.



