Meshive GPU Cloud logoMeshive

Your models. Your endpoint. OpenAI‑compatible.

Pick a model — get a dedicated HTTPS API in minutes. No framework setup, no infra to manage.

main.py
from openai import OpenAI

client = OpenAI(
    base_url="https://llama-a1b2.inference.meshive.ai/v1",
    api_key="sk_...",
)

resp = client.chat.completions.create(
    model="llama-3.1-8b-instruct",
    messages=[{"role": "user", "content": "Hello!"}],
    stream=True,  # SSE streaming, out of the box
)

Pick a model. We handle the rest.

Frameworks, GPU matching, serving config — already done. Just pick a model.

Text

LLMs & Embeddings

QwenDeepSeekKimiGLMGemma
Served by vLLM — pre-configured

Image

Text-to-Image & Image-to-Image

Qwen-ImageSDXLFLUX.1
Served by ComfyUI — pre-configured

Video

Text-to-Video & Image-to-Video

WanLTX-Video
Served by ComfyUI — pre-configured
Serverless

Model Catalog

Most popular models are already registered — click to deploy. Or register your own.

Model templates55Text · Image · Video
Custom models+Any HF repo
CompatibilityOpenAI SDKDrop-in client
All 55TextImageVideoSpeech soonEmbedding soon
ModelContexttext onlyMin VRAM
Q
Qwen3-0.6BTextBF16
HFQwen/Qwen3-0.6B
41K
7 GB
Q
Qwen-ImageImageT2I
HFQwen/Qwen-Image
—
32 GB
W
Wan2.1-T2V-1.3BVideoT2V
HFWan-AI/Wan2.1-T2V-1.3B-Diffusers
—
10 GB
Q
Qwen3-1.7BTextBF16
HFQwen/Qwen3-1.7B
41K
9 GB
F
FLUX.1-schnellImageT2I
HFblack-forest-labs/FLUX.1-schnell
—
24 GB
•••+50 more model templates
55 official templates — text, image & video, ready to deploy in one click.

Any Hugging Face repo. Deployed.

Public or private — paste a repo and deploy. The backend derives everything else.

Paste a repo — we derive the rest

The backend reads the Hugging Face repo and auto-fills parameter count, context length, and config. Nothing to tune.

The engine is picked for you

LLMs run on vLLM; image & video run on SGLang Diffusion — auto-resolved on the server. No framework choice to make.

Public or private

Public repos work out of the box. For gated or private models, add a saved Hugging Face token.

Register custom modelAuto-configured

Text in, text out · Served by vLLM, auto-resolved on the server.

meta-llama/Llama-3.1-8B-Instruct

Backend reads the repo and auto-fills its config.

Private model
Use saved tokenhf-prod

An endpoint that's yours alone

Dedicated replicas, keys you control, and live metrics — built for production traffic.

Dedicated HTTPS endpoint

{slug}.inference.meshive.ai/v1, TLS by default. Your replicas serve only you — no noisy neighbors, no shared rate limits.

API keys you control

Issue, revoke, and rotate keys. Set per-key RPM and daily token limits.

Autoscale 1 → 10 replicas

Least-loaded load balancing across replicas. A price cap puts you in control of spend.

No cold starts

At least one replica stays warm at all times — your first request responds instantly.

Meta-Llama-3-8B-Instruct-AWQ

vLLMchat
Running
Active replicas
2 / 3
TTFT P95182 ms
Total tok/s1,284
Concurrent14
Error rate0%
GPU 2× RTX 4090 · 24GB · 61%
POSTllama-a1b2.inference.meshive.ai/v1
Playground$0.16/hr

Replica status

Meta-Llama-3-8B-Instruct-AWQ · read-only

Active2 / 3
Queue0
Endpoint P95214ms
GPU61%
Per-replica · 2
R1Running1× RTX 4090 · 24GB
$0.08/hr
TTFT176ms
Avg tok/s648
Conc8
KV cache42%
R2Running1× RTX 4090 · 24GB
$0.08/hr
TTFT191ms
Avg tok/s636
Conc6
KV cache38%

Live metrics for every deployment — TTFT, throughput, and per-replica GPU utilization, right in the console.

Fire and forget

Long-running image & video jobs? Submit async, get an HMAC-signed webhook when it's done. Every run lands in the Artifacts gallery and is delivered to your own S3, R2, or MinIO bucket — or use ours.

CodePlayground
# Code or Playground — same job
job = client.images.generate(
    model="flux.1-dev",
    prompt="neon city at dawn",
)  # async
queuedjob_3f9a…
POST /async

Meshive

queue → GPU
on complete
Artifacts
your-bucket/outputs/flux/job_3f9a.png
saved

Choose where your outputs land

Every finished run shows up in the Artifacts gallery — and is copied to the storage you pick.

Image & video · set at deploy or switch anytime from the serving card

Meshive managed

10-day

Zero setup. Outputs auto-delete after 10 days.

Your bucket

Kept

S3, R2, or MinIO — any S3-compatible store. Kept permanently.

See it run before you commit.

Chat with an LLM, generate an image, render a video — right in the Playground.

ChatImageVideo