Your models. Your endpoint. OpenAI‑compatible.
Pick a model — get a dedicated HTTPS API in minutes. No framework setup, no infra to manage.
from openai import OpenAI client = OpenAI( base_url="https://llama-a1b2.inference.meshive.ai/v1", api_key="sk_...", ) resp = client.chat.completions.create( model="llama-3.1-8b-instruct", messages=[{"role": "user", "content": "Hello!"}], stream=True, # SSE streaming, out of the box )
Pick a model. We handle the rest.
Frameworks, GPU matching, serving config — already done. Just pick a model.
Text
LLMs & Embeddings
Image
Text-to-Image & Image-to-Image
Video
Text-to-Video & Image-to-Video
Model Catalog
Most popular models are already registered — click to deploy. Or register your own.
Any Hugging Face repo. Deployed.
Public or private — paste a repo and deploy. The backend derives everything else.
Paste a repo — we derive the rest
The backend reads the Hugging Face repo and auto-fills parameter count, context length, and config. Nothing to tune.
The engine is picked for you
LLMs run on vLLM; image & video run on SGLang Diffusion — auto-resolved on the server. No framework choice to make.
Public or private
Public repos work out of the box. For gated or private models, add a saved Hugging Face token.
Text in, text out · Served by vLLM, auto-resolved on the server.
Backend reads the repo and auto-fills its config.
An endpoint that's yours alone
Dedicated replicas, keys you control, and live metrics — built for production traffic.
Dedicated HTTPS endpoint
{slug}.inference.meshive.ai/v1, TLS by default. Your replicas serve only you — no noisy neighbors, no shared rate limits.
API keys you control
Issue, revoke, and rotate keys. Set per-key RPM and daily token limits.
Autoscale 1 → 10 replicas
Least-loaded load balancing across replicas. A price cap puts you in control of spend.
No cold starts
At least one replica stays warm at all times — your first request responds instantly.
Meta-Llama-3-8B-Instruct-AWQ
Replica status
Meta-Llama-3-8B-Instruct-AWQ · read-only
$0.08/hr
$0.08/hr
Live metrics for every deployment — TTFT, throughput, and per-replica GPU utilization, right in the console.
Fire and forget
Long-running image & video jobs? Submit async, get an HMAC-signed webhook when it's done. Every run lands in the Artifacts gallery and is delivered to your own S3, R2, or MinIO bucket — or use ours.
# Code or Playground — same job job = client.images.generate( model="flux.1-dev", prompt="neon city at dawn", ) # async
Meshive
queue → GPU



Choose where your outputs land
Every finished run shows up in the Artifacts gallery — and is copied to the storage you pick.
See it run before you commit.
Chat with an LLM, generate an image, render a video — right in the Playground.
