Webinar, 15 Oct: Who controls your enterprise AI? Join us live →

Sovereign inference.
Run any model. One API.

The AI inference platform for 300+ open models, on Australian infrastructure or yours. ~70% lower cost than closed-source APIs.*

*Compared with closed-source API list prices. Savings vary by workload.

POST  api.xinference.co/v1/chat/completions
client = OpenAI(
    base_url="https://api.xinference.co/v1",
    api_key="XINFERENCE_API_KEY",
)

# two lines. that's the migration.

Trusted in production by

Siemens Yum! Brands AIA TFC OpticalComms XW Bank
300+
Open models
100+
Teams trust Xinference
9.4K
GitHub stars

Why teams switch

Most teams start on the Model API and move to dedicated compute as usage grows.

See customer results→

No lock-in

300+ open models behind one OpenAI-compatible API. Swap models, hardware or clouds without rebuilding your applications.

Cost certainty

Predictable pricing per token, with the option to scale to a set cost per GPU hour as your workloads demand. We measure your savings with you before you commit.

Private by design

Runs in Australia or on infrastructure you control. No training on your data, no retention by default.

One control plane

The observability and governance enterprises need, built in: per-request logs, live TTFT and TPOT monitoring, role-based access, audit logs and SSO.

In production.

Quick-service restaurants

One platform behind 50+ AI use cases.

Yum! runs 50+ AI use cases across KFC, Pizza Hut and Taco Bell, at 1.52M requests a day across dual data centres.

35–45%
lower infra cost
Read the case study →
Financial services

6,000-user AI on one pooled cluster.

360,000 requests per day on 72 pooled GPUs, serving 6,000 users.

40–50%
lower infra cost
Read the case study →
Insurance

A regulated AI platform, compliant by design.

500K+ daily requests, replacing a self-managed vLLM and Kubernetes stack.

~30%
lower AI cost
Read the case study →

What you can build on Xinference.

Most production workloads already run on open models, with the same or similar results as closed-source APIs.

Enterprise RAG

Dense retrieval over your own documents, with embedding, rerank and the LLM on one runtime.

Conversational AI

Customer assistants and internal helpdesks, sharing one pooled cluster with the rest of your workloads.

Agents and function calling

Multi-step reasoning and tool use, on frontier open models such as Kimi 3 and DeepSeek 4.1 Flash. You can also build them with Xagent, which runs agents on your own models.

Coding assistance

IDE copilots and code generation, on long-context models served with streaming.

Summarisation and extraction

Document parsing, summarisation and classification, including high-volume batch pipelines.

Speech and multimodal

Real-time transcription, low-latency voices for calls and agents, and image generation with FLUX, SDXL and SD3.

Run any model. Keep the inference sovereign.