Serve LLMs, embeddings, and media models from one Apache-2.0 platform.
One-box alternative to juggling Ollama plus separate embedding servers.
Xinference launches, schedules, and serves many models from one platform: LLMs, embeddings, rerankers, image and audio models behind an OpenAI-shaped API with a built-in chat UI. Apache-2.0 and very active.
multi-model serving · built-in UI · GPU scheduling · OpenAI-compatible API