serving.
Listed 4Updated Sep 11, 2026
Sorted by Most popular
- vLLMHigh-throughput model server for GPU inference.apache-2.0 · linux, docker58k
- SGLangFast model serving with KV-cache reuse for agent and reasoning workloads.apache-2.0 · linux, macos, +236k
- TensorRT-LLMNVIDIA’s kernel-level serving stack for maximum GPU inference throughput.other · linux, windows, +115k
- XinferenceServe LLMs, embeddings, and media models from one Apache-2.0 platform.apache-2.0 · linux, macos, +29.6k