Skip to content

OpenAI API.

Local OpenAI-compatible servers, gateways, and fine-tuning stacks for your own models.

Listed 26Catalog Sep 1, 2026

Sorted by Most popular

  1. OllamaLocal model runner with a one-line CLI and a REST API.mit · linux, macos, +2152k
  2. DifyVisual LLM app builder with RAG, agents, and a backend.other · linux, web, +1118k
  3. llama.cppC/C++ runtime that made local GGUF models practical.mit · linux, macos, +187k
  4. nanoGPTKarpathy’s ~300-line GPT trainer — the cleanest way to learn LLM training.mit · linux, macos, +163k
  5. LLaMA-FactoryWeb UI and YAML recipes for fine-tuning many open LLMs.apache-2.0 · linux, windows, +158k
  6. vLLMHigh-throughput model server for GPU inference.apache-2.0 · linux, docker58k
  7. LlamaIndexThe MIT Python framework for building document ingestion and RAG pipelines.mit · linux, macos, +152k
  8. exoLink everyday devices into one cluster and run models too big for any single machine.apache-2.0 · linux, macos47k
  9. UnslothFaster LoRA fine-tuning with lower VRAM on consumer GPUs.apache-2.0 · linux, windows, +145k
  10. DeepSpeedMicrosoft’s Apache-2.0 ZeRO optimizer for multi-GPU, multi-node LLM training.apache-2.0 · linux, macos, +243k
  11. GraphRAGMicrosoft’s MIT library for graph-based RAG over whole document corpora.mit · linux, macos, +136k
  12. SGLangFast model serving with KV-cache reuse for agent and reasoning workloads.apache-2.0 · linux, macos, +236k
  13. LocalAIOpenAI-compatible API that can sit in front of several local backends.mit · linux, macos, +235k
  14. LiteLLMProxy that translates many model APIs into one OpenAI-shaped client.mit · linux, macos, +230k
  15. MLXApple’s framework for running and training models on Mac unified memory.mit · macos28k
  16. llamafileSingle-file LLM runtime from Mozilla. Download, chmod, run.apache-2.0 · linux, macos, +123k
  17. MLC-LLMCompile-and-serve LLMs on Metal, Vulkan, WebGPU, and phones — no CUDA required.apache-2.0 · linux, macos, +423k
  18. PEFTHugging Face’s LoRA/QLoRA library that makes big-model tuning fit one GPU.apache-2.0 · linux, macos, +122k
  19. TRLHugging Face’s Apache-2.0 trainers for SFT, DPO, PPO, and reward modeling.apache-2.0 · linux, macos, +119k
  20. LangfuseTraces, prompts, and evals for LLM apps.mit · linux, web, +116k
  21. TensorRT-LLMNVIDIA’s kernel-level serving stack for maximum GPU inference throughput.other · linux, windows, +115k
  22. shell_gptCLI that runs an LLM as a shell command for scripts, refactors, and answers.mit · linux, macos, +112k
  23. AxolotlYAML-driven fine-tuning toolkit for serious training runs.apache-2.0 · linux, docker11k
  24. PhoenixLocal LLM tracing and evals workspace from Arize — pip install and open.other · linux, macos, +211k
  25. XinferenceServe LLMs, embeddings, and media models from one Apache-2.0 platform.apache-2.0 · linux, macos, +29.6k
  26. torchtunePyTorch-native library for fine-tuning LLMs with recipes.bsd-3-clause · linux7k