OpenAI API.
Local OpenAI-compatible servers, gateways, and fine-tuning stacks for your own models.
Listed 26Catalog Sep 1, 2026
Sorted by Most popular
- OllamaLocal model runner with a one-line CLI and a REST API.mit · linux, macos, +2152k
- DifyVisual LLM app builder with RAG, agents, and a backend.other · linux, web, +1118k
- llama.cppC/C++ runtime that made local GGUF models practical.mit · linux, macos, +187k
- nanoGPTKarpathy’s ~300-line GPT trainer — the cleanest way to learn LLM training.mit · linux, macos, +163k
- LLaMA-FactoryWeb UI and YAML recipes for fine-tuning many open LLMs.apache-2.0 · linux, windows, +158k
- vLLMHigh-throughput model server for GPU inference.apache-2.0 · linux, docker58k
- LlamaIndexThe MIT Python framework for building document ingestion and RAG pipelines.mit · linux, macos, +152k
- exoLink everyday devices into one cluster and run models too big for any single machine.apache-2.0 · linux, macos47k
- UnslothFaster LoRA fine-tuning with lower VRAM on consumer GPUs.apache-2.0 · linux, windows, +145k
- DeepSpeedMicrosoft’s Apache-2.0 ZeRO optimizer for multi-GPU, multi-node LLM training.apache-2.0 · linux, macos, +243k
- GraphRAGMicrosoft’s MIT library for graph-based RAG over whole document corpora.mit · linux, macos, +136k
- SGLangFast model serving with KV-cache reuse for agent and reasoning workloads.apache-2.0 · linux, macos, +236k
- LocalAIOpenAI-compatible API that can sit in front of several local backends.mit · linux, macos, +235k
- LiteLLMProxy that translates many model APIs into one OpenAI-shaped client.mit · linux, macos, +230k
- MLXApple’s framework for running and training models on Mac unified memory.mit · macos28k
- llamafileSingle-file LLM runtime from Mozilla. Download, chmod, run.apache-2.0 · linux, macos, +123k
- MLC-LLMCompile-and-serve LLMs on Metal, Vulkan, WebGPU, and phones — no CUDA required.apache-2.0 · linux, macos, +423k
- PEFTHugging Face’s LoRA/QLoRA library that makes big-model tuning fit one GPU.apache-2.0 · linux, macos, +122k
- TRLHugging Face’s Apache-2.0 trainers for SFT, DPO, PPO, and reward modeling.apache-2.0 · linux, macos, +119k
- LangfuseTraces, prompts, and evals for LLM apps.mit · linux, web, +116k
- TensorRT-LLMNVIDIA’s kernel-level serving stack for maximum GPU inference throughput.other · linux, windows, +115k
- shell_gptCLI that runs an LLM as a shell command for scripts, refactors, and answers.mit · linux, macos, +112k
- AxolotlYAML-driven fine-tuning toolkit for serious training runs.apache-2.0 · linux, docker11k
- PhoenixLocal LLM tracing and evals workspace from Arize — pip install and open.other · linux, macos, +211k
- XinferenceServe LLMs, embeddings, and media models from one Apache-2.0 platform.apache-2.0 · linux, macos, +29.6k
- torchtunePyTorch-native library for fine-tuning LLMs with recipes.bsd-3-clause · linux7k