Self-hostable software: Inference Serving (7)
7 projects
vLLM
92.3k
A fast, easy-to-use library for LLM inference and serving.
Inference ServingLocal LLMsAI Tooling
2 Containers
LiteLLM
59.3k
LiteLLM is an open source AI gateway and Python SDK for calling 100+ LLM providers through a unified OpenAI-compatible interface.
API ManagementProxy ServersArtificial Intelligence+2
3 ContainersOpen Core
KoboldCpp
11.8k
KoboldCpp is a self-contained AI text-generation runner for GGML and GGUF models, with bundled UI, multiple inference modes, and broad support for image, audio, video, and vision tasks.
Inference ServingLocal LLMsAI Interfaces+2
Single Container
Text Generation Inference
10.9k
A Rust, Python, and gRPC server for deploying and serving large language models with high-performance text generation.
Inference ServingArtificial IntelligenceLocal LLMs+1
2 Containers
GoModel
1.2k
GoModel is a fast, lightweight AI gateway written in Go that provides unified OpenAI-compatible and Anthropic-compatible APIs across many LLM providers.
Artificial IntelligenceGenerative Artificial Intelligence GenaiAPI Management+5
6 Containers
kev
1.1k
Kev is a decision-model server built on a Qwen3 base model that reads a document once and returns calibrated probabilities for typed questions. It provides a TypeSafe-compatible API and a local playground.
Artificial IntelligenceLocal LLMsInference Serving
LLMKube
210
Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and OpenAI-compatible API.
Artificial IntelligenceLocal LLMsInference Serving+2
