Self-hostable software: Inference Serving (5)
5 projects
vLLM
88.4k
A fast, easy-to-use library for LLM inference and serving.
Inference ServingLocal LLMsAI Tooling
2 Containers
LiteLLM
55.8k
LiteLLM is an open source AI gateway and Python SDK for calling 100+ LLM providers through a unified OpenAI-compatible interface.
API ManagementProxy ServersArtificial Intelligence+2
3 ContainersOpen Core
KoboldCpp
11.4k
KoboldCpp is a self-contained AI text-generation runner for GGML and GGUF models, with bundled UI, multiple inference modes, and broad support for image, audio, video, and vision tasks.
Inference ServingLocal LLMsAI Interfaces+2
Single Container
Text Generation Inference
10.9k
A Rust, Python, and gRPC server for deploying and serving large language models with high-performance text generation.
Inference ServingArtificial IntelligenceLocal LLMs+1
2 Containers
LLMKube
186
Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and OpenAI-compatible API.
Artificial IntelligenceLocal LLMsInference Serving+1
