selfhostedworld.com logoselfhostedworld.com

Try describing what you need:

Self-hostable software: Inference Serving (7)

7 projects
vLLM logo

vLLM

92.3k
A fast, easy-to-use library for LLM inference and serving.
Inference ServingLocal LLMsAI Tooling
2 Containers
Details
LiteLLM logo

LiteLLM

59.3k
LiteLLM is an open source AI gateway and Python SDK for calling 100+ LLM providers through a unified OpenAI-compatible interface.
API ManagementProxy ServersArtificial Intelligence+2
3 ContainersOpen Core
Details
KoboldCpp logo

KoboldCpp

11.8k
KoboldCpp is a self-contained AI text-generation runner for GGML and GGUF models, with bundled UI, multiple inference modes, and broad support for image, audio, video, and vision tasks.
Inference ServingLocal LLMsAI Interfaces+2
Single Container
Details
Text Generation Inference logo

Text Generation Inference

10.9k
A Rust, Python, and gRPC server for deploying and serving large language models with high-performance text generation.
Inference ServingArtificial IntelligenceLocal LLMs+1
2 Containers
Details
GoModel logo

GoModel

1.2k
GoModel is a fast, lightweight AI gateway written in Go that provides unified OpenAI-compatible and Anthropic-compatible APIs across many LLM providers.
Artificial IntelligenceGenerative Artificial Intelligence GenaiAPI Management+5
6 Containers
Details
kev logo

kev

1.1k
Kev is a decision-model server built on a Qwen3 base model that reads a document once and returns calibrated probabilities for typed questions. It provides a TypeSafe-compatible API and a local playground.
Artificial IntelligenceLocal LLMsInference Serving
Details
LLMKube logo

LLMKube

210
Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and OpenAI-compatible API.
Artificial IntelligenceLocal LLMsInference Serving+2
Details