LLM
★ 12.6kAbout
LLM is a command-line tool and Python library for working with large language models. It is designed for developers and power users who want to run prompts, chat interactively, extract structured...
Community Ratings
No ratings yetNew this week
Read recap →Weekly Recap — Oct 2, 2026 – Oct 9, 2026Oct 9, 2026
Features
- command-line prompts
- SQLite prompt/response logging
- embeddings generation
- structured content extraction
- tool execution
- interactive chat
- plugin-based model support
- local model support
Details
- Last Updated
- Oct 10, 2026
- Created
- Apr 1, 2023
- Install Methods
- package-managersource
- Requirements
- PythonpipHomebrewpipxuvOpenAI API key for OpenAI modelsGemini API key for Gemini modelsAnthropic API key for Anthropic modelsOllama for local models
- Backup & Export
- file-backup
- Privacy & Independence
- Cloud: Optional
- Deployment
Deployment: no container deployment found
Track your self-hosted stack
Bookmark software to try, rate tools you've used, and keep your collection in one place.
Metadata extracted from README on Mar 25, 2026
Related Software
Ollama
182.7k
Ollama is a tool for running open models locally and exposing them through a REST API for apps, agents, and integrations.
Artificial IntelligenceLocal LLMsGenerative AI+3
llamafile
26.2k
llamafile lets users distribute and run LLMs as a single-file executable with no installation. It also includes whisperfile for local speech-to-text transcription and translation.
Artificial IntelligenceLocal LLMsAI Tooling+2
LLM Harbor
3.2k
Harbor is a CLI and companion app for spinning up and managing a local LLM stack with pre-wired services like Open WebUI, llama.cpp, Ollama, vLLM, SearXNG, Speaches, and ComfyUI.
Artificial IntelligenceLocal LLMsAI Tooling+5
twinny
3.7k
Twinny is a free AI extension for Visual Studio Code that provides AI-assisted coding features.
Artificial IntelligenceAI Coding AssistantPermissive License+2
ollaya
1.3k
Ollaya runs open decision models locally and provides Ollama-like CLI, daemon, and API access for model serving.
Artificial IntelligenceLocal LLMsInference Serving+2
LLMKube
231
Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and OpenAI-compatible API.
Artificial IntelligenceLocal LLMsInference Serving+2
Text Generation Inference
10.9k
A Rust, Python, and gRPC server for deploying and serving large language models with high-performance text generation.
Inference ServingArtificial IntelligenceLocal LLMs+1
2 Containers
vLLM
93.5k
A fast, easy-to-use library for LLM inference and serving.
Inference ServingLocal LLMsAI Tooling
2 Containers
AnythingLLM
66.9k
An all-in-one AI application for chatting with documents, building AI agents, and running local or cloud LLM workflows.
Artificial IntelligenceAutomation ToolsSelf Hosting Solutions+7
Single Container
SillyTavern
34.3k
SillyTavern is a locally installed interface for interacting with text generation LLMs, image generation engines, and TTS voice models.
Artificial IntelligenceAI InterfacesCopyleft License+3
Single Container
LocalAI
49.5k
LocalAI is an open-source AI engine that runs LLMs and other multimodal models locally on a wide range of hardware, with OpenAI-compatible APIs and on-demand backends.
Generative AIPermissive LicenseArtificial Intelligence+4
Single Container
private-gpt
57.6k
PrivateGPT is an open-source API layer that turns local models into production AI applications.
Artificial IntelligenceAI RAG SystemsPermissive License+2
