alibaba/MNN · GitHub
vllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for LLMs
sgl-project/sglang: SGLang is a fast serving framework for large language models and vision language models.
AMD-AIG-AIMA/Instella: Fully Open Language Models with Stellar Performance
The Local AI Playground
fullmoon: local intelligence
Source-code: https://github.com/mainframecomputer/fullmoon-ios
GitHub - chatbox/chatbox: User-friendly Desktop Client App for AI Models/LLMs (GPT, Claude, Gemini, Ollama...)
Together AI – Fast Inference, Fine-Tuning & Training
Pinokio - Localhost Platform for Humans and AI
Source-code: https://github.com/pinokiocomputer/pinokio
Stability-AI/StableSwarmUI: StableSwarmUI, A Modular Stable Diffusion Web-User-Interface, with an emphasis on making powertools easily accessible, high performance, and extensibility.
microsoft/semantic-kernel: Integrate cutting-edge LLM technology quickly and easily into your apps
Jeffser/Alpaca · GitHub
BerriAI/litellm · GitHub
LLMLingua Series | Effectively Deliver Information to LLMs via Prompt Compression
Source-code: https://github.com/microsoft/LLMLingua
BentoML: Build, Ship, Scale AI Applications
Source-code: https://github.com/bentoml/OpenLLM
cumulo-autumn/StreamDiffusion · GitHub
Amazon Q - AWS
LangChain
Llama.cpp - Run LLM Inference in C/C++
Source-code: https://github.com/ggml-org/llama.cpp
Forefront: Powerful Language Models A Click Away