llama.cpp
Run language and vision-language models across local and cloud hardware.
What is llama.cpp?
llama.cpp helps users run language and vision-language models across local and cloud hardware. The product supports c and C++ inference engine and model server and command-line tools. Starting points include deploy a local model and compare inference performance on available hardware. Check the documentation, license, supported models, hardware requirements, and any separate API or hosting costs. Test on a small representative project before integrating it into an existing system.
Source: official product website. Reviewed .
What it helps you do
- C and C++ inference engine
- Model server and command-line tools
Where to start
- Deploy a local model
- Compare inference performance on available hardware
Before you choose
Check the documentation, license, supported models, hardware requirements, and any separate API or hosting costs. Test on a small representative project before integrating it into an existing system.
This profile is based on the provider's published information. We have not independently tested every feature.
Pricing
Check provider. Source code is available. Check the project license and documentation; model APIs, compute, hosting, and commercial services may have separate costs.
Visit llama.cppMore tools to consider
- Mistral Vibe (formerly Le Chat): Mistral AI's conversational chat interface for fast, multilingual AI interactions.
- Aleph Alpha: Build specialized language-model applications for organizations.
- Braintrust: AI evaluation and prompt management platform
- KoboldCpp: Run language models locally through a lightweight inference application.
- llama.cpp: Run language and vision-language models across local and cloud hardware.
- Llamafile: Package and run language models as portable local executables.
- LlamaIndex: Data framework for building LLM applications with custom knowledge.
- Portkey AI: AI gateway for managing LLM reliability, routing, and observability.
- Semantic Kernel: Microsoft open-source SDK for integrating LLMs into applications.
- Vellum AI: AI development platform for building, testing, and deploying LLM workflows.
- Beam Cloud
- Letta
- Liquid AI
- MindsDB
- Nomic Atlas
- Patronus AI
- Pezzo
- Pixtral 12B
- Poolside
- Predibase
- PromptLayer
- Qwen2.5-VL
- Ragie
- Reducto
- Sarvam AI
- Tensorlake
- TrainMyAI: Platform for training custom AI models on your proprietary data.
- Trieve
- turbopuffer
- Unstructured