llama.cpp

Run language and vision-language models across local and cloud hardware.

What is llama.cpp?

llama.cpp helps users run language and vision-language models across local and cloud hardware. The product supports c and C++ inference engine and model server and command-line tools. Starting points include deploy a local model and compare inference performance on available hardware. Check the documentation, license, supported models, hardware requirements, and any separate API or hosting costs. Test on a small representative project before integrating it into an existing system.

Source: official product website. Reviewed .

What it helps you do

  • C and C++ inference engine
  • Model server and command-line tools

Where to start

  1. Deploy a local model
  2. Compare inference performance on available hardware

Before you choose

Check the documentation, license, supported models, hardware requirements, and any separate API or hosting costs. Test on a small representative project before integrating it into an existing system.

This profile is based on the provider's published information. We have not independently tested every feature.

Pricing

Check provider. Source code is available. Check the project license and documentation; model APIs, compute, hosting, and commercial services may have separate costs.

Visit llama.cpp

More tools to consider

Explore llama.cpp alternatives