llama.cpp

llama.cpp is an open-source inference engine written in C and C++ for running large language models on personal computers.

1 article
Last mentioned

llama.cpp is a project Georgi Gerganov released in 2023. It runs models on CPUs and many kinds of GPUs, loads quantized models in the GGUF file format so they can run with little memory, and is widely used as the foundation of local LLM tools.

This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.

Articles covering this entry

The model is available on Hugging Face in GGUF, the file format used by the llama.cpp inference engine, in two builds:


© 2026 AIPOST. All rights reserved.

AIPOST is an AI publication covering practical AI, AI security, performance, startups, health, ethics and industry news. No account is needed, and our privacy policy explains how we handle personal information.