llama.cpp
llama.cpp is an open-source inference engine written in C and C++ for running large language models on personal computers.
1 article
Last mentioned llama.cpp is a project Georgi Gerganov released in 2023. It runs models on CPUs and many kinds of GPUs, loads quantized models in the GGUF file format so they can run with little memory, and is widely used as the foundation of local LLM tools.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.