[GitHub Trending] ggml-org/llama.cpp
9.2 relevance
Score Breakdown
technical depth 9
novelty 5
actionability 9
community 10
strategic 8
personal 9
Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.
Essential LLM inference tool, always relevant.
Summary
llama.cpp, the flagship C/C++ inference engine for the ggml library, runs LLMs on Apple Silicon, x86, RISC-V, RISC-V, and NVIDIA GPUs with 1.5-bit to 8-bit quantization and CPU+GPU hybrid inference. Recent additions include Hugging Face cache migration, multimodal support in llama-server, VS Code/Vim FIM plugins, and native GGUF support on Hugging Face Inference Endpoints. It now supports the gpt-oss model in native MXFP4 format, developed in collaboration with NVIDIA.
Author
ggml-org