Transformers now runs llama.cpp GGUF quants natively
On September 22, 2026, Hugging Face added native GGUF support to Transformers, the llama.cpp quantized format behind Ollama, LM Studio, and Jan. If you run local models on Apple Silicon from Python, adopt `from_pretrained` with a GGUF file and drop the homegrown conversions.