FR
live
tag

#transformers

Transformers now runs llama.cpp GGUF quants natively

On September 22, 2026, Hugging Face added native GGUF support to Transformers, the llama.cpp quantized format behind Ollama, LM Studio, and Jan. If you run local models on Apple Silicon from Python, adopt `from_pretrained` with a GGUF file and drop the homegrown conversions.

NeoMME Fuses Text and Images in a Single Bidirectional Transformer

Hcompany ships NeoMME, a 260M–800M multilingual multimodal encoder that processes text and images in one Transformer with no separate vision tower. For visual document retrieval, its Retriever variant reaches the ViDoRe v3 Pareto frontier with a 255× smaller index.

Type at least two characters.

↑ ↓ navigate ↵ open esc dismiss