Transformers now runs llama.cpp quants
Synopsis
This Hugging Face post describes how transformers can now load GGUF quantized checkpoints directly through from_pretrained with a gguf_file argument: on Apple Silicon it reuses ggml Metal kernels distributed via the kernels library (packed-quantized weight matmul, fused normalization, flash attention, gated delta net, plus a home-grown topk for MoE routing) and trims generate overhead by dropping the redundant attention mask early and deferring the stopping check asynchronously; on a MacBook Pro M2 Max the authors measure generation speed close to llama.cpp across three checkpoints, report file-size tradeoffs for Q4_K_M, Q5_K_M and Q6_K, and state that packed inference is MPS-only with architecture coverage for Qwen3.5 dense and MoE (including compatible Qwen3.
Interpretation
GGUF checkpoints become a first-class load path in transformers: pass the Hub model_id and filename as gguf_file to from_pretrained, and everything afterwards uses the standard API (apply_chat_template, generate, or transformers serve exposing an OpenAI-compatible endpoint). GGUF previously belonged to a separate inference runtime; now the same quantized checkpoint can be loaded and generated from inside a PyTorch process. When weights stay packed, transformers automatically loads the ggml/Metal kernels and uses ggml-org/ggml-attn as the attention implementation; if that kernel cannot be fetched it falls back to "sdpa" with a warning, and "sdpa" can also be forced explicitly. Grounded in official loading examples and configuration notes, i.e. implementation and usage description rather than a controlled experiment.
Reusing the kernels together with two changes to the generate loop brings local generation speed close to llama.cpp. On the kernel side: reading packed quantized weights directly (avoiding expansion of the whole weight matrix before each decode), fused normalization, Metal flash attention, gated delta net, and a topk kernel aimed at the separate MoE routing bottleneck. On the loop side: #48814 drops an all-ones padding mask early for unpadded decoder-only inputs, and #47975 makes the stopping decision asynchronous and consumes it on the following step; both changes improve all transformers models, not only GGUF runs. Measured on a MacBook Pro M2 Max (32 GB unified memory, macOS 26.6, PyTorch 2.12.1, kernels 0.17.0) across three checkpoints; the llama.cpp column is llama-bench tg128 (128 decoded tokens, averaged over three repetitions, prompt processing excluded), while the transformers column is three warmed runs best-of-three for the same 128 tokens including prefill, so conditions are not identical. The post states that "Transformers is close to llama.cpp across all three checkpoints".
Quantization tiers come with a concrete size-versus-precision reference: Unsloth's Qwen3.5-4B goes from 8.42 GB at BF16 to 2.74 GB at Q4_K_M (Q6_K 3.53 GB, Q5_K_M 3.14 GB). Variants such as Q4_K_M mix tensor precisions, using mostly 4-bit weights while keeping sensitive tensors at higher precision; the post suggests starting with Q4_K_M and trying Q5_K_M or Q6_K when more memory is available. File sizes are checkable format facts; the quality tradeoff is explicitly described as depending on the model and the task, to be evaluated on the work you actually want the model to do.
The integration brings GGUF checkpoints into a PyTorch development workflow and sketches a path to more architectures and modalities: inspect intermediate activations, modify a forward pass, evaluate quantized checkpoints, validate GGUF conversions, use custom logits processors and stopping criteria, or dequantize with GgufConfig(dequantize=True) and continue with a standard training workflow. A kernel operates on tensors and does not require the whole model to come from a GGUF file, so the same attention, normalization and matmul kernels are described as reusable in other transformers models and loading workflows, potentially by computer vision, audio and multimodal models; this is especially useful for new architectures, research models and custom variants that may never get a dedicated llama.cpp implementation. A design and roadmap description; the post states each architecture still needs integration and validation and that the initial examples cover text generation.
Perspective
The work targets a single interactive conversation on Apple Silicon: the packed inference path is MPS-only for now, while GGUF import through dequantization remains a separate option, and supporting the file format does not imply packed kernels on every device; architecture coverage is Qwen3.5 dense and MoE, including compatible Qwen3.8 checkpoints. It lets developers working in Python and PyTorch load, debug, evaluate and convert-check quantized checkpoints in familiar tooling, and opens a route to bringing ggml kernels to architectures transformers already implements, potentially including reusable vision, audio and multimodal operators; for maximum local inference efficiency the post still points to llama.cpp.
The text describes the speed comparison in prose without the numeric values behind the chart, so horizontal verification requires returning to the original figure; the two protocols differ (one includes prefill, the other reports decode-only tg128 averaged over three repetitions), meaning "close" should be read as an overall impression on one machine. The benchmark script's note that "back-to-back runs decay by 10% or more" indicates sensitivity to thermal state, so other devices or cooling conditions warrant fresh measurement. Quality loss at each quantization tier depends on the model and the task, and the post advises evaluating on your own workload, so per-task differences remain for users to test; padding and batching paths and generate_batch on MPS are listed as follow-up work and are scope statements. In addition, this parse begins at "from_pretrained, and start generating on your own machine.", so the opening of the post is not included and the framing of intended audience and prerequisites may be slightly incomplete.
