Local LLM inference optimisations: from attention mechanisms to predictive decoding and software-model-hardware implementations.
On Sunday I write about speculative decoding, and immediately we get Qwen3.6 with MTP and support for llama.cpp: https://huggingface.co/unsloth/Qwen3.6-27B-MTP-GGUF
I just tested it and it looks really promising. I'll report back with some numbers.
On Sunday I write about speculative decoding, and immediately we get Qwen3.6 with MTP and support for llama.cpp: https://huggingface.co/unsloth/Qwen3.6-27B-MTP-GGUF
I just tested it and it looks really promising. I'll report back with some numbers.