Draft llama.cpp PR for DSA (Deepseek Sparse Attention)

#8
by whoisjeremylam - opened

Just in case anyone missed it, I stumbled across this PR.

The implementation seems pretty progressed!

@whoisjeremylam

For sure, I've been following fairydreaming for a while, and have that PR on my "morning briefings" list with 🀞 ...

Its been kinda wild with so many vibe coded forks and patches popping up for all the inference engines to chase turboquant kv-cache, DSV4, MTP, DFlash, etc...

Also wild unsloth released some Qwen3.6 MTP quants despite the mainline PR still being marked draft xD (it works pretty good though in my own testing).

Good seeing ya and cheers!

Sign up or log in to comment