Qwen3.8-27B Uncensored — Q4_1_G64 4.5 bpw for GP100

Q4_1_G64 · 4.5 bpw · 15.0 GiB · +10% no-spec vs Q4_1 · 128k wall −4.8%

GGUF pack of the JonathanColetti abliteration for 2× Tesla P100 (sm_60).

The served target is type-43 Q4_1_G64: mradermacher i1 Q4_1 body, output head left Q8_0, other Q4_1 weights refit to one (d,m) per 64 weights (4.5 bpw) with the local imatrix. Same HFMA2 decode as Q4_1; fewer DRAM sectors (36 B / 64 w vs 20 B / 32 w). The Q4_1 parent stays in the repo so stock llama.cpp still has a file it can load.

Why this type, and how a decode round actually spends time: DESIGN.md. It walks GP100 vs GP104 (Pascal Tuning Guide + Arafa), the lab DECODE_TIMELINE.md trust tags, why Q4_K is a 2.3× tax at the same 4.5 bpw, why G64 is a storage grouping not a new inner product, and which levers are closed.

The point of G64 is speed, not extra context. On the realistic vibe bar — a 128K session that grows, then compacts — G64 finishes first:

pack image graph-slots 128k+compact novel wall_sum
G64 llama-p100:pr-sm60-g64slot (a8c9f3a7b) 8 1532.79 s
Q4_1 same image, DFlash-3 / ngram 3..7 1 1609.65 s

G64 wins by 76.86 s (−4.8%). Q4_1 on --graph-slots 8 OOMed CUDA1 (~436 MiB) at the ~110k regrow; the Q4_1 number above is the run that finished.

Do not quote 100 tok/s exact-reuse, or a short HumanEval wall, as the headline. Those are below.

Two paths

Path A — stock llama.cpp. Load the Q4_1 target + Q4_1 DFlash + F16 mmproj. G64 is type 43; stock llama.cpp refuses it cleanly. Path A does not include LoopSpec (ngram-mod,draft-dflash) or the speed table.

Path B — this fork, latest rebase. thefallentree/llama.cpp-gp100 branch cursor/gemma4-mtp-host-6006 @ 958bce0bb (onto ggml-org master 930e2fa59, 2026-09-16). LoopSpec, sm_60 HFMA2, slot-release, empty-src sampling, GDN checkpoint eviction. Do not freeze on a8c9f3a7b / gp100-qwen38-g64slot (2026-09-05 archive). Do not track pr-sm60 HEAD — commit 31 (19d05ddb7) aborts Qwen LoopSpec in cpy.cu.

sm_60 only. P40 / GTX 10-series are sm_61; the HFMA2 path is gated off there.

Files

File Size Mix
Qwen3.8-27B-Uncensored.i1-Q4_1_G64-Q8_0-head.gguf 14.996 GiB Path B. 504× Q4_1_G64 + Q8_0 output.weight + Q4_1 token_embd. imatrix=yes
Qwen3.8-27B-Uncensored.i1-Q4_1-Q8_0-head.gguf 16.439 GiB Path A. i1 Q4_1 body, Q8_0 head, Q4_1 token_embd
Qwen3.8-27B-Uncensored-DFlash2-Q4_1-private-q4-head-exact-embd.gguf 2.611 GiB Served draft for both targets. token_embd left Q4_1
mmproj-Qwen3.8-27B-Uncensored-F16.gguf 0.864 GiB Vision projector, F16

Why / how: DESIGN.md. Convert any model + serve recipe: CONVERT.md. Header dump: TENSORS.md. Bench protocol + raw rows: BENCH.md.

The G64 DFlash draft is held. Production uses the Q4_1 draft.

Speed (Path B, 2× Tesla P100)

Same coding fixture, 512 gen, ignore-eos, GPU already warm. Token budgets from POST /tokenize. Prefill and decode are separate. Restart between sizes so novel ngram is clean.

Headline: 128K session + compact

tools/bench_vibe_128k_compact.py: warmup, grow 1k → pin-keep → 16k → 64k → 128k, compact (keep 4k + tail 16k), regrow ~110k, compact again. Score is the sum of novel walls. Prefix cache on (token-ID prefixes, --slot-prompt-similarity 0.05, --ctx-checkpoints 8 --checkpoint-min-step 0). mmproj is loaded (disables ctx_shift).

G64 1532.79 s vs Q4_1 1609.65 s (Q4_1 needed --graph-slots 1).

Exact-reuse decode (G64 constant-w8 vs Q4_1 cap 7)

Same fixture, same g64slot binary, restart between sizes. G64 reuse is faster at 16k and 64k.

prompt pack prefill tok/s novel decode exact-reuse reuse accept
1k G64 g64slot 181 36.0 101.7 0.989
1k Q4_1 (older pr-sm60-ds) 201 46.8 60.8 0.70 (hash mismatch)
16k G64 g64slot 214.5 39.3 100.6 0.996
16k Q4_1 g64slot 193.6 53.1 91.3 0.991
64k G64 g64slot 194.7 33.4 69.3 0.869
64k Q4_1 g64slot 177.8 40.1 56.9 0.849 (hash mismatch)

No-spec llama-bench (same g64slot binary)

llama-bench -p 512 -n 128, tensor split, FA on, Q8 KV. G64 is faster at every depth:

depth Q4_1 tg128 G64 tg128 G64 Δ Q4_1 pp512 G64 pp512
0 32.22 35.69 +10.8% 207.9 228.9
1024 32.57 36.10 +10.8% 210.5 231.3
16384 29.14 32.07 +10.1% 195.6 214.8

Honest short-prompt loss

G64 constant-w8 (DFlash-7 / ngram 7) loses the short novel-wall bar. Publish it:

task Q4_1 cap 7 G64 w8 (g64slot)
coding-1k novel wall 15.63–16.07 s 19.88 s
HumanEval/0–7 novel wall sum 71.21 s 82.92 s

1k accept 0.286 vs Q4_1 ~0.53 — DFlash-7 rejects most drafts on a short prompt. Do not serve cap 3 / DFlash-2 to hide this; those arms do not beat Q4_1 on the 128k bar.

Wiki PPL (shipped files)

Same g64slot llama-perplexity, eval/wiki.test.raw, 40 chunks, tensor split:

file PPL
Q4_1 6.0409 ± 0.144
G64 6.1198 ± 0.147 (+1.31%)

Under the +2% publish gate. Older G64 emulation (different GGUF) was +1.0% (6.113 vs 6.051).

Run (Path B)

git clone https://github.com/thefallentree/llama.cpp-gp100
cd llama.cpp-gp100 && git checkout cursor/gemma4-mtp-host-6006   # 958bce0bb; not g64slot
# build llama-server with CUDA (sm_60)

./build/bin/llama-server \
  --model Qwen3.8-27B-Uncensored.i1-Q4_1_G64-Q8_0-head.gguf \
  --model-draft Qwen3.8-27B-Uncensored-DFlash2-Q4_1-private-q4-head-exact-embd.gguf \
  --mmproj mmproj-Qwen3.8-27B-Uncensored-F16.gguf \
  --jinja \
  --spec-type ngram-mod,draft-dflash \
  --spec-draft-n-max 7 --spec-draft-n-min 0 \
  --spec-ngram-mod-n-min 7 --spec-ngram-mod-n-max 7 \
  --backend-sampling --graph-slots 8 \
  --ctx-checkpoints 8 --checkpoint-min-step 0 \
  --slot-prompt-similarity 0.05 \
  --device CUDA0,CUDA1 --split-mode tensor --tensor-split 1,1 \
  --device-draft CUDA1 \
  --cache-type-k q8_0 --cache-type-v q8_0 \
  --cache-type-k-draft q4_0 --cache-type-v-draft q4_0 \
  --ctx-size 131072 --flash-attn on

Path A: same command with the Q4_1 target, --spec-draft-n-max 3, --spec-ngram-mod-n-min 3 --spec-ngram-mod-n-max 7, and --graph-slots 1 if you want the 128k+compact path to finish on 2×16 GB.

Context

Default is 131072 + Q8_0 K/V. Native qwen35.context_length is 262144.

G64 saves ~2 GiB of weights. That does not buy a larger Q8 window for multi-turn chat. 163840 and 196608 Q8 decode the first high-n_kv turn and OOM CUDA1 on the cached second turn (leftover FA workspace stays resident). 256K Q8 loads but idle free is ~598 MiB.

Capacity mode is 256K with Q4_0 KV (./run.sh qwen-long in the lab scripts). Q4 KV PPL was +0.67% vs Q8 on an earlier measurement. Do not raise Q8 past 131072 on 2× P100.

Why / how (short)

GP100 is not GP104: two schedulers/SM, no DP4A, HFMA2 is 2× FP32 only as half2, I2F/F2F are quarter-rate, HBM2 ~606 GB/s measured, PCIe only. Q4_K is the same 4.5 bpw as G64 and a ~2.3× decode-round tax here because its scales are extracted every row iteration. Q4_0’s 18-byte block fails 4-byte alignment on half its blocks. G64 is Q4_1’s aligned (d,m) + nibbles stretched to 64 weights (36 B / 64 w). Same A16 HFMA2 inner product; fewer DRAM sectors.

A decode round on this pair is accept (~10 µs) + draft on CUDA1 (GPU0 idle) + issue-bound verify. ~600 launches/round are real; their cost is 0.64 ms. CUDA graphs replay launches and measured flat. The legal speed clock is predicted_per_token_ms, not the old wall-clock slot harness.

Full write-up, with the Pascal Tuning Guide, Arafa, DECODE_TIMELINE.md trust tags, and the lab architecture reviews: DESIGN.md.

To G64 a different model (not this 27B pack): CONVERT.md — scripts/g64/convert_to_g64.py inspect|convert|serve on the fork. Not Qwen-specific. Default mix is G64 body / Q8_0 head / Q4_1 token_embd.

Architecture

qwen35, 27B, 65 blocks (48 gated-attn+SSM, 17 full-attn), d_model 5120, FFN 17408, vocab 248320, context 262144, native MTP on blk.64. Dense.

HFMA2 takes Q4_1 / Q4_1_G64 / Q4_K / Q8_0 with ne[0]%32==0, ne[1]%4==0, weight ne[2]==1. Every quantized 2-D weight here passes.

Related work (not this stack)

Minima (2026-09) reports NVFP4 W4A4 on Qwen3.8-27B with vLLM + FP8 KV on Blackwell SM120. The overlapping claim is that GDN linears are the easy half to quantize — this G64 pack already 4-bit those weights, gates included. It is not NVFP4, not FP8 KV, and P100 cannot run that recipe. Do not copy those numbers onto this card.

Provenance

Base Qwen/Qwen3.8-27B (Apache 2.0)
Abliteration JonathanColetti/Qwen3.8-27B-Uncensored
i1 Q4_1 body mradermacher/Qwen3.8-27B-Uncensored-i1-GGUF
G64 target that Q4_1, refit to Q4_1_G64 (imatrix); output.weight stays Q8_0; token_embd stays Q4_1
DFlash2 incoai/Qwen3.8-27B-DFlash2 / z-lab
Runtime thefallentree/llama.cpp-gp100 cursor/gemma4-mtp-host-6006 @ 958bce0bb

License

Apache 2.0, inherited from Qwen/Qwen3.8-27B. Local inference and research. Abliterated / uncensored — not a safety-tuned deployment model.

Downloads last month
13,398
GGUF
Model size
4B params
Architecture
dflash
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(29)
this model

Paper for fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF