Instructions to use fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1 # Run inference directly in the terminal: llama cli -hf fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1 # Run inference directly in the terminal: llama cli -hf fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1 # Run inference directly in the terminal: ./llama-cli -hf fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1 # Run inference directly in the terminal: ./build/bin/llama-cli -hf fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1
Use Docker
docker model run hf.co/fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1
- LM Studio
- Jan
- vLLM
How to use fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1
- Ollama
How to use fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF with Ollama:
ollama run hf.co/fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1
- Unsloth Desktop
- Pi
How to use fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF with Docker Model Runner:
docker model run hf.co/fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1
- Lemonade
How to use fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1
Run and chat with the model
lemonade run user.Qwen3.8-27B-Uncensored-GP100-GGUF-Q4_1
List all available models
lemonade list
- Hermes Agent
How to use fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF:Q4_1" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.8-27B Uncensored — Q4_1_G64 4.5 bpw for GP100
Q4_1_G64 · 4.5 bpw · 15.0 GiB · +10% no-spec vs Q4_1 · 128k wall −4.8%
GGUF pack of the JonathanColetti abliteration for 2× Tesla P100 (sm_60).
The served target is type-43 Q4_1_G64: mradermacher i1 Q4_1 body, output head left Q8_0, other Q4_1 weights refit to one (d,m) per 64 weights (4.5 bpw) with the local imatrix. Same HFMA2 decode as Q4_1; fewer DRAM sectors (36 B / 64 w vs 20 B / 32 w). The Q4_1 parent stays in the repo so stock llama.cpp still has a file it can load.
Why this type, and how a decode round actually spends time: DESIGN.md. It walks GP100 vs GP104 (Pascal Tuning Guide + Arafa), the lab DECODE_TIMELINE.md trust tags, why Q4_K is a 2.3× tax at the same 4.5 bpw, why G64 is a storage grouping not a new inner product, and which levers are closed.
The point of G64 is speed, not extra context. On the realistic vibe bar — a 128K session that grows, then compacts — G64 finishes first:
| pack | image | graph-slots | 128k+compact novel wall_sum |
|---|---|---|---|
| G64 | llama-p100:pr-sm60-g64slot (a8c9f3a7b) |
8 | 1532.79 s |
| Q4_1 | same image, DFlash-3 / ngram 3..7 | 1 | 1609.65 s |
G64 wins by 76.86 s (−4.8%). Q4_1 on --graph-slots 8 OOMed CUDA1 (~436 MiB) at the ~110k regrow; the Q4_1 number above is the run that finished.
Do not quote 100 tok/s exact-reuse, or a short HumanEval wall, as the headline. Those are below.
Two paths
Path A — stock llama.cpp. Load the Q4_1 target + Q4_1 DFlash + F16 mmproj. G64 is type 43; stock llama.cpp refuses it cleanly. Path A does not include LoopSpec (ngram-mod,draft-dflash) or the speed table.
Path B — this fork, latest rebase. thefallentree/llama.cpp-gp100 branch cursor/gemma4-mtp-host-6006 @ 958bce0bb (onto ggml-org master 930e2fa59, 2026-09-16). LoopSpec, sm_60 HFMA2, slot-release, empty-src sampling, GDN checkpoint eviction. Do not freeze on a8c9f3a7b / gp100-qwen38-g64slot (2026-09-05 archive). Do not track pr-sm60 HEAD — commit 31 (19d05ddb7) aborts Qwen LoopSpec in cpy.cu.
sm_60 only. P40 / GTX 10-series are sm_61; the HFMA2 path is gated off there.
Files
| File | Size | Mix |
|---|---|---|
Qwen3.8-27B-Uncensored.i1-Q4_1_G64-Q8_0-head.gguf |
14.996 GiB | Path B. 504× Q4_1_G64 + Q8_0 output.weight + Q4_1 token_embd. imatrix=yes |
Qwen3.8-27B-Uncensored.i1-Q4_1-Q8_0-head.gguf |
16.439 GiB | Path A. i1 Q4_1 body, Q8_0 head, Q4_1 token_embd |
Qwen3.8-27B-Uncensored-DFlash2-Q4_1-private-q4-head-exact-embd.gguf |
2.611 GiB | Served draft for both targets. token_embd left Q4_1 |
mmproj-Qwen3.8-27B-Uncensored-F16.gguf |
0.864 GiB | Vision projector, F16 |
Why / how: DESIGN.md. Convert any model + serve recipe: CONVERT.md. Header dump: TENSORS.md. Bench protocol + raw rows: BENCH.md.
The G64 DFlash draft is held. Production uses the Q4_1 draft.
Speed (Path B, 2× Tesla P100)
Same coding fixture, 512 gen, ignore-eos, GPU already warm. Token budgets from POST /tokenize. Prefill and decode are separate. Restart between sizes so novel ngram is clean.
Headline: 128K session + compact
tools/bench_vibe_128k_compact.py: warmup, grow 1k → pin-keep → 16k → 64k → 128k, compact (keep 4k + tail 16k), regrow ~110k, compact again. Score is the sum of novel walls. Prefix cache on (token-ID prefixes, --slot-prompt-similarity 0.05, --ctx-checkpoints 8 --checkpoint-min-step 0). mmproj is loaded (disables ctx_shift).
G64 1532.79 s vs Q4_1 1609.65 s (Q4_1 needed --graph-slots 1).
Exact-reuse decode (G64 constant-w8 vs Q4_1 cap 7)
Same fixture, same g64slot binary, restart between sizes. G64 reuse is faster at 16k and 64k.
| prompt | pack | prefill tok/s | novel decode | exact-reuse | reuse accept |
|---|---|---|---|---|---|
| 1k | G64 g64slot |
181 | 36.0 | 101.7 | 0.989 |
| 1k | Q4_1 (older pr-sm60-ds) |
201 | 46.8 | 60.8 | 0.70 (hash mismatch) |
| 16k | G64 g64slot |
214.5 | 39.3 | 100.6 | 0.996 |
| 16k | Q4_1 g64slot |
193.6 | 53.1 | 91.3 | 0.991 |
| 64k | G64 g64slot |
194.7 | 33.4 | 69.3 | 0.869 |
| 64k | Q4_1 g64slot |
177.8 | 40.1 | 56.9 | 0.849 (hash mismatch) |
No-spec llama-bench (same g64slot binary)
llama-bench -p 512 -n 128, tensor split, FA on, Q8 KV. G64 is faster at every depth:
| depth | Q4_1 tg128 | G64 tg128 | G64 Δ | Q4_1 pp512 | G64 pp512 |
|---|---|---|---|---|---|
| 0 | 32.22 | 35.69 | +10.8% | 207.9 | 228.9 |
| 1024 | 32.57 | 36.10 | +10.8% | 210.5 | 231.3 |
| 16384 | 29.14 | 32.07 | +10.1% | 195.6 | 214.8 |
Honest short-prompt loss
G64 constant-w8 (DFlash-7 / ngram 7) loses the short novel-wall bar. Publish it:
| task | Q4_1 cap 7 | G64 w8 (g64slot) |
|---|---|---|
| coding-1k novel wall | 15.63–16.07 s | 19.88 s |
| HumanEval/0–7 novel wall sum | 71.21 s | 82.92 s |
1k accept 0.286 vs Q4_1 ~0.53 — DFlash-7 rejects most drafts on a short prompt. Do not serve cap 3 / DFlash-2 to hide this; those arms do not beat Q4_1 on the 128k bar.
Wiki PPL (shipped files)
Same g64slot llama-perplexity, eval/wiki.test.raw, 40 chunks, tensor split:
| file | PPL |
|---|---|
| Q4_1 | 6.0409 ± 0.144 |
| G64 | 6.1198 ± 0.147 (+1.31%) |
Under the +2% publish gate. Older G64 emulation (different GGUF) was +1.0% (6.113 vs 6.051).
Run (Path B)
git clone https://github.com/thefallentree/llama.cpp-gp100
cd llama.cpp-gp100 && git checkout cursor/gemma4-mtp-host-6006 # 958bce0bb; not g64slot
# build llama-server with CUDA (sm_60)
./build/bin/llama-server \
--model Qwen3.8-27B-Uncensored.i1-Q4_1_G64-Q8_0-head.gguf \
--model-draft Qwen3.8-27B-Uncensored-DFlash2-Q4_1-private-q4-head-exact-embd.gguf \
--mmproj mmproj-Qwen3.8-27B-Uncensored-F16.gguf \
--jinja \
--spec-type ngram-mod,draft-dflash \
--spec-draft-n-max 7 --spec-draft-n-min 0 \
--spec-ngram-mod-n-min 7 --spec-ngram-mod-n-max 7 \
--backend-sampling --graph-slots 8 \
--ctx-checkpoints 8 --checkpoint-min-step 0 \
--slot-prompt-similarity 0.05 \
--device CUDA0,CUDA1 --split-mode tensor --tensor-split 1,1 \
--device-draft CUDA1 \
--cache-type-k q8_0 --cache-type-v q8_0 \
--cache-type-k-draft q4_0 --cache-type-v-draft q4_0 \
--ctx-size 131072 --flash-attn on
Path A: same command with the Q4_1 target, --spec-draft-n-max 3, --spec-ngram-mod-n-min 3 --spec-ngram-mod-n-max 7, and --graph-slots 1 if you want the 128k+compact path to finish on 2×16 GB.
Context
Default is 131072 + Q8_0 K/V. Native qwen35.context_length is 262144.
G64 saves ~2 GiB of weights. That does not buy a larger Q8 window for multi-turn chat. 163840 and 196608 Q8 decode the first high-n_kv turn and OOM CUDA1 on the cached second turn (leftover FA workspace stays resident). 256K Q8 loads but idle free is ~598 MiB.
Capacity mode is 256K with Q4_0 KV (./run.sh qwen-long in the lab scripts). Q4 KV PPL was +0.67% vs Q8 on an earlier measurement. Do not raise Q8 past 131072 on 2× P100.
Why / how (short)
GP100 is not GP104: two schedulers/SM, no DP4A, HFMA2 is 2× FP32 only as half2, I2F/F2F are quarter-rate, HBM2 ~606 GB/s measured, PCIe only. Q4_K is the same 4.5 bpw as G64 and a ~2.3× decode-round tax here because its scales are extracted every row iteration. Q4_0’s 18-byte block fails 4-byte alignment on half its blocks. G64 is Q4_1’s aligned (d,m) + nibbles stretched to 64 weights (36 B / 64 w). Same A16 HFMA2 inner product; fewer DRAM sectors.
A decode round on this pair is accept (~10 µs) + draft on CUDA1 (GPU0 idle) + issue-bound verify. ~600 launches/round are real; their cost is 0.64 ms. CUDA graphs replay launches and measured flat. The legal speed clock is predicted_per_token_ms, not the old wall-clock slot harness.
Full write-up, with the Pascal Tuning Guide, Arafa, DECODE_TIMELINE.md trust tags, and the lab architecture reviews: DESIGN.md.
To G64 a different model (not this 27B pack): CONVERT.md — scripts/g64/convert_to_g64.py inspect|convert|serve on the fork. Not Qwen-specific. Default mix is G64 body / Q8_0 head / Q4_1 token_embd.
Architecture
qwen35, 27B, 65 blocks (48 gated-attn+SSM, 17 full-attn), d_model 5120, FFN 17408, vocab 248320, context 262144, native MTP on blk.64. Dense.
HFMA2 takes Q4_1 / Q4_1_G64 / Q4_K / Q8_0 with ne[0]%32==0, ne[1]%4==0, weight ne[2]==1. Every quantized 2-D weight here passes.
Related work (not this stack)
Minima (2026-09) reports NVFP4 W4A4 on Qwen3.8-27B with vLLM + FP8 KV on Blackwell SM120. The overlapping claim is that GDN linears are the easy half to quantize — this G64 pack already 4-bit those weights, gates included. It is not NVFP4, not FP8 KV, and P100 cannot run that recipe. Do not copy those numbers onto this card.
Provenance
| Base | Qwen/Qwen3.8-27B (Apache 2.0) |
| Abliteration | JonathanColetti/Qwen3.8-27B-Uncensored |
| i1 Q4_1 body | mradermacher/Qwen3.8-27B-Uncensored-i1-GGUF |
| G64 target | that Q4_1, refit to Q4_1_G64 (imatrix); output.weight stays Q8_0; token_embd stays Q4_1 |
| DFlash2 | incoai/Qwen3.8-27B-DFlash2 / z-lab |
| Runtime | thefallentree/llama.cpp-gp100 cursor/gemma4-mtp-host-6006 @ 958bce0bb |
License
Apache 2.0, inherited from Qwen/Qwen3.8-27B. Local inference and research. Abliterated / uncensored — not a safety-tuned deployment model.
- Downloads last month
- 13,398
4-bit
8-bit
Model tree for fallentree/Qwen3.8-27B-Uncensored-GP100-GGUF
Base model
Qwen/Qwen3.8-27B