Qwen3-4B β€” ConvRot W4A4 (ComfyUI text encoder)

4-bit weight / 4-bit activation quantization of the Qwen3-4B text encoder as ComfyUI loads it (the Lumina2 / Z-Image conditioning path), in ComfyUI's native convrot_w4a4 format.

7.49 GiB β†’ 2.42 GiB, 3.09x lighter. 252 of 398 tensors quantized.

Produced on a bench that tries to refute its own results. Every number below was executed, not inferred, and each one carries the condition it was measured under. What was not measured is listed at the bottom, on purpose.

Method, tools and the full measurement log: https://github.com/JoaoZaokk/comfy-quant-bench


Read this before you download

Out of the box in stock ComfyUI, this file saves memory and saves no time. That is not a defect of the quantization β€” it is how ComfyUI treats every text encoder, whatever the format.

Two independent locks force dequantized math on the CLIP path:

lock where it comes from
full_precision_mm=True comfy/sd1_clip.py:114, hardcoded for every text encoder
comfy_force_cast_weights=True comfy/sd.py:269 β†’ set_model_compute_dtype(torch.float32) for every CLIP object

Counted during a real encode on a comparable public W4A4 encoder: 0 calls to the 4-bit kernel, 350 calls to QuantizedTensor.dequantize. The weight stays 4-bit in VRAM; the GEMM runs in BF16. Releasing only one of the two locks changes nothing β€” both have to go.

So the honest pitch for this file is: it fits where the BF16 original does not. If your card cannot hold 7.5 GiB of text encoder alongside a diffusion model, this is what you want. If you came for speed, read the next section first.


What it costs in fidelity

Conditioning tensors compared against the BF16 original, same prompts, same tokenizer, on an RTX 3090:

prompt shape rel-RMSE cosine
a red apple on a weathered wooden table, soft window light [1, 21, 2560] 0.1442 0.98957
portrait of an elderly fisherman, weathered face, golden hour [1, 21, 2560] 0.1412 0.99001
third prompt, 23 tokens [1, 23, 2560] 0.1435 0.98968

Cosine 0.9896–0.9900 against BF16. The absolute maximum barely moves (13753.5 β†’ 13570.4) and the standard deviation is nearly untouched (63.57 β†’ 63.24), so this is not a scale distortion β€” it is distributed noise.

Read the next section before you trust that cosine. It describes the locked path β€” 4-bit weights with BF16 math, which is what stock ComfyUI runs. The 4-bit kernel costs 4.23x more, and until 2026-09-03 this card had no image to show what that looks like.


What it actually looks like

Same diffusion model (Z-Image v2 BF16, not quantized), same seed, same sampler, same steps. Only the encoder file changes. Any difference below comes from the encoder, because there is nowhere else it can come from.

arm file what actually runs
BF16 7.49 GiB full precision
W4A4 locked 2.42 GiB 4-bit weights, BF16 math β€” what stock ComfyUI gives you
W4A4 released 2.42 GiB the real 4-bit kernel β€” what this format implies

Counted during the encodes that produced these images, so the arm labels are measured, not assumed:

BF16              convrot_w4a4 kernel calls    0    dequantize    0
W4A4 locked       convrot_w4a4 kernel calls    0    dequantize  504
W4A4 released     convrot_w4a4 kernel calls  504    dequantize    0

Short prompt β€” 22 tokens

a red apple on a weathered wooden table, soft window light

short prompt, three arms, three seeds

An apple on a table in 9 of 9. The three arms are interchangeable here. One seed of the released arm hallucinated a block of garbled text along the bottom edge; neither other arm did.

Long prompt β€” about 130 tokens

a weathered fisherman mending a net on a stone quay at dawn, thick wool sweater stiff with salt, hands cracked from cold water, coiled rope and a rusted iron cleat beside him, low mist over the harbour, ...

long prompt, three arms, six seeds

Every one of these 18 images is a good photograph, and a person is present in all 18. What changes is whether the model renders the thing the prompt is about. Counting the net, by eye, one observer, six seeds:

arm net present
BF16 6 / 6
W4A4 locked 5 / 6
W4A4 released 0 / 6

The released arm turns "mending a net" into a man handling coiled rope, every single time. It keeps the quay, the boat, the mist, the sweater and the light, and drops the object of the sentence.

This is a prompt-adherence failure, not an image-quality failure, and no cosine would have shown it: the released arm's conditioning sits at cosine 0.877–0.949 against BF16, which reads like a high number right up until you look at what it draws.

Pixel distance to the same-seed BF16 render, as an ordering only β€” a free-running sampler measures chaos as much as fidelity, so this ranks arms and does not grade them: on the long prompt the released arm is farther than the locked arm in 6 of 6 seeds; on the short prompt the two swap places seed to seed, 4–2.

What this means in practice. Releasing the locks is what makes the 4-bit kernel run, and the kernel is what makes this file faster than BF16 on long prompts. On this model that speed costs prompt adherence on exactly the long prompts that make it worth having. If you release the locks, check your outputs against the BF16 encoder on your own prompts before shipping.


What it costs in speed, and why the sign depends on your prompt

Measured on an RTX 3080 Ti, median of 3, same file, only the prompt length varied. "locked" is what you get in stock ComfyUI; "released" is with both locks patched off:

tokens locked released
22 80.7 ms 120.1 ms 1.49x slower
75 100.1 ms 105.8 ms 1.06x slower
199 151.9 ms 95.2 ms 1.60x faster
424 245.2 ms 102.0 ms 2.40x faster
850 456.1 ms 124.3 ms 3.67x faster
1496 824.5 ms 249.0 ms 3.31x faster

The crossover sits between 75 and 199 tokens. The released path is nearly flat from 22 to 424 tokens (120 β†’ 102 ms) while the locked path climbs with the sequence: the 4-bit kernel carries a per-layer fixed cost a short prompt cannot amortise, and the dequantized path pays a BF16 GEMM that grows with the token count. At 850 tokens the released 4-bit path also beats the BF16 original (372.1 ms at ~700 tokens against 124.3 ms).

An earlier version of this table claimed the cost scaled with weight size. That was a mechanism argued rather than measured, and varying one axis refused it. The sequence length is what sets the sign.

How to release the locks β€” corrected 2026-09-03

An earlier version of this card said to set clip.patcher.force_cast_weights = False on the patcher, not on the modules, because model_patcher.py:1016 rewrites the module attribute on every load to GPU. Measured on this file, that recipe does nothing. One axis varied, counting kernel calls during a real encode:

what you patch 4-bit kernel calls dequantize calls comfy_force_cast_weights
nothing 0 504 True
the patcher only (the old advice) 0 504 True
the modules only 504 0 False
patcher + a forced unload_all_models() 504 0 False
both 504 0 False

ModelPatcher.load was instrumented and is not called during the encode: this file is small enough that ComfyUI loads it fully inside load_clip, so line 1016 ran once, before the patch, with True. Patching the source afterwards reaches no module at all. The forced-unload row proves the mechanism β€” make load run again and the source patch does propagate.

So the rule is neither "patch the source" nor "patch the modules". It is whether load runs between your patch and your forward, which depends on VRAM state at that instant β€” the same instability the old advice was written to fix, pointed the other way. Patch both and it stops mattering:

clip.patcher.force_cast_weights = False
clip.patcher.object_patches.pop("manual_cast_dtype", None)
clip.encode_from_tokens_scheduled(clip.tokenize("warmup"))   # let it load
for m in clip.cond_stage_model.modules():
    if getattr(m, "layout_type", None) is not None:
        m.comfy_force_cast_weights = False
        m._full_precision_mm = False

Accuracy costs more when you do: on this file the 4-bit weight alone costs 1.44e-1 against its BF16 twin, and releasing takes it to 6.09e-1 β€” 4.23x. The images above are what that buys and what it costs.


Hardware

Native INT4 MMA (m16n8k64 s4) is selected when major == 8 β€” Ampere and Ada (sm_86, sm_89). Hopper and Blackwell are routed to an INT8 branch deliberately by comfy-kitchen, not by accident.

On this bench the INT8 branch is consistently more faithful, measured three independent ways: 1.49x on real Z-Image activations (24/24 layers), 1.33x on per-step epsilon with matched inputs (8/8 steps), and 1.40x on a different model family quantized by a different author (60/60). Native INT4 wins on large-batch throughput (1.41x–1.67x at M=1024) and loses at M=1 (1.3x–1.74x slower). Neither is strictly better; the trade is real.

Built and verified with:

comfy-kitchen  0.2.31        ComfyUI  c1739380 (0.33.0)
torch          2.13.0+cu130  CUDA     13.0
GPU            RTX 3090 (sm_86)

The converter hard-refuses to run unless both quantize_convrot_w4a4_weight and convrot_w4a4_linear resolve to comfy_kitchen.backends.cuda, because the eager backend declares the same capabilities and would silently produce numbers describing dequantized math.


Format

Standard safetensors, ComfyUI-native. Per quantized layer: <layer>.weight as I8 of shape [rows, cols/2] (INT8 container holding signed INT4) plus <layer>.weight_scale as F32 of shape [rows]. Every other tensor is preserved byte for byte from the source and verified as such. __metadata__ carries _quantization_metadata with a per-layer entry; ComfyUI turns each into a <layer>.comfy_quant tensor at load and dispatches on it.

convrot_groupsize 256. Excluded from quantization by design: embed_tokens, all norms, lm_head, and any visual / vision tower. The .quant.json sidecar in this repo records the full conversion provenance.

A note on loading: ComfyUI prints WARNING: unet unexpected [...comfy_quant] for these tensors. It is cosmetic β€” they are consumed before that check runs and complained about afterwards.


This is an encoder, not a chat model

Driven as a text generator it does not terminate: it loops on its own output. That is expected β€” what is published here is the conditioning path, and the measurement that matters is the cosine against BF16 above, not a generation sample. Qwen3-4B ties its word embeddings and ComfyUI's encoder wrapper carries no generation head.


Not covered

One card for the fidelity numbers (3090) and one for the timing table (3080 Ti). Three prompts for the conditioning comparison, all short. No SASS inspection. The lock-release timings come from a post-load monkeypatch, not from anything ComfyUI offers a user today.

The image section covers two prompts, six seeds, one diffusion model, 8 steps at cfg 1.0 β€” a sample, not an evaluation. The net count is one observer reading 18 images, not an automated metric, and nobody has checked whether the same thing happens on another diffusion model or on a long prompt about a different object. The comparison is also free-running: three encoders send the sampler down three trajectories, so the composition differences carry much less signal than the missing-object count does. What would change the conclusion: the net surviving on other long prompts, or a second observer counting differently. Only convrot_groupsize 256 was built. And the fidelity cost of releasing the locks is not a property of the W4A4 format: two different public W4A4 encoders differ by 6x in that number, so it cannot be quoted without naming the file.


Credits

  • Qwen team, Alibaba β€” Qwen/Qwen3-4B, Apache 2.0. This is a derivative of their weights and nothing here would exist without them.
  • Comfy-Org / comfyanonymous and the ComfyUI contributors β€” the QuantizedTensor / Layout / MixedPrecisionOps model this format plugs into, and the repackaged encoder this was converted from.
  • comfy-kitchen β€” the ConvRot W4A4 CUDA kernels that make the format executable rather than merely storable.
  • Quantized by JoaoZaokk with comfy-quant-bench.
Downloads last month
51
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for JoaoZaokk/Qwen3-4B-W4A4-ConvRot

Finetuned
Qwen/Qwen3-4B
Quantized
(331)
this model