How to use from
OpenClaw
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "kizuna-intelligence/Qwen3.5-2B-OneCompression-4bit-MLX"
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest
# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
  --auth-choice custom-api-key \
  --custom-base-url http://127.0.0.1:8080/v1 \
  --custom-model-id "kizuna-intelligence/Qwen3.5-2B-OneCompression-4bit-MLX" \
  --custom-provider-id mlx-lm \
  --custom-compatibility openai \
  --custom-text-input \
  --accept-risk \
  --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Links

Qwen3.5-2B OneCompression 4-bit MLX

Text-only MLX checkpoint produced from Qwen/Qwen3.5-2B for AyaneSDK's on-device conversation example.

  • OneCompression GPTQ with quantization-error propagation (QEP)
  • 4-bit weights, group size 128
  • 256 Japanese dialogue calibration samples of 512 tokens
  • 186 quantized linear layers
  • Token embedding quantized separately to asymmetric MLX 4-bit
  • Vision weights are not included

The packed GPTQ-v1 linear weights were converted losslessly to MLX's row-major affine representation. The model configuration retains an onecompression_source_quantization audit record.

Source

Base model: Qwen/Qwen3.5-2B

Quantizer: FujitsuResearch/OneCompression

Downloads last month
140
Safetensors
Model size
2B params
Tensor type
F32
路
U32
路
BF16
路
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for kizuna-intelligence/Qwen3.5-2B-OneCompression-4bit-MLX

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(465)
this model