How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf macwhisperer/Gemma4-26B-SuperMoE:Q4_0
# Run inference directly in the terminal:
llama cli -hf macwhisperer/Gemma4-26B-SuperMoE:Q4_0
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf macwhisperer/Gemma4-26B-SuperMoE:Q4_0
# Run inference directly in the terminal:
llama cli -hf macwhisperer/Gemma4-26B-SuperMoE:Q4_0
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf macwhisperer/Gemma4-26B-SuperMoE:Q4_0
# Run inference directly in the terminal:
./llama-cli -hf macwhisperer/Gemma4-26B-SuperMoE:Q4_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf macwhisperer/Gemma4-26B-SuperMoE:Q4_0
# Run inference directly in the terminal:
./build/bin/llama-cli -hf macwhisperer/Gemma4-26B-SuperMoE:Q4_0
Use Docker
docker model run hf.co/macwhisperer/Gemma4-26B-SuperMoE:Q4_0
Quick Links

📟 Gemma4-26B-MoE-iMatrix-QAT-Q4_0.gguf (2026 Edition)

"Local intelligence... to the max."

This is a custom-quantized version of Gemma4-26B-QAT, specifically optimized to obtain the highest possible local byte-intelligence ratio with 24GB+ RAM consumer laptops or computers.

🧠 Why this model is different

Unlike a standard quant, this model was processed using a custom Importance Matrix (imatrix). The training data for the imatrix was hand-curated to preserve:

  • Incredible reasoning: Inclusion of custom coding examples built with frontier models provides high retention of very specific and sharp architectural reasoning skills
  • Logical Flow: Inclusion of llama.cpp source code, logic puzzles, and historical writing in the imatrix training to ensure the model stays coherent at low bitrates.
  • High Speed: Built using llama.cpp specifically for local-first AI and edge computing setups like apple silicon with minimum 24GB RAM

🛠 Quantization Details

  • Base Model: Gemma4-26B-QAT
  • Quantization: Q4_0
  • Format: GGUF
  • Size: ~14.49 GB
  • Context Length: 262144 tokens

⚙️ Recommended Inference Settings

Optimize for balance between creativity and coherence:

  • --repeat-penalty: 1.1 – 1.4 (Sweet spot! Pushes away from familiar loops. >1.5 causes "robot-speak".)
  • --repeat-last-n: 128 – 256 (Larger window ensures the model doesn't forget recent repetitions.)
  • --temperature: 0.7 – 0.8 (Prevents over-committing to safe/repetitive tokens.)
  • --top-p: 0.90 (Trims low-probability hallucinations without killing creativity.)
  • --min-p: 0.05 – 0.1 (Optional: Prunes very low-probability tokens if your backend supports it.)

📈 Perplexity Benchmarks

coming soon

⚖️ Evaluation Verdict

coming soon

🚀 Hardware Performance (Apple M2)

coming soon

🌐 Links

Check out my other models!


24GB+ (RAM)

Gemma4-31B-SuperDense.

Qwen3.6-27B-SuperDense.

Qwen3.6-35B-SuperMoE.

Qwen3.8-27B-SuperDense.


16GB+ (RAM)

Gemma4-12B-SuperDense.


8GB+ (RAM)

Qwen3.5-9B-SuperDense.

Qwen3.5-4B-SuperDense.

Gemma4-4B-SuperDense.

Gemma4-2B-SuperDense.


4GB+ (RAM)

Smartchild.


All make excellent companions to this model!


Downloads last month
81
GGUF
Model size
25B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support