NeoHorse-1-4B-GGUF

NeoHorse-1-4B is a 4-billion-parameter causal language model from TokenRhythm, post-trained from Qwen3.5-4B as an initial prototype on the path toward recursive self-improvement (RSI), targeting text-based agent harnesses, tool use, coding, and instruction following. Like its 9B sibling, its core innovation is a routing harness that assigns tasks to a heterogeneous model pool, records tool interactions and outcomes, estimates capability demand, and feeds that signal back into shaping the next training mixture via routing-guided curriculum SFT and on-policy distillation, backed by rigorous deduplication, decontamination, and six-dimensional semantic data evaluation. This release contains language-model weights only (vision weights excluded, repackaged for text-only inference) and retains a 262,144-token native context window extensible to 1,010,000. Across a ten-benchmark evaluation against Qwen3.5-4B, Gemma-4-E4B-it, Nanbeige-4.2-3B, Agents-A1-4B, and Spark-X2.5-4B, NeoHorse-1-4B posts the best overall macro average (64.87 vs. 58.94 for its Qwen3.5-4B base, a +5.93 gain), with the largest improvements on agentic benchmarks like VitaBench (+10.50), WorkBuddy Bench (+9.79), HumanEval (+9.75), and QwenClawBench (+6.21), alongside a strong tau2-Bench score of 88.46 (best in the comparison set). It's servable via SGLang or vLLM with Qwen3-style reasoning and tool-call parsers, and is released under the Apache License 2.0.

Model Files

File Name Quant Type File Size File Link
NeoHorse-1-4B.BF16.gguf BF16 8.42 GB Download
NeoHorse-1-4B.Q3_K_L.gguf Q3_K_L 2.42 GB Download
NeoHorse-1-4B.Q3_K_M.gguf Q3_K_M 2.26 GB Download
NeoHorse-1-4B.Q3_K_S.gguf Q3_K_S 2.07 GB Download
NeoHorse-1-4B.Q4_0.gguf Q4_0 2.54 GB Download
NeoHorse-1-4B.Q4_K_M.gguf Q4_K_M 2.71 GB Download
NeoHorse-1-4B.Q4_K_S.gguf Q4_K_S 2.56 GB Download
NeoHorse-1-4B.Q5_0.gguf Q5_0 2.99 GB Download
NeoHorse-1-4B.Q5_K_M.gguf Q5_K_M 3.07 GB Download
NeoHorse-1-4B.Q5_K_S.gguf Q5_K_S 2.99 GB Download

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
2,356
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/NeoHorse-1-4B-GGUF

Finetuned
Qwen/Qwen3.5-4B
Quantized
(13)
this model

Collection including prithivMLmods/NeoHorse-1-4B-GGUF