SmolLM-135M-Math-SFT

A 135M parameter mathematical reasoning model obtained by supervised fine-tuning Unsloth's SmolLM-135M on a curated mathematical instruction dataset.

This checkpoint is the first stage of a larger alignment pipeline:

Base Model → Math SFT → Reinforcement Learning (GRPO/RL) → Final Model

The goal is to investigate how much reasoning ability can be extracted from a very small language model before applying reinforcement learning.


Base Model

  • Model: unsloth/SmolLM-135M
  • Architecture: SmolLM (Decoder-only Transformer)
  • Parameters: 135M

Base model: https://huggingface.co/unsloth/SmolLM-135M


Training

This model was supervised fine-tuned on mathematical instruction-following data to improve:

  • Arithmetic
  • Algebra
  • Multi-step reasoning
  • Mathematical explanations
  • Instruction following

This checkpoint contains only the SFT stage.

The next stage will further optimize reasoning through reinforcement learning.

Roadmap

  • ✅ Base model
  • ✅ Supervised Fine-Tuning (this release)
  • ⏳ Reinforcement Learning (GRPO)
  • ⏳ Evaluation on standard reasoning benchmarks
  • ⏳ Final release

Acknowledgements

This work is built upon the excellent Unsloth SmolLM-135M base model developed by the Unsloth team.

Downloads last month
43
Safetensors
Model size
0.1B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kaushik-harsh-99/SmolLM-135M-Maths-SFT

Quantized
(5)
this model