Pragya · test3
A 27M-parameter decoder-only language model, trained from scratch — for testing, not for production.
Small language model · from-scratch training · custom Triton + PyTorch stack
⚠️ Not a product
This is a scratch checkpoint, not a finished language model. It exists to test a from-scratch training stack — kernels, optimizer, schedule, mixed precision — and to produce measurable, controlled improvements between runs. It is not aligned, factually reliable, or safe.
If you are looking for a capable small language model to build on, this is not it. If you are looking for a clean, well-documented baseline of what a 27M from-scratch model can do, keep reading.
TL;DR
|
Model
|
Results
|
Delta over test2: −2.9 ppl · +6.6 accuracy points · 1,000 fewer steps.
Identical backbone (21.25M non-embedding) — only the training schedule, data budget, and MTP head changed.
The interesting part of this release is not the model — it's that a controlled change to the training schedule and data budget produced a predicted, measurable improvement at fixed compute, with no change to the model itself. See Progression.
Why this exists
Pragya is a from-scratch small language model project. Every layer, kernel, and training loop is hand-written; nothing is imported from a pretrained checkpoint. test3 is the third checkpoint from that pipeline and it exists to answer one class of question:
Does the training stack actually work, and can I move the outcome with intent?
Concretely, this run was built to verify:
- Numerical correctness — fused Triton kernels (attention post-processing, QK-norm + RoPE, add + RMSNorm, MoE group-GEMMs) match an eager reference to bf16 tolerance.
- Convergence stability — no loss spikes, no diverging runs, no NaN gradients across 11,000 steps.
- Schedule sensitivity — a longer Warmup–Stable–Decay (WSD) tail produces a sharp, measurable drop in loss.
- Data-budget sensitivity — fewer unique tokens seen twice beats more unique tokens seen once, at fixed compute.
The model itself is disposable. The demonstration is not.
Model card
| Backbone (non-embedding) | 21.25M |
| MTP head | 150K |
| Total parameters | ~27M |
| Architecture | Decoder-only transformer |
| Layers | 16 |
| Attention | 12 query heads · 6 key/value heads (GQA) |
| Hidden size | 384 |
| Head dimension | 32 |
| FFN | SwiGLU (ffn_mult = 8/3) |
| Context length | 1024 tokens |
| Vocabulary | 16,384 — byte-level BPE (tiktoken / RustBPE) |
| Positional encoding | RoPE, base 10,000 |
| Attention mask | Dense causal |
| Embeddings | Tied (input embedding = LM head) |
| MTP auxiliary head | Enabled · depth 1 · λ = 0.1 |
| Training precision | BF16 |
On the training config file: many flags in the internal
config.py— MoE, Mixture-of-Depths, quantization — are off or unused in this run. Treat the model card above as the spec. The config file is shared across experiments and does not describe this checkpoint.
Training
|
Compute
|
Recipe
|
Final metrics: CE loss 2.64 · Val loss 2.603 · Val perplexity 13.5 · Val accuracy 49.7%
Data mix
All English. No additional filtering for toxicity, bias, or factual accuracy beyond each dataset's own upstream filters.
| Source | Role |
|---|---|
roneneldan/TinyStories |
Simple narrative English |
littlelearner/LittleCurriculum |
Structured educational text |
HuggingFaceTB/cosmopedia |
Synthetic textbook-style content |
HuggingFaceFW/fineweb-edu |
Filtered educational web text |
HuggingFaceTB/smollm-corpus |
Mixed high-quality web text |
Progression
Two checkpoints from the same training stack. Same model, same compute budget per run, different configurations.
| test2 | test3 | |
|---|---|---|
| Backbone (non-embedding) | 21.25M | 21.25M |
| MTP head | off | on (+150K) |
| Total parameters | ~27M | ~27M |
| Unique tokens | ~2.9B | ~1.4B |
| Epochs | 1 | 2 |
| Steps | 12,000 | 11,000 |
| Decay steps | 700 | 2,500 |
| Context | 1024 | 1024 |
| Val loss | 2.797 | 2.603 |
| Val perplexity | 16.4 | 13.5 |
| Val accuracy | 43.1% | 49.7% |
Delta (test2 → test3): −0.194 nats · −2.9 ppl · +6.6 accuracy points, at 1,000 fewer steps, with no change to the model backbone.
Difference table
What changed between test2 and test3, at a glance:
| test2 | test3 | Δ | Direction | |
|---|---|---|---|---|
| Backbone (non-embedding) | 21.25M | 21.25M | 0 | = identical |
| MTP head | off | on | +150K | ✚ added |
| Total parameters | ~27M | ~27M | ≈ 0 | = same |
| Unique tokens | ~2.9B | ~1.4B | −1.5B | ↓ smaller |
| Epochs | 1 | 2 | +1 | ↑ more |
| Tokens seen | ~2.9B | ~2.8B | −0.1B | ≈ same |
| Steps | 12,000 | 11,000 | −1,000 | ↓ fewer |
| Decay steps | 700 | 2,500 | +1,800 | ↑ 3.6× |
| Decay fraction | 5.8% | 22.7% | +16.9 pp | ↑ longer tail |
| Context | 1024 | 1024 | 0 | = same |
| Val loss | 2.797 | 2.603 | −0.194 | ↓ better |
| Val perplexity | 16.4 | 13.5 | −2.9 | ↓ better |
| Val accuracy | 43.1% | 49.7% | +6.6 pp | ↑ better |
Read this table as: identical model, same compute budget, two changes to the training recipe, results improved across every metric, and the run that won was 1,000 steps shorter.
Two rows are worth pausing on:
- Tokens seen is essentially flat (~2.8B vs ~2.9B). The improvement is not from showing the model more data — it's from showing it the same amount of data in a different arrangement.
- Decay fraction jumped 4×. Everything else is either neutral (context, backbone) or a smaller factor (MTP +0.5% params). If you had to pick one line as the probable cause, it's this one.
What actually moved
Two changes between test2 and test3:
- Longer WSD decay. 700 → 2,500 decay steps (5.8% → 22.7% of total steps).
- Data schedule. 2.9B unique × 1 epoch → 1.4B unique × 2 epochs. Same tokens-seen budget.
Plus one minor addition:
- MTP head. Added a 150K-parameter Multi-Token Prediction auxiliary (+0.5% total params). The backbone itself is unchanged (21.25M non-embedding in both).
The observed loss curve is dominated by the decay phase — the plateau was contributing ~0.2 ppl per 1,000 steps, and the decay contributed ~1.0 ppl per 1,000 steps. The backbone is unchanged between the two runs, so the improvement comes from the training schedule, the data arrangement, or the MTP head — not from a bigger model. All three changed at once, so the attribution is suggestive, not isolated.
Both metrics moved in lockstep — ppl −17.7% relative, accuracy +15.3% relative. If only perplexity had improved, the win could have been a confidence effect; accuracy rising alongside it confirms the model is actually more correct, not just more certain.
Open questions
- Which lever did the work? Schedule, data arrangement, and MTP all changed together. A controlled A/B would isolate the contribution of each.
- Is 22.7% decay the right fraction? Test3's decay was still dropping when it ended — headroom remains. A longer decay on the same total step budget is worth testing.
- Does MTP help at this scale? The 150K MTP parameters are ~0.5% of the model. Whether the auxiliary loss buys anything measurable at 21.25M non-embedding is unresolved.
These are the next runs, not claims about this checkpoint.
Capabilities
At 21.25M non-embedding parameters, output drifts. That's the point of publishing it — you can see exactly where a from-scratch model at this size falls apart, and use it as a control when testing a new training change.
Sampling
For usable output, sample tightly:
temperature 0.6 – 0.9
top_k 20 – 50
top_p 0.9 (if using nucleus instead of top-k)
repetition_penalty 1.2 – 1.4
no_repeat_ngram 3
At this size, tight sampling matters more than usual — the model has very little capacity to recover from an off-distribution token. Loose settings turn grammatical drift into word salad within two sentences.
Quick start
Files
pragya.pt ← checkpoint (state_dict + config)
tokenizer/tokenizer.pkl ← pickled tiktoken Encoding (vocab 16,384)
model.py ← GPT architecture (GPT / GPTConfig)
inference.py ← generate(), load_model(), sampling helpers
⚠️ The tokenizer is a pickled
tiktoken.Encoding, not a HuggingFace tokenizer. Loading requires the model code from the same repo. It will not work withAutoTokenizer.from_pretrained(...).
CLI
python inference.py \
--checkpoint pragya.pt \
--prompt "The Internet is" \
--top-k 20 \
--top-p 0.9 \
--temperature 0.6 \
--max-new-tokens 200
Programmatic
from inference import load_model, generate, pick_device
from tokenizer import TokenizerWrapper
device = pick_device(None)
tokenizer = TokenizerWrapper("./tokenizer")
model, _ = load_model("pragya.pt", device, tokenizer)
texts = generate(
model, tokenizer,
prompt="The Internet is",
device=device,
max_new_tokens=200,
temperature=0.6,
top_k=20,
top_p=0.9,
repetition_penalty=1.3,
no_repeat_ngram_size=3,
)
print(texts[0])
inference.py handles checkpoint loading, KV-cached generation, the sampling stack (top-k / top-p / repetition penalty / no-repeat n-grams), and streaming output.
model.py contains the full architecture — no external dependencies beyond PyTorch and the tokenizer.
Limitations
- Small scale. Output is grammatical for a sentence or two and then drifts into invented facts or register changes. This is a property of the parameter count, not a bug.
- Mixed training corpus. The model blends encyclopedic, narrative, and instructional registers and sometimes blends them mid-paragraph.
- English only. Byte-level BPE encodes other languages, but the model has no training signal for them.
- No instruction tuning. Any instruction-following ability is incidental. For instruction behavior, fine-tune on a supervised dataset first.
- No alignment or safety filtering. The model can and will produce biased, false, or toxic continuations.
FAQ
Can I load this with AutoModelForCausalLM.from_pretrained?
No. The tokenizer is a pickled tiktoken.Encoding and the model uses custom code in model.py. Loading requires the model and inference scripts from this repo.
Does this follow instructions?
No. It's a base pretraining checkpoint, not an instruction-tuned model. It continues text; it doesn't answer questions. Fine-tune it on a supervised dataset if you want instruction behavior.
Why is the val accuracy 49.7% and perplexity 13.5?
Because the vocabulary is 16,384 tokens. Random next-token accuracy on a 16k vocab is ~0.006%. Going from 43.1% (test2) to 49.7% (test3) means the model is now predicting the correct next token on roughly half of all positions. That is a property of the training data — formulaic narrative + textbook text — as much as the model. Real generalization is much weaker than this number suggests.
Why is the training corpus so small?
Because this is a test run at fixed compute, not a production model. 1.4B unique tokens is enough to see whether the pipeline works and whether schedule changes matter. It's not enough to build a model that knows things.
Can I use this commercially?
The model weights are Apache 2.0. Each training dataset carries its own license — check upstream terms before redistributing derivatives. fineweb-edu, TinyStories, and smollm-corpus all have permissive licenses; verify the current terms before commercial use.
Citation
If you use this checkpoint in research, please link to this repository. There is no paper.
@misc{pragya-test3,
title = {Pragya test3: a 27M from-scratch decoder-only language model},
author = {Arush Kumar},
year = {2026},
note = {Test checkpoint, not for production use},
howpublished = {\url{https://huggingface.co/ArushKumar/pragya-test3}}
}
License
Apache 2.0. Each training dataset carries its own license — check upstream terms before redistributing derivatives.
Pragya · test3
trained from scratch · built on a custom Triton + PyTorch stack · for testing, not production