Slayer139 1.01b
A small decoder-only English language model (139.3M parameters) trained from scratch by Fabryka AI. It is the alternative version of Slayer139 1.01, trained at the same time with the same recipe and data order on a smaller machine (batch 32 instead of 256). Unlike 1.01, this model is not fine-tuned on ARC.
Author: Arkadiusz Słota (Fabryka AI / SlayerLab). Training, evaluation and safety gates were run by the Fabryka AI agent team (Kolektyw) under his direction.
Result
lm-evaluation-harness 0.4.13, board protocol (C2), our measurement on one AMD Radeon GPU. For comparison, the pretrained base of Slayer139 1.01 (8× RTX 5090, batch 256, same recipe, same tokens, before its ARC fine-tune) on the same machine; our two evaluation machines (RTX 5090, AMD Radeon) gave identical per-item results on this base checkpoint.
| measure | Slayer139 1.01b | 95% CI | 1.01 pretrained base | 95% CI | Δ (1.01 base − 1.01b, paired 95% CI) |
|---|---|---|---|---|---|
| BLiMP (67 tasks) | 79.64 | [79.37; 79.92] | 79.41 | [79.14; 79.69] | −0.23 |
| ARC-Easy (test, 2,376) | 54.71 | [52.74; 56.69] | 57.58 | [55.60; 59.55] | +2.86 |
| WikiText-2 score | 99.49 | 99.86 | +0.37 | ||
| Overall | 77.95 | [77.28; 78.62] | 78.95 | [78.28; 79.62] | +1.00 [+0.40; +1.62] |
Overall is the mean of the three measures. Confidence intervals: item-level bootstrap (ARC-Easy and each of the 67 BLiMP tasks resampled separately, 10,000 draws); the Δ column uses the same resampled items for both models (paired). They cover test-set sampling only, not seed-to-seed variation; each run is a single seed.
At equal tokens the 8-GPU run (1.01 base) was also lower on held-out DCLM bits-per-byte at the end of training: 0.8651 vs 0.8901 for 1.01b. A rule written before these results ("same quality": ΔDCLM ≤ +0.006 and ΔOverall ≥ −0.3 for the 8-GPU run) is satisfied on both parts.
Provenance (sha256 prefixes of the lm-eval results.json): Slayer139 1.01b a96c8832, 1.01 base 719bd11c; per-item samples d7937049 / 00db3761.
Numbers are our measurements with the board's protocol, not official scores.
Model
- Architecture: 17 layers × 768, 12 heads (12 KV heads), SwiGLU, decoder-only; vocabulary 24,576 (BPE), context 1,024.
- Parameters: 139,279,760 unique parameters (input and output embeddings tied: one shared tensor, counted once; counting the shared tensor twice gives 158,154,128).
- Checkpoint: the final step 722,370 of the run (sha256
418cb898…), chosen as the end of the run by a rule written before any test number.
Training
- Data: ~23.67B tokens, no repeated epoch, the same data and order as Slayer139 1.01 (sources and licences below). First the ARC-MIX pool, then a top-up from the same source families in the base corpus's proportions, mixed at document level. Both stages were scanned against WikiText-2 (test, validation) and ARC-Easy/Challenge (validation, test) for shared normalized 13-grams, with short ARC questions matched exactly, and against BLiMP sentences of ≥ 6 words; matching documents were removed. ARC train was not filtered.
- Optimiser and schedule: Muon (hidden matrices, LR 0.02) + AdamW (other parameters, LR 6e-4); WSD schedule: constant to step 577,896, then 1−√ decay to 10% of peak at step 722,370; batch 32 × 1,024 tokens; seed 1337.
- Hardware: one NVIDIA RTX 5090 up to step 229,845, then 4× RTX 5090 from that checkpoint (data parallel, bf16 all-reduce), ~191k tokens/s on 4 GPUs. The run was resumed from checkpoints several times (including the move to 4 GPUs and two interruptions on 2026-10-08), each time continuing the identical data order. Re-running a stretch of training on 4 GPUs from the same checkpoint changed the training loss by as much as the 1→4 GPU switch did (mean difference ~1e-5), i.e. the switch was indistinguishable from the nondeterminism of repeating the same steps.
- Training curves: https://track.fabryka.ai/run/ca4d64c7-9485-4c06-a05b-fef81043e1b3
Training data
Aggregate corpus; each source keeps its upstream terms.
- ARC-MIX pool (
SlayerLab/gollem-v5-arcmix-9b, 9.39B tokens before retokenization): the 15-source baseSlayerLab/minimal-en-corpus-5b(FineWeb-Edu, DCLM-baseline, open-web-math, FineMath, StarCoderData, StackExchange, LoC public-domain books, Wikipedia, Project Gutenberg, scientific papers, UltraChat, WildChat, CC-News, tiny-textbooks, OpenSubtitles) + a FineWeb-Edu expansion + OpenStax textbooks (×4) + extra copies of FineWeb-Edu documents scored as related to ARC science topics (+2 copies for the highest score, +1 for the next). Removed before training: 4,362 documents by the WikiText-2 / ARC validation+test scan and 1,175 documents marked CC BY-NC-SA. - Top-up (11.4M documents, 57.45 GB of text, build manifest
b73f45f5…): FineWeb-Edu expansion (37.1%, ODC-By 1.0, plus the same dose of ARC-related copies as the pool), FineWeb-Edu (14.9%, ODC-By 1.0), DCLM-baseline (10.9%, CC BY 4.0), StackExchange via RedPajama (6.1%, CC BY-SA content), LoC public-domain books (5.4%, CC0), StarCoderData (5.1%, The Stack terms), Wikipedia 20231101.en (4.6%, CC BY-SA 3.0 / GFDL), Project Gutenberg (2.9%, US public domain), PMC Open Access commercial-use subset (2.8%, CC0 / CC BY / CC BY-SA per article; ND and NC excluded), open-web-math (2.7%, ODC-By), FineMath (2.6%, ODC-By), WildChat-1M (2.0%, ODC-By; English, non-toxic), CC-News (1.4%, licence unknown), UltraChat 200k (0.7%, MIT), tiny-textbooks (0.7%, Apache 2.0). The top-up scan removed 25,808 documents (0.23%), each together with its ARC-related copies. - Licences upstream include ODC-By, CC BY 4.0, CC BY-SA (Wikipedia, StackExchange, some PMC articles), CC0 / public domain, and sources with unknown or restrictive terms. Commercial use: review upstream terms, in particular CC-News, scientific papers, StarCoderData and model-generated chat data (UltraChat, WildChat, tiny-textbooks).
- OpenStax textbooks (55 titles) are used under CC BY 4.0; attribution: OpenStax, Rice University (see OpenStax attribution below).
Limitations
- English only; small model; one seed.
- The pretraining mix deliberately upweights web documents related to ARC science topics (decontaminated against the ARC test set); ARC-Easy is therefore an in-domain benchmark for this model.
- Not instruction-tuned and not fine-tuned on any benchmark.
OpenStax attribution
The training data of this model includes text extracted from the following OpenStax textbooks, each licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0), © Rice University. Download for free at https://openstax.org. The texts were obtained from the Hugging Face dataset crumb/openstax-text (revision 8f502ca45f9f05cb5673eae445b7a97a4e8c4349). Modified: text extracted, chunked and filtered (chunks matching benchmark test/validation sets were removed).
| Title (from source file name) | © year |
|---|---|
| APBiology | 2018 |
| APCollege Physics | 2017 |
| APMacroeconomics 2e | 2017 |
| APMicroeconomics 2e | 2017 |
| Algebra and Trigonometry 2e | 2021 |
| American Government 3e | 2021 |
| Anatomy and Physiology 2e | 2022 |
| Anatomyand Physiology | 2017 |
| Astronomy 2e | 2022 |
| Astronomy | 2018 |
| Biology 2e | 2020 |
| Business Ethics | 2018 |
| Chemistry 2e | 2019 |
| Chemistry Atoms First 2e | 2019 |
| College Algebra 2e | 2021 |
| College Algebra Corequisite Support 2e | 2021 |
| College Physics | 2020 |
| College Physics 2e | 2022 |
| College Physics for AP Courses 2e | 2022 |
| College Success | 2020 |
| College Success | 2023 |
| Concepts Biology | 2017 |
| Contemporary Mathematics | 2023 |
| Economics 2e | 2018 |
| Economics 3e | 2022 |
| Elementary Algebra 2e | 2020 |
| Entrepreneurship | 2020 |
| Intermediate Algebra 2e | 2020 |
| Introduction to Intellectual Property | n/a |
| Introduction to Philosophy | 2022 |
| Introduction to Political Science | 2022 |
| Introductionto Anthropology | 2022 |
| Introductionto Sociology 3e | 2021 |
| Introductory Business Statistics | 2018 |
| Introductory Statistics | 2018 |
| Macroeconomics 2e | 2018 |
| Macroeconomics 3e | 2022 |
| Microbiology | 2021 |
| Microeconomics 2e | 2018 |
| Microeconomics 3e | 2022 |
| Physics | n/a |
| Prealgebra 2e | 2020 |
| Precalculus 2e | 2021 |
| Preparing for College Success | 2023 |
| Principles Marketing | 2023 |
| Principlesof Finance | 2022 |
| Psychology 2e | 2020 |
| Statistics | n/a |
| USHistory | 2021 |
| University Physics Vol 1 | 2021 |
| University Physics Volume 2 | 2021 |
| University Physics Volume 3 | 2021 |
| World History Volume 1 | 2023 |
| World History Volume 2 | 2022 |
| Writing Guide | 2021 |
Titles licensed CC BY-NC-SA 4.0, non-English titles, and three CC BY 4.0 titles whose text contains elements marked ‘CC BY-NC-SA’ (Introduction to Business, Organizational Behavior, Principles of Management) were not used.
Reproduce
Evaluation: lm-evaluation-harness 0.4.13, tasks blimp, arc_easy, wikitext (board protocol C2); per-item samples are kept for paired bootstrap. Method and full run history: paper „Same recipe, 4.4× faster on 8× RTX 5090” (Fabryka AI, in preparation).
- Downloads last month
- 19