Ahmetemiiii2 commited on
Commit
e06289d
·
verified ·
1 Parent(s): 81f6f92

Töz-1 (toz1-sft-g.pt)

Browse files
Files changed (3) hide show
  1. README.md +29 -6
  2. config.json +1 -1
  3. generation_config.json +1 -1
README.md CHANGED
@@ -71,10 +71,30 @@ Sistem mesajı ile eğitilmedi. Bağlam penceresi 1024 token.
71
 
72
  ## Değerlendirme
73
 
74
- Aşağıdaki setler bizim hazırladığımız ve **eğitimde hiç görülmeyen** donuk setlerdir
75
- (veri setinde `olcum/`, puanlayıcılar `uretim/` altında). Genel Türkçe benchmark'larda
76
- 195M bir modelin skoru şans seviyesine yakın olduğundan, modelin gerçekten öğrenmesini
77
- istediğimiz davranışları ölçen setler kullandık.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
78
 
79
  | set | ne ölçüyor | Töz-1 |
80
  |---|---|---|
@@ -82,7 +102,7 @@ istediğimiz davranışları ölçen setler kullandık.
82
  | kısıt (230) | talimat takibi (IFEval tarzı): biçim / tam doğru | **%75.2 / %40.4** |
83
  | tanım (90) | ham modelin bildiği ama SFT'de görmediği varlıklarda doğru kategori | **%25.6** |
84
 
85
- Tek bir eğitim koşusunun rastgeleliği bu setlerde ±3–4 puan oynatabiliyor.
86
 
87
  ## Eğitim
88
 
@@ -140,7 +160,10 @@ Töz-1 is a 195M-parameter Turkish-only language model trained **from scratch**
140
  (own tokenizer, 4.42B pre-training tokens on 2× T4 + one RTX 3060 Ti), followed by a
141
  101-second SFT on 12.3K mostly programmatically generated examples. It handles casual chat,
142
  common facts, follow-up questions, single-digit arithmetic and simple formatting
143
- instructions, and it confidently makes factual errors — treat it as a small-model
 
 
 
144
  experiment, not a knowledge source. The architecture matches Qwen3 exactly (RMSNorm,
145
  QK-norm, SwiGLU, RoPE, tied embeddings), so it loads with plain `transformers>=4.51`;
146
  the weights are **not** derived from Qwen. License: Apache-2.0.
 
71
 
72
  ## Değerlendirme
73
 
74
+ ### Genel Türkçe benchmark'lar
75
+
76
+ Hepsi aynı scriptle (`mmlu_olc.py`, `turblimp_olc.py`), 0-shot, şıkların/cümlelerin
77
+ olasılığı karşılaştırılarak ölçüldü. Karşılaştırma için ufakzeka-1-base de **aynı scriptle**
78
+ ölçüldü.
79
+
80
+ | benchmark | şans | Töz-1 (sohbet) | Töz-1-Base | ufakzeka-1-base |
81
+ |---|---|---|---|---|
82
+ | **TurBLiMP** — Türkçe dilbilgisi, 16 konu × 1.000 cümle çifti | %50 | %91.2 | **%93.2** | %92.0 |
83
+ | **Turkish MMLU** — alibayram/turkish_mmlu, 6.200 sınav sorusu, 5 şık (cloze, uzunluk normalize) | %20 | %24.1 | %25.0 | — |
84
+ | Turkish MMLU — aynı set, şık harfi (A–E) | %20 | %20.5 | %21.1 | — |
85
+
86
+ * TurBLiMP'te Töz-1-Base, 13.5B tokenla eğitilmiş ufakzeka-1-base'in (151M) önünde; Töz-1 4.42B
87
+ tokenla eğitildi. Dilbilgisi az veriyle doyan bir yetenek; dünya bilgisi gerektiren testlerde
88
+ daha çok token gören modellerin önde olması beklenir.
89
+ * Turkish MMLU (KPSS, TUS, ALES vb.) 195M bir model için çok zor: sonuç şansın üstünde
90
+ (6.200 soruda ±0.5 puan gürültü) ama düşük. Şık harfi biçimi küçük modellerde şans
91
+ seviyesinde kalır. ufakzeka-1 bu seti bizim scriptimizle ölçülmediği için "—".
92
+ * SFT, bilgi testinde (MMLU) skoru korudu; dilbilgisi testinde 2 puan düşürdü.
93
+
94
+ ### Kendi donuk setlerimiz
95
+
96
+ Eğitimde hiç görülmeyen, modelin öğrenmesini istediğimiz davranışları ölçen setler
97
+ (veri setinde `olcum/`, puanlayıcılar `uretim/` altında):
98
 
99
  | set | ne ölçüyor | Töz-1 |
100
  |---|---|---|
 
102
  | kısıt (230) | talimat takibi (IFEval tarzı): biçim / tam doğru | **%75.2 / %40.4** |
103
  | tanım (90) | ham modelin bildiği ama SFT'de görmediği varlıklarda doğru kategori | **%25.6** |
104
 
105
+ Tek bir eğitim koşusunun rastgeleliği bu küçük setlerde ±3–4 puan oynatabiliyor.
106
 
107
  ## Eğitim
108
 
 
160
  (own tokenizer, 4.42B pre-training tokens on 2× T4 + one RTX 3060 Ti), followed by a
161
  101-second SFT on 12.3K mostly programmatically generated examples. It handles casual chat,
162
  common facts, follow-up questions, single-digit arithmetic and simple formatting
163
+ instructions, and it confidently makes factual errors. On TurBLiMP (Turkish grammar
164
+ minimal pairs) the base model scores 93.2% vs 92.0% for ufakzeka-1-base (13.5B tokens), both
165
+ measured with the same script; on Turkish MMLU (6,200 exam questions) it scores 25.0%
166
+ (chance 20%). Treat it as a small-model
167
  experiment, not a knowledge source. The architecture matches Qwen3 exactly (RMSNorm,
168
  QK-norm, SwiGLU, RoPE, tied embeddings), so it loads with plain `transformers>=4.51`;
169
  the weights are **not** derived from Qwen. License: Apache-2.0.
config.json CHANGED
@@ -46,7 +46,7 @@
46
  },
47
  "sliding_window": null,
48
  "tie_word_embeddings": true,
49
- "transformers_version": "5.2.0",
50
  "use_cache": true,
51
  "use_sliding_window": false,
52
  "vocab_size": 32768,
 
46
  },
47
  "sliding_window": null,
48
  "tie_word_embeddings": true,
49
+ "transformers_version": "5.17.0",
50
  "use_cache": true,
51
  "use_sliding_window": false,
52
  "vocab_size": 32768,
generation_config.json CHANGED
@@ -7,5 +7,5 @@
7
  "repetition_penalty": 1.1,
8
  "temperature": 0.7,
9
  "top_p": 0.9,
10
- "transformers_version": "5.2.0"
11
  }
 
7
  "repetition_penalty": 1.1,
8
  "temperature": 0.7,
9
  "top_p": 0.9,
10
+ "transformers_version": "5.17.0"
11
  }