Model is unstable due to lacking training / Model niestabilny ze względu na niewystarczające szkolenie

A fine-tuned version of Qwen2.5 3B Instruct, made for (at least understanding) Polish. The model was trained using the PolQA dataset. Available in multiple quantization variants, ranging from Q2_K_L up to F16. It should be used only for basic translation and conversations (From Polish to English and vice versa). Training was done on relatively low-end hardware (M1 chip, 8GB of unified memory, and 256GB of internal storage) for a kind of short time (an hour and 10 minutes), hence the low quality. The model can use anywhere from 1.3 to 6.3GB of RAM (depending on quantization, here Q2_K_L versus F16 with a context window of 8192).

Dostrojony wariant Qwen2.5 3B Instruct, przeznaczony do (co najmniej zrozumienia) języka polskiego. Model został przetrenowany za pomocą zbioru danych PolQA. Dostępny w kilku wariantach kwantyzacji, od Q2_K_L, aż do F16. Model powinien być używany jedynie do podstawowego tłumaczenia i rozmów (z polskiego na angielski i zamiennie). Trening był wykonany na sprzęcie o całkiem niskiej mocy obliczeniowej (czip Apple M1, 8GB zunifikowanego RAMu, 256GB pamięci wewnętrznej) przez krótki czas (godzina i dziesięć minut), stąd niska jakość. Model może używać of 1.3 do 6.3GB RAMu (zależy od kwantyzacji, tutaj Q2_K_L porównany z F16 z okienkiem kontekstu 8192).

Downloads last month
131
GGUF
Model size
3B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for piskle/Qwen2.5-3B-Instruct-Polski-GGUF

Base model

Qwen/Qwen2.5-3B
Quantized
(299)
this model

Dataset used to train piskle/Qwen2.5-3B-Instruct-Polski-GGUF