view reply Yes it’s published here https://doi.org/10.13140/RG.2.2.27323.99362 The only variable in these results is the training checkpoint.
view reply It’s reinforcement learning that causes the reasoning problems. I’ve done controlled experiments and based models reason quite effectively. Reinforcement learning disturbs the weights dramatically and reasoning is one of the most damaged results
view post Post 93 Released MAiRY Star, a companion/perosna Ai fine tuned with SFT Studio Pro dataset , no RLHF or reward optimization. Access at https://triadai.tech See translation Reply