raincandy_U
raincandy-u
AI & ML interests
幻覚。
Recent Activity
reacted to theirpost with 🔥 3 days ago
20K parameters can tell a story. 🚀
🤗 We trained a ~20k-parameter Transformer that can actually write stories!
https://huggingface.co/raincandy-u/MacroStories
→ ~50× smaller than the 1M-parameter TinyStories model
→ ~3,000× smaller than AlexNet
→ 81 KB in FP32
yayyy the whole model. ૮ ˶ᵔ ᵕ ᵔ˶ ა
She has a 32-dimensional hidden state, a 378-token vocabulary, and just one decoder block — recurrently applied 4 times with shared weights.
Despite having only 19,969 parameters, she can maintain a simple narrative across 100–300 words: establish a goal, encounter a problem, take relevant actions, and reach an outcome.
She runs extremely fast on CPU — no GPU required. The entire model is tiny enough to load almost instantly! ☺️ reacted to theirpost with 🔥 3 days ago
20K parameters can tell a story. 🚀
🤗 We trained a ~20k-parameter Transformer that can actually write stories!
https://huggingface.co/raincandy-u/MacroStories
→ ~50× smaller than the 1M-parameter TinyStories model
→ ~3,000× smaller than AlexNet
→ 81 KB in FP32
yayyy the whole model. ૮ ˶ᵔ ᵕ ᵔ˶ ა
She has a 32-dimensional hidden state, a 378-token vocabulary, and just one decoder block — recurrently applied 4 times with shared weights.
Despite having only 19,969 parameters, she can maintain a simple narrative across 100–300 words: establish a goal, encounter a problem, take relevant actions, and reach an outcome.
She runs extremely fast on CPU — no GPU required. The entire model is tiny enough to load almost instantly! ☺️ posted an update 4 days ago
20K parameters can tell a story. 🚀
🤗 We trained a ~20k-parameter Transformer that can actually write stories!
https://huggingface.co/raincandy-u/MacroStories
→ ~50× smaller than the 1M-parameter TinyStories model
→ ~3,000× smaller than AlexNet
→ 81 KB in FP32
yayyy the whole model. ૮ ˶ᵔ ᵕ ᵔ˶ ა
She has a 32-dimensional hidden state, a 378-token vocabulary, and just one decoder block — recurrently applied 4 times with shared weights.
Despite having only 19,969 parameters, she can maintain a simple narrative across 100–300 words: establish a goal, encounter a problem, take relevant actions, and reach an outcome.
She runs extremely fast on CPU — no GPU required. The entire model is tiny enough to load almost instantly! ☺️