Post
2387
20K parameters can tell a story. ๐
๐ค We trained a ~20k-parameter Transformer that can actually write stories!
raincandy-u/MacroStories
โ ~50ร smaller than the 1M-parameter TinyStories model
โ ~3,000ร smaller than AlexNet
โ 81 KB in FP32
yayyy the whole model. เซฎ หถแต แต แตหถ แ
She has a 32-dimensional hidden state, a 378-token vocabulary, and just one decoder block โ recurrently applied 4 times with shared weights.
Despite having only 19,969 parameters, she can maintain a simple narrative across 100โ300 words: establish a goal, encounter a problem, take relevant actions, and reach an outcome.
She runs extremely fast on CPU โ no GPU required. The entire model is tiny enough to load almost instantly! โบ๏ธ
๐ค We trained a ~20k-parameter Transformer that can actually write stories!
raincandy-u/MacroStories
โ ~50ร smaller than the 1M-parameter TinyStories model
โ ~3,000ร smaller than AlexNet
โ 81 KB in FP32
yayyy the whole model. เซฎ หถแต แต แตหถ แ
She has a 32-dimensional hidden state, a 378-token vocabulary, and just one decoder block โ recurrently applied 4 times with shared weights.
Despite having only 19,969 parameters, she can maintain a simple narrative across 100โ300 words: establish a goal, encounter a problem, take relevant actions, and reach an outcome.
She runs extremely fast on CPU โ no GPU required. The entire model is tiny enough to load almost instantly! โบ๏ธ