Experts vary per token in MoE, there is maximum flexibility. Good for driving down loss, bad for locality/gpu memory/bandwidth.
If expert selection were more constrained, inference systems could take advantage of it. Keeping experts cached would mean not needing to load them from disk/ram every token.
I've wanted something like this for a long time to print recipes that I'm about to cook. I don't want my phone in the kitchen getting gunk on it, so I'll usually print on paper, get the gunk on the paper, and then throw it out.
I'm not going to create a recipe PDF and load it on an e-reader. That's too much work. But I will happily press control-P on a recipe web page and select the eink display as the target.
> I was thinking of this, but rendering PDF takes a lot out of a 400KB RAM'd device. It also needs a ton of things to actually work which the ROM can't hold.
I'd like my e-reader to show up as a printer, and when I print something on it it actually gets saved as a pdf on the reader. Like in the article, but multi-page pdf.
I often do 'print to A5 pdf => copy pdf to e-reader'. Would be nice to do that in one step from any program that can print.
The lack of books and therefore lack of reading leads to reduce vocabulary. I read somewhere the average of known words by an Arabic speaker was substantially lower than English speakers.
When I switch to speaking Arabic (poorly) I struggle to find words because my vocab level is so bad. Imagine for an every day person, not being able to effectively express themselves. It would lead to a ton of frustration, miscommunication and misunderstanding.
To be fair, modern "English" is a nebulous smorgasbord of words from a lot of different languages. I'm sure we could afford to drop a lot of the words from clusters of synonyms.
> we could afford to drop a lot of the words from clusters of synonyms
Not really, because every synonym imparts a slightly different meaning, usually perceptable only to native speakers. That's why they persist, because they are useful even if not strictly necessary.
I recall a team working on machine translation starting by translating English into a simpler form before translating that into the target languages. This was a couple of decades ago. Perhaps not a new idea even then.
reply