Some personal experiments around routing tokens to different autoregressive attention, akin to mixture-of-experts - View it on GitHub
Star
115
Rank
231497