A concise but complete full-attention transformer with a set of promising experimental features from various papers - View it on GitHub
Star
5276
Rank
6338