Systematic evaluation framework that automatically rates overthinking behavior in large language models. - View it on GitHub
Star
4
Rank
2858665