A multi-document benchmark for evaluating LLMs on three forms of scientific reasoning (deduction, induction, causal abduction), with parametric control over inference complexity and premise obfuscation. - View it on GitHub
Star
8
Rank
1892176