A procedurally generated video benchmark for epistemic restraint in vision-language models: each physics scenario pairs an answerable video with an unanswerable one, and the PECS metric scores a model only when it answers the knowable cases and abstains on the unknowable ones. - View it on GitHub
Star
0
Rank
14369279