A procedurally generated video benchmark for epistemic restraint in vision-language models: each physics scenario pairs an answerable video with an unanswerable one, and the PECS metric scores a model only when it answers the knowable cases and abstains on the unknowable ones. -
View it on GitHub