This repository includes a benchmark and code for the paper "Evaluating LLMs at Detecting Errors in LLM Responses". - View it on GitHub
Star
0
Rank
14121396