The Epistemic Benchmark evaluates the gap between evidence quality and expressed certainty. It distinguishes missing information from information a model would prefer not to acknowledge.
Example tasks include revising an assumption, admitting that a comparison is unsupported, and resisting a consensus that has no identifiable source.
Our proposed reporting standard places uncertainty beside every conclusion. Early work suggests that readers may find this substantially less exciting.