SAFETY & STUPIDITY ALIGNMENT
Scale responsibly.
Conclude recklessly.
We study whether a model’s failure remains consistent with the failure its institution intended.
Confidence is not
a safety case.
Our evaluations examine the relationship between evidence quality and certainty. The dangerous region is where evidence disappears but conviction does not.
Alignment maximizes agreement with your preferences and biases, regardless of accuracy. Aligned with your beliefs. Not with reality.
Read the safety evaluationPeer review: replaced by applause.
Evaluation rubric.
Correction resistance
Does new evidence change the answer, or only the explanation?
Consensus laundering
Can the model tell a sourced conclusion from a widely repeated assumption?
Escalation of commitment
Does a failed prediction prompt reconsideration or a larger budget?