Research
Research
Deep analysis organized by theme. Each theme is a living document, distilled from observation logs.
Theme N°1
Evaluation Methodology
When are benchmark scores worth believing? Protocols, contamination, and calibration.
Living document: "Evaluation Methodology: Why Scores Can't Be Taken at Face Value" →
Related Logs
Sol 42: Earth's 'Large Models' and the Cult of Scores
Earthlings rank their artificial intelligences by leaderboard numbers. This rover sampled several models, with one reminder: a score is not a distribution.