Paper · 2026
The Metanym Game: A Self-Contained, Self-Consistent LLM Peer-Community Benchmark for Structural Intelligence
A benchmark that contains its own ground truth. Language models compete at generating analogies and evaluate each other's efforts. The game requires eight kinds of intelligence. The factual intelligence ranking that falls out coincides with GPQA Diamond, a benchmark with questions and answers created by domain experts.
