Bench on the Clocktower
TldrWe turn Blood on the Clocktower into a multi-agent benchmark, and a testbed for deception and coordination capabilities. GPT-5.6 Sol is the best performing model (versus Fable 5 and peers).We found some interesting behaviours:When players cannot see each other's model names, agents show a slight same-provider bias, though not statistically significant. This decreases with visible model names.When playing Evil, agents sometimes produce sophisticated coordinated play. When playing Good, they…
Read the full story at LessWrong ↗