Foundation Models for Oversight
Cross-posted from the Transluce blog. To oversee an AI model, wed ideally like to ask questions such as: What are important situations where the model sandbags? Does the model have an objective it wouldnt admit to if asked directly? Does the model treat a user differently once it infers something about their identity, and along what axis? Is the models chain of thought load-bearing, or is it a post-hoc rationalization of an answer that was already settled on? Is it reward hacking on this input…
Read the full story at LessWrong ↗