Foundation Models for Oversight

LESSWRONGneutral2026-07-28 16:37:28 UTC
AdYour ad here[email protected]

Cross-posted from the Transluce blog. To oversee an AI model, wed ideally like to ask questions such as: What are important situations where the model sandbags? Does the model have an objective it wouldnt admit to if asked directly? Does the model treat a user differently once it infers something about their identity, and along what axis? Is the models chain of thought load-bearing, or is it a post-hoc rationalization of an answer that was already settled on? Is it reward hacking on this input…

Read the full story at LessWrong ↗
AdYour ad here[email protected]