Model Hermeneutics: Monitoring Closed-Weight Models with Open-Weight Internals
LessWrong
Read the full articleResearchers propose a method called Model Hermeneutics to monitor closed-weight AI models by analyzing open-weight internal components, enabling better oversight without direct access to full model weights.
