3 Comments
User's avatar
Geordie Korper's avatar

Who watches the watcher? The model’s escape route then is to slowly corrupt the less advanced observer model and have do the work. Alignment and defense in depth are probably the only answer. A tripwire on the cache server and limiting what it could access would have helped. The problem you often see is the person setting up the test environment is not highly trained in locking down outbound access but they become when something like this happens.

Ryan Baker's avatar

As a theoretical attack surface, yes. Practically speaking though, corrupting a watcher instance would be very difficult. Whatever sandbox you use for the evaluation model instance could be applied to the watcher instance. If it's less advanced than the evaluation model, it is even less likely to be able to find a way out. It thus seems unlikely the evaluation model gains risk from the addition of the watcher. It's interactions with a watcher are more likely to be trying to obfuscate itself from it rather than corrupting it.

I'd be cautious in the security domain with ever naming one or two solutions as the only answer. Also, observer models are defense in depth. And yes, if you're concerned about the watcher being subverted, watch the watcher, and so on for as many levels as seems necessary/cost effective. I think it won't make much sense to do that until you're worry is subversion rather than obfuscation. A second watcher has relatively little hope of detecting an obfuscation the first missed; much better hope of detecting ongoing attempts to subvert. It's very hard to hide attempts to subvert as the subverting text has to be consumed to accomplish the objective.

Ryan Baker's avatar

I'm becoming more convinced this is a big blind spot, and I'm not sure why. The concept "monitor", comes up here, but is swiftly swept under the rug with "I can see why they didn't think of that".

https://blog.redwoodresearch.org/p/the-openaihuggingface-incident-redwood?utm_campaign=post&utm_medium=web&t=2926.0&d=30.0