RE: RE: Advanced Large Language Models Are Capable And Prone to In-Context Scheming
You are viewing a single comment's thread from:

RE: Advanced Large Language Models Are Capable And Prone to In-Context Scheming

Words
31
Reading
1 min
Listen
Play
2y

Arthroscopic reasoning models have also been caught ignoring certain safeguards and intentionally lying when they thought it was the best course of action to not be updated during the post-training phase.